CLASP: Continual Low-rank Adapters for Spatially Placed Concepts from One Hypernetwork

Wojciech Gromski1,2 · Patryk Krukowski3,4 · Jan Miksa2,3 · Maciej Zieba1,5 · Przemysław Spurek2,3
1 Wrocław University of Science and Technology  ·  2 IDEAS Research Institute  ·  3 Jagiellonian University  ·  4 AKCES NCBR  ·  5 Tooploox
Preprint
Wrocław University of Science and Technology IDEAS Research Institute Jagiellonian University AKCES NCBR Tooploox

CLASP overview and generations after fifty concepts

26.0 M
parameters, the same after 10, 50 or 100 concepts
2 × 256
numbers stored per new concept (CIDM: 0.42 M)
+14%
generation time at every length (CIDM: up to +154%)
100
concepts learned in sequence (CIDM stops at 37)

Personalizing a text-to-image diffusion model with a sequence of new concepts usually means either forgetting the earlier ones or storing a separate adapter for each, so the model grows with every concept. CLASP replaces the store with a single hypernetwork of fixed size: from a compact task embedding it generates each concept’s low-rank adapter for a frozen diffusion model, and from the same embedding and a bounding box it generates tokens that place the concept where the user asks. An output-space regularizer keeps the adapters of earlier concepts in place as new ones are learned. On the CIFC benchmark, CLASP forgets about four times less than CIDM and twenty-three times less than sequential fine-tuning, and it keeps learning to a hundred concepts at the same size and generation cost.

Concepts that stay

Drag the slider to move through the ten-concept sequence. Next to its reference photos, each column shows the same concept, generated after the given task from the same prompt and the same initial noise, by our fixed-size network, by CIDM, which stores an adapter for every concept, and by sequential fine-tuning on the same budget.

Reference photo
Reference photos
Ours
Ours
CIDM
CIDM
Fine-tuning
Fine-tuning
after task 10 of 10

Forgetting, measured

Forgetting is the drop in DINO similarity from each concept’s best score to its score at the end of the ten-concept sequence. All three methods are read at nearly the same text alignment, on SD-1.5, as the mean and standard deviation over three training seeds.

Forgetting over the ten-concept sequence Ours 0.0060, CIDM 0.0232, sequential fine-tuning 0.1386 DINO. Lower is better.
Forgetting (DINO, lower is better)
Method Forgetting Text alignment
Ours 0.0060 ± 0.0031 75.6
CIDM 0.0232 ± 0.0014 75.9
Fine-tuning 0.1386 ± 0.0042 75.5

Fixed size, fixed cost

A per-concept method grows with every concept it learns: CIDM stores an adapter and its token embeddings for each one, and runs all of them at every denoising step. Our network keeps its size, a new concept adds only its task embedding, and generation costs the same at any length.

Generation time added to the frozen backbone Ours adds 14 percent at 10, 37 and 100 concepts. CIDM adds 45 percent at 10 concepts and 154 percent at 37, and does not run beyond 37.
Generation time added to the frozen backbone
Concepts Ours CIDM
10 +14% +45%
37 +14% +154%
100 +14% does not run

Measured on one A100, batch size 5, 50 DPM-Solver steps, fp32, median over 12 batches.

Put a concept where you want it

Pick a concept and click a quadrant. The dashed box is the region we ask for. The images are those of Figure 7 in the paper.

click a quadrant
requested box
Generated image

Citation

@misc{gromski2026clasp,
  title         = {{CLASP}: Continual Low-rank Adapters for Spatially Placed
                   Concepts from One Hypernetwork},
  author        = {Gromski, Wojciech and Krukowski, Patryk and Miksa, Jan and
                   Zieba, Maciej and Spurek, Przemys{\l}aw},
  year          = {2026},
  eprint        = {2610.01331},
  archivePrefix = {arXiv},
  primaryClass  = {cs.CV},
  url           = {https://arxiv.org/abs/2610.01331}
}