STAGE: Subspace-Targeted Affine Generative Erasure

for Text-to-3D Models

Karol Dziekan1Przemysław Spurek1,2Dawid Malarz1,2

1 Jagiellonian University    2 IDEAS Research Institute

Left: TRELLIS generates a voxel structure in stage G_S and fills it with material and fine geometry in stage G_L. Right: a table under the three edit modes; editing G_S changes its shape, editing G_L changes its material.
Concepts localize to different stages of TRELLIS. TRELLIS first fixes the voxel scaffold in a structural stage GS and then fills it with material, color and fine geometry in a latent stage GL. Editing only GS (Mode-S) changes shape but not material, editing only GL (Mode-L) changes material but not shape, and object identity needs both stages (Mode-SL).

Abstract

Concept erasure suppresses a target concept while preserving behavior on unrelated inputs. Existing closed-form methods were designed for 2D image diffusion and assume a single generative pathway, so one edit must cover geometry and texture at once. Native 3D generators, which synthesize structured 3D representations directly rather than by lifting 2D samples, violate this assumption. We show that shape and object concepts must be erased in the structural stage of the pipeline and material concepts in the appearance stage. We therefore formulate erasure in native text-to-3D as a stage-aware editing problem and introduce STAGE, a training-free, closed-form framework. STAGE confines each edit to the low-dimensional subspace spanned by the differences between erase and anchor embeddings, and relaxes the norm-preserving (orthogonal) constraint of prior editors into a least-squares affine correction that maps target activations onto safe anchors subject to a penalty on the displacement of retained prompts. The correction applies to the structural stage, the appearance stage, or both. We find that the stage an edit must reach is determined by concept type. On TRELLIS, the standard open native 3D generator, across 15 shape, material, and object concepts, STAGE reaches 66.7 on a composite score that balances forgetting the target concept against preserving everything else, aggregating CLIP-based semantic and physical metrics, versus 53.2 for the strongest adapted baseline.

Method

The STAGE pipeline: paired erase and anchor prompts are encoded, their differences span a contrast subspace, and a closed-form edit of the cross-attention key and value projections is folded into the structural stage, the appearance stage, or both.
STAGE folded into TRELLIS. STAGE encodes paired prompts, builds the contrast subspace from the differences between erase and anchor embeddings, and solves the closed-form correction in one step. The resulting weights replace WK and WV in the cross-attention of GS, GL, or both, with no retraining.
01

Contrast subspace

Erase and anchor prompts differ only in the concept word. The leading singular directions of their token-level differences span the subspace the edit may act in.

02

Affine correction

A least-squares affine map moves erase embeddings onto their anchors, while anchors, the null prompt and other retained prompts stay in place.

03

Stage-aware folding

The correction is folded exactly into the cross-attention of the stage the concept lives in, in seconds and without training.

T(c) = c + N(D⊤c) + n0   ⟶   W′ = W + (WN)D⊤,   b′ = b + Wn0

Mode-S · shapes Mode-L · materials Mode-SL · objects
Three editing primitives: an unconstrained least-squares update moves unrelated prompts, a rotation must move the anchor to reach it, and the STAGE correction acts only along the erase direction.
Why an affine correction. A least-squares update of the whole projection (UCE) is not confined to a subspace, so unrelated prompts drift. A rotation (OCE) preserves distances, so it cannot move the target onto the anchor without moving the anchor too. STAGE acts only along the erase direction and keeps anchors and retained prompts in place.

Results

We evaluate on 15 concepts of TRELLIS, five shapes, five materials and five objects, with the 3D Unlearning Score (3D-US), which combines forgetting (F) of the erased concept and preservation (P) of all other concepts, measured with CLIP on renders and with Uni3D on point clouds or a material classifier. All methods edit the same stages.

Method Shape Material Object Overall
FP3D-US FP3D-US FP3D-US 3D-US
UCE 0.610.8566.9 0.220.8430.8 0.570.6551.9 47.5
OCE 0.730.6767.9 0.360.7441.8 0.830.4153.2 53.2
STAGE 0.700.9177.7 0.450.9256.4 0.900.6167.7 66.7

F and P lie in [0, 1], 3D-US in [0, 100], higher is better. Axis values are means over the five concepts of the axis, and Overall is the geometric mean of the three axes.

Qualitative comparison

One concept per axis erased by UCE, OCE and STAGE on the same prompt and seed: cube to sphere, stone to wood, table to car.
One concept per axis. Shape (cubic → round), material (stone → wood) and object (table → car), all methods on the same prompt and seed.

Redirection across seeds

A chair prompt erased toward teddy bear with seven random seeds per method: STAGE produces a teddy bear every time.
Chair → teddy bear under seven random seeds. STAGE yields a teddy bear under every seed, while UCE and OCE mostly produce distorted or unrelated objects.
Mean forgetting against mean preservation for every configuration of the three hyperparameter sweeps; the STAGE configurations lie above those of both baselines.

Forgetting and preservation

Every configuration of the hyperparameter sweeps of the three methods, placed by its mean forgetting and preservation. The STAGE configurations form a front above both baselines, and even at its default hyperparameters STAGE scores above both tuned baselines (60.1 vs 53.2).

BibTeX

@misc{dziekan2026stage,
  title  = {{STAGE}: Subspace-Targeted Affine Generative Erasure for Text-to-3D Models},
  author = {Dziekan, Karol and Spurek, Przemys{\l}aw and Malarz, Dawid},
  year   = {2026}
}