We trained a character into a platform. It kept 0% of its markers.
If you ask anyone — including us, four days before the run — how to get the strongest character lock on a generation platform, the answer is the same: train the character in. Fine-tune it, make it a native object, pay the training cost and collect the consistency. That is what the platforms recommend, and it is what intuition says.
So we measured it. One stylised character with declared, checkable identity markers — hair colour, a face mark, a specific collar palette. Two conditions on the same platform: the trained-in version, and plain reference images with the written specification. Same scenes, same judge.
The result
The trained version kept 0% of the declared markers. Hair colour wrong, face mark gone, collar palette replaced. The plain-reference condition kept 100%. We assumed a mistake and re-ran it on a second platform: same direction, same conclusion.
The mechanism, as far as the runs let us see it: training compresses a character into the model’s own style space. For a stylised character, the nearest point in that space is a different character — smoothed, prettified, regressed toward what the model already likes to draw. Reference images at generation time do not get compressed; they sit next to the prompt as evidence the model has to reconcile on every call.
Why a similarity score hides this
The trained condition scored well on whole-image similarity. It reproduced composition, lighting and rendering style beautifully. It just wasn’t the same character — and a single similarity number cannot tell those two things apart. This is why every result we publish is reported on axes that can disagree: face identity, whole-subject similarity, prompt adherence, and how rigid the output set became. The platform with the highest single score in our main run had the lowest identity — it got its score by repeating the reference framing and ignoring three of the five scenes.
What changed because of this
The per-platform notes that travel with a registered character are built on references-plus-specification, not on training, and their wording is written from measured behaviour rather than platform documentation. Where training is the right call — it sometimes is, for photoreal humans on specific platforms — the note says so explicitly, with the run that justifies it.
Two of the four headline findings in our benchmark contradicted what we ourselves believed when we designed it. That is the strongest argument we know for measuring instead of assuming — and for publishing the misses along with the hits.
我們把角色訓練進平台,它保留了 0% 的特徵
問任何人,包括做實驗前四天的我們自己,要怎麼在生成平台上把角色鎖得最牢,答案都一樣:把角色訓練進去。微調它,讓它變成平台原生的物件,付訓練費,換一致性。平台是這麼建議的,直覺也是這麼說的。
所以我們量了。一隻風格化角色,事先寫好可以核對的身分特徵:髮色、一個臉部記號、一組特定的衣領配色。同一個平台上兩種條件:訓練進去的版本,以及普通參考圖加書面規格。相同場景、相同評分方式。
結果
訓練版保留了 0% 的聲明特徵:髮色錯、臉部記號不見、衣領配色被換掉。普通參考圖那組保留了 100%。我們以為自己弄錯了,換到第二個平台重跑:同樣的方向、同樣的結論。
就實驗看得到的範圍,機制是這樣:訓練把角色壓進模型自己的風格空間。對風格化角色來說,那個空間裡最近的點是另一隻角色,被抹平、被修飾、被拉回模型本來就喜歡畫的樣子。生成時附上的參考圖不會被壓縮,它們就在 prompt 旁邊,是模型每一次呼叫都得對照的證據。
為什麼相似度分數看不出來
訓練版在整圖相似度上分數很好。它漂亮地重現了構圖、光線和渲染風格,只是不是同一隻角色,而一個相似度數字分不出這兩件事。這就是為什麼我們公開的每個結果都拆成可以彼此矛盾的軸:人臉身分、整體相似、prompt 遵從度、輸出集合的僵硬程度。主實驗裡單一分數最高的平台,身分分數最低;它是靠重複參考圖的取景、忽略五個場景裡的三個拿到分數的。
這件事改變了什麼
跟著登記角色走的平台備註,都建立在「參考圖+規格」上,不建立在訓練上;措辭依量到的行為寫,不照平台文件寫。訓練確實是對的選擇時(某些平台上的寫實人像偶爾如此),適配包會明講,並附上支持它的那次實驗。
量測基準裡四個主要發現,有兩個和我們設計它時所相信的相反。這是「量測而不是假設」,以及「落空的和命中的一起公開」最有力的理由。