Measurement note 01 · 2026-08-06
Method note · Cross-platform character consistency
One character, seven arms across five independent base models, the same five scenes each, five metrics — plus three further rounds reported separately because their conditions differ. Every number came from one day's work, and the limits of that are stated before the results.
The useful part of this note is not the ranking. It is the measurement setup — and the three places where our own earlier conclusions turned out to be wrong.
Where this sits now. OpenStela is a registry and market. This note documents a measurement setup — how face and whole-subject similarity were scored, normalised and where they break — from the work that preceded the registry. None of it runs on the registry today: the platform records declarations and does not grade works for likeness. It is kept for the record; it is not a product feature.
量測筆記 01 · 2026-08-06
方法筆記 · 跨平台角色一致性
一隻角色、五個獨立底座模型上的七個實驗組、每組相同的五個場景、五項指標;另外三輪後續實驗因為條件不同,分開回報。所有數字都出自一天的工作,這件事的限制會先講,再講結果。
這份筆記有用的地方不是排名,而是量測怎麼設置,以及我們自己先前的結論被證明錯誤的三個地方。
這一頁現在的位置。OpenStela 是登錄簿與市集。這份筆記記錄的是登錄簿之前的一套量測設置:人臉與整體相似度怎麼打分、怎麼正規化、在哪裡會失效。這些現在都不在登錄簿上執行,平台記錄聲明、不替作品打「像不像」的分數。留在這裡當紀錄,不是產品功能。
Method
Identity similarity alone is easy to game. A model that ignores the prompt and reproduces the reference framing every time will score extremely well on any embedding metric, because the output resembles the reference — that is the whole trick. The run below contains a clean example of exactly that.
So every arm is reported on three axes that can disagree:
| Axis | Instrument | Reads |
|---|---|---|
| Identityis it the same person | SFace & ArcFace, on the largest detected face | Same-person thresholds 0.363 / 0.30 |
| Whole subjectdoes it look like the character | Unicom ViT-B/16, background removed | Normalised between floor and ceiling |
| Prompt adherencedid it do what was asked | CLIP-T-C, image against its own scene text | Compared against a deliberately mismatched baseline |
| Rigiditydid it just repeat itself | Within-set similarity across the five outputs | High values mean the composition collapsed |
Normalisation. Raw cosine values are meaningless on their own, so each is placed between two measured endpoints. The ceiling is how similar the character's own reference images are to each other (0.7236 here) — the practical upper bound. The floor is how similar this character is to unrelated characters (0.1951). Prompt adherence gets the same treatment: the same images scored against the wrong scene text give a baseline of ≈0.199, which is what "ignored the prompt" actually looks like.
Two rules that came out of getting this wrong before. Photoreal humans are judged on face metrics, not whole-image embeddings — the latter are dominated by pose, lighting and framing. Non-human and stylised characters have no usable face, so they are judged on Unicom with the background removed and replaced by a uniform grey; without that step the score tracks the background rather than the character.
方法
只看身分相似度很容易作弊。一個不理 prompt、每次都照抄參考圖取景的模型,在任何嵌入指標上都會拿高分,因為輸出本來就長得像參考圖,這就是整個把戲。下面的實驗裡就有一個乾淨的例子。
所以每個實驗組都用三條可以互相矛盾的軸回報:
| 軸 | 儀器 | 怎麼讀 |
|---|---|---|
| 身分是不是同一個人 | SFace 與 ArcFace,取偵測到的最大那張臉 | 同人門檻 0.363 / 0.30 |
| 整體主體像不像這隻角色 | Unicom ViT-B/16,去背 | 在地板與天花板之間正規化 |
| Prompt 遵從有沒有照做 | CLIP-T-C,圖對自己的場景文字 | 和刻意錯配的基準比較 |
| 僵硬度是不是只在重複自己 | 五張輸出之間的組內相似度 | 數值高代表構圖塌縮 |
正規化。原始 cosine 值本身沒有意義,所以每個值都放在兩個實測端點之間。天花板是角色自己的參考圖彼此有多像(這裡是 0.7236),也就是實務上的上限。地板是這隻角色和不相干角色有多像(0.1951)。Prompt 遵從度同樣處理:同一批圖對錯的場景文字打分,得到約 0.199 的基準,那就是「不理 prompt」實際的樣子。
兩條從先前做錯學到的規則。寫實人像用人臉指標評,不用整圖嵌入,後者會被姿勢、光線、取景主導。非人與風格化角色沒有可用的臉,所以先去背、換成均勻灰底,再用 Unicom 評;少了這一步,分數跟著背景走,不跟角色。
Results
Photoreal human character. Same five scene prompts, same reference set, same ceiling and floor. Reference-image count differs per arm because it is each platform's native limit — flattening that would test something no one actually does.
| Arm | Base model | Unicom | SFace ≥.363 | ArcFace ≥.30 | Prompt | Rigidity |
|---|---|---|---|---|---|---|
| GPT Image 2reference · 3 img | OpenAI | 83.3% | 0.5152 | 0.5290 | 0.2360 | 0.6395 |
| Flux LoRA + anchorfine-tune · 5 img | Black Forest Labs | 90.3% | 0.4583 | 0.4723 | 0.2432 | 0.6110 |
| Seedream 4reference · 3 img | ByteDance | 87.2% | 0.4865 | 0.4523 | 0.2506 | 0.6177 |
| Character packmultimodal · 5 img | Google Gemini | 61.4% | 0.4173 | 0.4369 | 0.2447 | 0.4242 |
| Runway Gen-4reference · 3 img | Runway | 85.5% | 0.3396 | 0.3972 | 0.2401 | 0.6478 |
| Flux LoRA alonefine-tune · no text | Black Forest Labs | 44.6% | 0.3601 | 0.3506 | 0.2675 | 0.2815 |
| Flux 1.1 proreference · 1 img | Black Forest Labs | 96.9% | 0.3251 | 0.2581 | 0.2149 | 0.8731 |
Ceiling 0.7236 · floor 0.1951 · prompt-adherence mismatch baseline ≈0.199. Tick marks on each gauge show the same-person threshold for that instrument. Prompt and rigidity columns are raw values, comparable within this table only.
Base-model overlap must be disclosed. Platforms are not independent evidence just because they are different products. Higgsfield's character system runs on Google's Nano Banana Pro — the same family as the Gemini arm. Three of the seven arms above sit on Flux. Counting brands instead of base models inflates apparent coverage.
結果
寫實人像角色。相同的五個場景 prompt、相同的參考圖集、相同的天花板與地板。各組參考圖張數不同,因為那是各平台的原生上限;硬拉平,測到的會是一個沒人真的在做的東西。
| 實驗組 | 底座模型 | Unicom | SFace ≥.363 | ArcFace ≥.30 | Prompt | 僵硬度 |
|---|---|---|---|---|---|---|
| GPT Image 2參考圖 · 3 張 | OpenAI | 83.3% | 0.5152 | 0.5290 | 0.2360 | 0.6395 |
| Flux LoRA + anchor微調 · 5 張 | Black Forest Labs | 90.3% | 0.4583 | 0.4723 | 0.2432 | 0.6110 |
| Seedream 4參考圖 · 3 張 | ByteDance | 87.2% | 0.4865 | 0.4523 | 0.2506 | 0.6177 |
| 角色包多模態 · 5 張 | Google Gemini | 61.4% | 0.4173 | 0.4369 | 0.2447 | 0.4242 |
| Runway Gen-4參考圖 · 3 張 | Runway | 85.5% | 0.3396 | 0.3972 | 0.2401 | 0.6478 |
| Flux LoRA alone微調 · 無文字 | Black Forest Labs | 44.6% | 0.3601 | 0.3506 | 0.2675 | 0.2815 |
| Flux 1.1 pro參考圖 · 1 張 | Black Forest Labs | 96.9% | 0.3251 | 0.2581 | 0.2149 | 0.8731 |
天花板 0.7236 · 地板 0.1951 · prompt 遵從錯配基準 ≈0.199。儀表上的刻度是該儀器的同人門檻。Prompt 與僵硬度兩欄是原始值,只能在本表內比較。
底座模型重疊要揭露。平台不會因為是不同產品就變成獨立證據。Higgsfield 的角色系統跑在 Google 的 Nano Banana Pro 上,和 Gemini 組是同一家族;上面七組有三組建在 Flux 上。數品牌而不數底座模型,會誇大表面上的覆蓋範圍。
Findings
It has the highest whole-subject score in the table (96.9%) and the lowest identity score (0.2581, below threshold). Its rigidity value of 0.8731 is higher than the character's own reference set — the five outputs resemble each other more than the references do. Visual inspection matches: all five are the same half-body studio portrait, and three of the five scene descriptions were simply not executed. The high score is a by-product of composition collapse, not consistency.
A LoRA trained on the character's five reference images, prompted with only a trigger word and a scene, failed the identity threshold (0.3601) and produced a long-haired woman for one of the five scenes. Nothing in the prompt stated the character's gender; the weights alone did not hold against the scene text.
Adding the character pack's written anchor to the same LoRA moved it to 0.4583 and 90.3%. The conclusion is not "don't fine-tune" — it is that the written specification is doing work the weights do not do, and the two compose.
Sweeping Midjourney's Omni Reference weight across five points, identity and prompt adherence both rise together (0.3448→0.4834 and 0.2032→0.2260). At the low end, prompt adherence sits at the mismatched-text baseline — the output has no measurable relationship to the scene description at all. Turning the dial down buys nothing.
None of the five sweep outputs produced the requested full-body street shot. Swapping in a full-body reference image made the framing tighter, not wider. Removing the facial micro-details from the anchor — implant line, heterochromia, temple etching — while keeping silhouette and garment widened the shot immediately.
The exchange rate is poor: prompt adherence rose 7% while ArcFace identity fell 67%. Wide framing is bought with facial identity.
Two platforms, same result. Training an anime character into Higgsfield's Soul ID completed successfully and retained 0% of its declared identity markers; attaching the same images as plain references retained 100%. Binding the same character into Kling's Element system produced no measurable gain over not binding it.
The mechanism explains it: these systems lock the face, and a stylised character's identity mostly lives elsewhere — hair colour, a mark, a garment palette.
Three arms scored high on whole-subject similarity while failing the face threshold: Runway (85.5% / 0.3396), Flux (96.9% / 0.2581), and a full-body-reference variant (99.2% / 0.2418). In each case inspection showed the same styling on a different face. For photoreal humans, whole-image embeddings are not a substitute for a face metric.
發現
它拿到全表最高的整體主體分數(96.9%)和最低的身分分數(0.2581,低於門檻)。僵硬度 0.8731,比角色自己的參考圖集還高,五張輸出彼此比參考圖之間更像。目視結果一致:五張全是同一張半身棚拍肖像,五個場景描述有三個根本沒被執行。高分是構圖塌縮的副產品,不是一致性。
用角色的五張參考圖訓練 LoRA,只用觸發詞加場景來下 prompt,身分門檻沒過(0.3601),五個場景裡有一個生出了一位長髮女性。prompt 裡沒有任何地方寫到角色的性別,光靠權重擋不住場景文字。
把角色包的書面錨句加到同一個 LoRA 上,分數升到 0.4583 與 90.3%。結論不是「別微調」,而是書面規格在做權重沒做的事,兩者可以疊加。
把 Midjourney 的 Omni Reference 權重掃過五個點,身分和 prompt 遵從度一起上升(0.3448→0.4834、0.2032→0.2260)。在低端,prompt 遵從度落在錯配文字的基準上,輸出和場景描述完全沒有可量測的關係。把旋鈕轉低,什麼都換不到。
五個掃描輸出沒有一個生出要求的全身街拍。換成全身參考圖,取景反而更緊。把錨句裡的臉部微細節拿掉(植入線、異色瞳、太陽穴蝕刻),留下輪廓和服裝,畫面立刻拉寬。
匯率很差:prompt 遵從度升 7%,ArcFace 身分掉 67%。寬取景是拿臉部身分換來的。
兩個平台,同一個結果。把一隻動漫角色訓練進 Higgsfield 的 Soul ID,訓練成功完成,卻保留了 0% 的聲明身分特徵;同一批圖當普通參考圖附上,保留 100%。把同一隻角色綁進 Kling 的 Element 系統,相較於不綁,沒有可量測的增益。
機制說得通:這些系統鎖的是臉,而風格化角色的身分多半在別的地方:髮色、一個記號、一組服裝配色。
三組在整體主體相似度上拿高分、卻沒過人臉門檻:Runway(85.5% / 0.3396)、Flux(96.9% / 0.2581)、一個全身參考圖變體(99.2% / 0.2418)。每一組目視都是同樣的造型換了一張臉。對寫實人像來說,整圖嵌入不能取代人臉指標。
Field notes
--ow 0–1000 parameter. Typing --ow 200 leaves the slider
untouched.custom_reference_id.
Passing --soul-id raises no error and quietly returns an image with no
character applied — it scored at the floor.completedsuccess. A poller waiting on the wrong string never exits.
And a completed training is not a usable one: verify by generating.現場筆記
--ow 0–1000 參數不同步。打 --ow 200,滑桿不會動。custom_reference_id。傳 --soul-id 不會出錯,只會悄悄回一張沒套角色的圖,分數落在地板。completedsuccess。等錯字串的輪詢永遠不會結束。而且訓練完成不等於能用,要靠生成來驗證。Corrections
These are included because a benchmark with no failed hypotheses has not been audited.
Reference strength trades character fidelity against compositional
freedom, so it should be tuned down for scene-heavy shots.
Both rise together across the full range. There is no trade to make; the dial should sit at maximum.
Composition is locked to the reference image's framing and cannot be
released.
It is not the reference — a full-body reference produced a tighter crop. It is the proportion of facial detail in the prompt, and it can be released by rewriting the anchor.
Kling exposes a three-step reference-strength control that would let us
test whether the trade-off generalises.
No such control exists in the product. The claim came from a third-party article and was not verified before planning around it.
更正
收錄這些,是因為一份沒有落空假設的基準,等於沒有被檢查過。
參考強度是在角色忠實度和構圖自由之間取捨,所以場景重的鏡頭應該調低。
兩者在整個範圍內一起上升。沒有什麼好取捨的,旋鈕應該放在最大。
構圖被鎖在參考圖的取景上,解不開。
不是參考圖的問題,全身參考圖反而生出更緊的裁切。是 prompt 裡臉部細節的比例,改寫錨句就能解開。
Kling 提供三段式參考強度控制,可以用來測試取捨是否普遍成立。
產品裡沒有這樣的控制項。這個說法來自第三方文章,在計畫之前沒有先驗證。
Limits
Every image, prompt, and score in this note came from scripts that take the character specification as input and write the measurement table as output. The setup is designed to be re-run against new platforms as they appear, which is the only way a note like this stays true for longer than a quarter.
限制
這份筆記裡每一張圖、每一段 prompt、每一個分數,都出自以角色規格為輸入、以量測表為輸出的腳本。整套設置是為了對新平台重跑而設計的,那是這樣一份筆記能維持超過一季仍然為真的唯一方法。