OpenStela

Measurement note 01 · 2026-08-06

Method note · Cross-platform character consistency

What survives when a character moves between platforms

One character, seven arms across five independent base models, the same five scenes each, five metrics — plus three further rounds reported separately because their conditions differ. Every number came from one day's work, and the limits of that are stated before the results.

The useful part of this note is not the ranking. It is the measurement setup — and the three places where our own earlier conclusions turned out to be wrong.

Where this sits now. OpenStela is a registry and market. This note documents a measurement setup — how face and whole-subject similarity were scored, normalised and where they break — from the work that preceded the registry. None of it runs on the registry today: the platform records declarations and does not grade works for likeness. It is kept for the record; it is not a product feature.

量測筆記 01 · 2026-08-06

方法筆記 · 跨平台角色一致性

角色在平台之間移動時,什麼撐得下來

一隻角色、五個獨立底座模型上的七個實驗組、每組相同的五個場景、五項指標;另外三輪後續實驗因為條件不同,分開回報。所有數字都出自一天的工作,這件事的限制會先講,再講結果。

這份筆記有用的地方不是排名,而是量測怎麼設置,以及我們自己先前的結論被證明錯誤的三個地方。

這一頁現在的位置。OpenStela 是登錄簿與市集。這份筆記記錄的是登錄簿之前的一套量測設置:人臉與整體相似度怎麼打分、怎麼正規化、在哪裡會失效。這些現在都不在登錄簿上執行,平台記錄聲明、不替作品打「像不像」的分數。留在這裡當紀錄,不是產品功能。

Scope

What this is

  • A complete record of one controlled run, including the setup that makes it reproducible
  • A measurement method that separates identity retention from prompt adherence — two things that are routinely conflated
  • Platform behaviours that are not in any vendor's documentation

What this is not

  • Not a ranking. n=5 images per arm, one character, one seed per point
  • Not a claim that any platform is better. Conditions differ per platform by design, because each has a different native mechanism
  • Not independent. We built the character pack being measured. Read the limitations

範圍

這是什麼

  • 一次受控實驗的完整紀錄,包含讓它可以重現的設置
  • 一套把「身分保留」和「prompt 遵從」拆開的量測方法,這兩件事經常被混為一談
  • 沒有寫在任何廠商文件裡的平台行為

這不是什麼

  • 不是排名。每組 n=5 張圖、一隻角色、每個點一個種子
  • 不主張哪個平台比較好。各平台條件刻意不同,因為每個平台的原生機制不一樣
  • 不獨立。被量的角色包是我們自己做的,請看「限制」那一節

Method

Three numbers, and why one is never enough

Identity similarity alone is easy to game. A model that ignores the prompt and reproduces the reference framing every time will score extremely well on any embedding metric, because the output resembles the reference — that is the whole trick. The run below contains a clean example of exactly that.

So every arm is reported on three axes that can disagree:

AxisInstrumentReads
Identityis it the same person SFace & ArcFace, on the largest detected face Same-person thresholds 0.363 / 0.30
Whole subjectdoes it look like the character Unicom ViT-B/16, background removed Normalised between floor and ceiling
Prompt adherencedid it do what was asked CLIP-T-C, image against its own scene text Compared against a deliberately mismatched baseline
Rigiditydid it just repeat itself Within-set similarity across the five outputs High values mean the composition collapsed

Normalisation. Raw cosine values are meaningless on their own, so each is placed between two measured endpoints. The ceiling is how similar the character's own reference images are to each other (0.7236 here) — the practical upper bound. The floor is how similar this character is to unrelated characters (0.1951). Prompt adherence gets the same treatment: the same images scored against the wrong scene text give a baseline of ≈0.199, which is what "ignored the prompt" actually looks like.

Two rules that came out of getting this wrong before. Photoreal humans are judged on face metrics, not whole-image embeddings — the latter are dominated by pose, lighting and framing. Non-human and stylised characters have no usable face, so they are judged on Unicom with the background removed and replaced by a uniform grey; without that step the score tracks the background rather than the character.

方法

三個數字,以及為什麼一個永遠不夠

只看身分相似度很容易作弊。一個不理 prompt、每次都照抄參考圖取景的模型,在任何嵌入指標上都會拿高分,因為輸出本來就長得像參考圖,這就是整個把戲。下面的實驗裡就有一個乾淨的例子。

所以每個實驗組都用三條可以互相矛盾的軸回報:

軸儀器怎麼讀
身分是不是同一個人 SFace 與 ArcFace,取偵測到的最大那張臉 同人門檻 0.363 / 0.30
整體主體像不像這隻角色 Unicom ViT-B/16,去背 在地板與天花板之間正規化
Prompt 遵從有沒有照做 CLIP-T-C,圖對自己的場景文字 和刻意錯配的基準比較
僵硬度是不是只在重複自己 五張輸出之間的組內相似度 數值高代表構圖塌縮

正規化。原始 cosine 值本身沒有意義,所以每個值都放在兩個實測端點之間。天花板是角色自己的參考圖彼此有多像(這裡是 0.7236),也就是實務上的上限。地板是這隻角色和不相干角色有多像(0.1951)。Prompt 遵從度同樣處理:同一批圖對錯的場景文字打分,得到約 0.199 的基準,那就是「不理 prompt」實際的樣子。

兩條從先前做錯學到的規則。寫實人像用人臉指標評,不用整圖嵌入,後者會被姿勢、光線、取景主導。非人與風格化角色沒有可用的臉,所以先去背、換成均勻灰底,再用 Unicom 評;少了這一步,分數跟著背景走,不跟角色。

Results

One character, seven arms

Photoreal human character. Same five scene prompts, same reference set, same ceiling and floor. Reference-image count differs per arm because it is each platform's native limit — flattening that would test something no one actually does.

ArmBase model Unicom SFace ≥.363 ArcFace ≥.30 Prompt Rigidity
GPT Image 2reference · 3 img OpenAI 83.3% 0.5152 0.5290 0.2360 0.6395
Flux LoRA + anchorfine-tune · 5 img Black Forest Labs 90.3% 0.4583 0.4723 0.2432 0.6110
Seedream 4reference · 3 img ByteDance 87.2% 0.4865 0.4523 0.2506 0.6177
Character packmultimodal · 5 img Google Gemini 61.4% 0.4173 0.4369 0.2447 0.4242
Runway Gen-4reference · 3 img Runway 85.5% 0.3396 0.3972 0.2401 0.6478
Flux LoRA alonefine-tune · no text Black Forest Labs 44.6% 0.3601 0.3506 0.2675 0.2815
Flux 1.1 proreference · 1 img Black Forest Labs 96.9% 0.3251 0.2581 0.2149 0.8731

Ceiling 0.7236 · floor 0.1951 · prompt-adherence mismatch baseline ≈0.199. Tick marks on each gauge show the same-person threshold for that instrument. Prompt and rigidity columns are raw values, comparable within this table only.

Base-model overlap must be disclosed. Platforms are not independent evidence just because they are different products. Higgsfield's character system runs on Google's Nano Banana Pro — the same family as the Gemini arm. Three of the seven arms above sit on Flux. Counting brands instead of base models inflates apparent coverage.

結果

一隻角色,七個實驗組

寫實人像角色。相同的五個場景 prompt、相同的參考圖集、相同的天花板與地板。各組參考圖張數不同,因為那是各平台的原生上限;硬拉平,測到的會是一個沒人真的在做的東西。

實驗組底座模型 Unicom SFace ≥.363 ArcFace ≥.30 Prompt 僵硬度
GPT Image 2參考圖 · 3 張 OpenAI 83.3% 0.5152 0.5290 0.2360 0.6395
Flux LoRA + anchor微調 · 5 張 Black Forest Labs 90.3% 0.4583 0.4723 0.2432 0.6110
Seedream 4參考圖 · 3 張 ByteDance 87.2% 0.4865 0.4523 0.2506 0.6177
角色包多模態 · 5 張 Google Gemini 61.4% 0.4173 0.4369 0.2447 0.4242
Runway Gen-4參考圖 · 3 張 Runway 85.5% 0.3396 0.3972 0.2401 0.6478
Flux LoRA alone微調 · 無文字 Black Forest Labs 44.6% 0.3601 0.3506 0.2675 0.2815
Flux 1.1 pro參考圖 · 1 張 Black Forest Labs 96.9% 0.3251 0.2581 0.2149 0.8731

天花板 0.7236 · 地板 0.1951 · prompt 遵從錯配基準 ≈0.199。儀表上的刻度是該儀器的同人門檻。Prompt 與僵硬度兩欄是原始值,只能在本表內比較。

底座模型重疊要揭露。平台不會因為是不同產品就變成獨立證據。Higgsfield 的角色系統跑在 Google 的 Nano Banana Pro 上,和 Gemini 組是同一家族;上面七組有三組建在 Flux 上。數品牌而不數底座模型,會誇大表面上的覆蓋範圍。

Findings

What the numbers say

Flux 1.1 pro is the clearest argument for reading three columns

It has the highest whole-subject score in the table (96.9%) and the lowest identity score (0.2581, below threshold). Its rigidity value of 0.8731 is higher than the character's own reference set — the five outputs resemble each other more than the references do. Visual inspection matches: all five are the same half-body studio portrait, and three of the five scene descriptions were simply not executed. The high score is a by-product of composition collapse, not consistency.

A fine-tune without a written spec is not enough

A LoRA trained on the character's five reference images, prompted with only a trigger word and a scene, failed the identity threshold (0.3601) and produced a long-haired woman for one of the five scenes. Nothing in the prompt stated the character's gender; the weights alone did not hold against the scene text.

Adding the character pack's written anchor to the same LoRA moved it to 0.4583 and 90.3%. The conclusion is not "don't fine-tune" — it is that the written specification is doing work the weights do not do, and the two compose.

Reference strength is not a trade-off dial

Sweeping Midjourney's Omni Reference weight across five points, identity and prompt adherence both rise together (0.3448→0.4834 and 0.2032→0.2260). At the low end, prompt adherence sits at the mismatched-text baseline — the output has no measurable relationship to the scene description at all. Turning the dial down buys nothing.

Framing is controlled by the density of facial detail in the prompt

None of the five sweep outputs produced the requested full-body street shot. Swapping in a full-body reference image made the framing tighter, not wider. Removing the facial micro-details from the anchor — implant line, heterochromia, temple etching — while keeping silhouette and garment widened the shot immediately.

The exchange rate is poor: prompt adherence rose 7% while ArcFace identity fell 67%. Wide framing is bought with facial identity.

Closed character systems do not help stylised characters

Two platforms, same result. Training an anime character into Higgsfield's Soul ID completed successfully and retained 0% of its declared identity markers; attaching the same images as plain references retained 100%. Binding the same character into Kling's Element system produced no measurable gain over not binding it.

The mechanism explains it: these systems lock the face, and a stylised character's identity mostly lives elsewhere — hair colour, a mark, a garment palette.

"Looks like them" and "is them" come apart, repeatedly

Three arms scored high on whole-subject similarity while failing the face threshold: Runway (85.5% / 0.3396), Flux (96.9% / 0.2581), and a full-body-reference variant (99.2% / 0.2418). In each case inspection showed the same styling on a different face. For photoreal humans, whole-image embeddings are not a substitute for a face metric.

發現

數字說了什麼

Flux 1.1 pro 是「要讀三欄」最清楚的例子

它拿到全表最高的整體主體分數(96.9%)和最低的身分分數(0.2581,低於門檻)。僵硬度 0.8731,比角色自己的參考圖集還高,五張輸出彼此比參考圖之間更像。目視結果一致:五張全是同一張半身棚拍肖像,五個場景描述有三個根本沒被執行。高分是構圖塌縮的副產品,不是一致性。

沒有書面規格的微調,不夠

用角色的五張參考圖訓練 LoRA,只用觸發詞加場景來下 prompt,身分門檻沒過(0.3601),五個場景裡有一個生出了一位長髮女性。prompt 裡沒有任何地方寫到角色的性別,光靠權重擋不住場景文字。

把角色包的書面錨句加到同一個 LoRA 上,分數升到 0.4583 與 90.3%。結論不是「別微調」,而是書面規格在做權重沒做的事,兩者可以疊加。

參考強度不是一個取捨旋鈕

把 Midjourney 的 Omni Reference 權重掃過五個點,身分和 prompt 遵從度一起上升(0.3448→0.4834、0.2032→0.2260)。在低端,prompt 遵從度落在錯配文字的基準上,輸出和場景描述完全沒有可量測的關係。把旋鈕轉低,什麼都換不到。

取景由 prompt 裡臉部細節的密度決定

五個掃描輸出沒有一個生出要求的全身街拍。換成全身參考圖,取景反而更緊。把錨句裡的臉部微細節拿掉(植入線、異色瞳、太陽穴蝕刻),留下輪廓和服裝,畫面立刻拉寬。

匯率很差:prompt 遵從度升 7%,ArcFace 身分掉 67%。寬取景是拿臉部身分換來的。

封閉角色系統對風格化角色沒有幫助

兩個平台,同一個結果。把一隻動漫角色訓練進 Higgsfield 的 Soul ID,訓練成功完成,卻保留了 0% 的聲明身分特徵;同一批圖當普通參考圖附上,保留 100%。把同一隻角色綁進 Kling 的 Element 系統,相較於不綁,沒有可量測的增益。

機制說得通:這些系統鎖的是臉,而風格化角色的身分多半在別的地方:髮色、一個記號、一組服裝配色。

「看起來像他」和「就是他」一再分道揚鑣

三組在整體主體相似度上拿高分、卻沒過人臉門檻:Runway(85.5% / 0.3396)、Flux(96.9% / 0.2581)、一個全身參考圖變體(99.2% / 0.2418)。每一組目視都是同樣的造型換了一張臉。對寫實人像來說,整圖嵌入不能取代人臉指標。

Field notes

Behaviour not found in any documentation

Midjourney · web slider ≠ --ow
The Omni Reference slider runs 1–100 and does not sync with the documented --ow 0–1000 parameter. Typing --ow 200 leaves the slider untouched.
Higgsfield · silent flag rejection
The parameter that attaches a trained character is custom_reference_id. Passing --soul-id raises no error and quietly returns an image with no character applied — it scored at the floor.
Higgsfield · terminal state is completed
Not success. A poller waiting on the wrong string never exits. And a completed training is not a usable one: verify by generating.
Kling · element binding requires a start frame
Removing the start image makes the binding control disappear entirely, so the character system cannot be tested in isolation on this platform.
Video · aggregate scores are meaningless
Same platform, same start frame, two scenes: 7.4% where the camera pulled away into a crowd, 92.2% where it stayed close. Averaging those produces a number that describes neither.
Replicate · 6 requests per minute
Exceeding it returns 429 mid-batch, which looks like a model failure until you read the body.

現場筆記

任何文件裡都找不到的行為

Midjourney · 網頁滑桿 ≠ --ow
Omni Reference 滑桿的範圍是 1–100,和文件寫的 --ow 0–1000 參數不同步。打 --ow 200,滑桿不會動。
Higgsfield · 旗標錯了不會報錯
掛上已訓練角色的參數是 custom_reference_id。傳 --soul-id 不會出錯,只會悄悄回一張沒套角色的圖,分數落在地板。
Higgsfield · 終態是 completed
不是 success。等錯字串的輪詢永遠不會結束。而且訓練完成不等於能用,要靠生成來驗證。
Kling · 元素綁定需要起始影格
拿掉起始圖,綁定控制項整個消失,所以在這個平台上沒辦法單獨測試角色系統。
影片 · 總分沒有意義
同平台、同起始影格、兩個場景:鏡頭拉遠進人群 7.4%,鏡頭留近 92.2%。把兩者平均,得到一個什麼都不描述的數字。
Replicate · 每分鐘 6 次請求
超過會在批次跑到一半時回 429,看起來像模型失敗,直到你去讀回應本文。

Corrections

Three things we had wrong

These are included because a benchmark with no failed hypotheses has not been audited.

Reference strength trades character fidelity against compositional freedom, so it should be tuned down for scene-heavy shots.

Both rise together across the full range. There is no trade to make; the dial should sit at maximum.

Composition is locked to the reference image's framing and cannot be released.

It is not the reference — a full-body reference produced a tighter crop. It is the proportion of facial detail in the prompt, and it can be released by rewriting the anchor.

Kling exposes a three-step reference-strength control that would let us test whether the trade-off generalises.

No such control exists in the product. The claim came from a third-party article and was not verified before planning around it.

更正

我們錯了的三件事

收錄這些,是因為一份沒有落空假設的基準,等於沒有被檢查過。

參考強度是在角色忠實度和構圖自由之間取捨,所以場景重的鏡頭應該調低。

兩者在整個範圍內一起上升。沒有什麼好取捨的,旋鈕應該放在最大。

構圖被鎖在參考圖的取景上,解不開。

不是參考圖的問題,全身參考圖反而生出更緊的裁切。是 prompt 裡臉部細節的比例,改寫錨句就能解開。

Kling 提供三段式參考強度控制,可以用來測試取捨是否普遍成立。

產品裡沒有這樣的控制項。這個說法來自第三方文章,在計畫之前沒有先驗證。

Limits

What would need to change to trust this further

  • Sample size. Five images per arm, one image per point on the sweep. Enough for direction, not for magnitude.
  • One character. The photoreal results come from a single subject. Anime findings come from a second one. Neither generalises on its own.
  • Home advantage. The reference images were generated with Gemini, which plausibly favours the Gemini arm. The LoRA was trained on those same images, which plausibly favours it in a different direction. The two biases are not netted out.
  • Prompt adherence is a narrow instrument. CLIP-T-C values sit between 0.21 and 0.27; only within-run comparisons against the mismatch baseline are meaningful.
  • Face metrics disagree at the boundary. One arm passed on ArcFace and failed on SFace. Both are reported rather than picking a winner.
  • One point had no detectable face — a dark, heavily shadowed portrait. The detector's known weakness, not a generation failure, but it makes the low end of the sweep less reliable.

Every image, prompt, and score in this note came from scripts that take the character specification as input and write the measurement table as output. The setup is designed to be re-run against new platforms as they appear, which is the only way a note like this stays true for longer than a quarter.

限制

要更信任這份筆記,需要改變什麼

  • 樣本數。每組五張圖,掃描每個點一張。足以確立方向,不足以確立量級。
  • 一隻角色。寫實結果來自單一主體,動漫發現來自第二隻。兩者單獨都推廣不了。
  • 主場優勢。參考圖是用 Gemini 生的,合理推測偏向 Gemini 組。LoRA 用同一批圖訓練,合理推測往另一個方向偏向它。兩個偏誤沒有互相抵銷。
  • Prompt 遵從是很窄的儀器。CLIP-T-C 值落在 0.21 到 0.27 之間,只有同一輪內對錯配基準的比較有意義。
  • 人臉指標在邊界上不一致。有一組 ArcFace 過了、SFace 沒過。兩個都回報,不挑贏家。
  • 有一個點偵測不到臉:一張昏暗、重陰影的肖像。這是偵測器的已知弱點,不是生成失敗,但它讓掃描的低端比較不可靠。

這份筆記裡每一張圖、每一段 prompt、每一個分數,都出自以角色規格為輸入、以量測表為輸出的腳本。整套設置是為了對新平台重跑而設計的,那是這樣一份筆記能維持超過一季仍然為真的唯一方法。