OpenStela

← Blog

Two kinds of character break platform training. Most PFPs are one of them.

A profile-picture NFT is a token pointing at one image. The image is fine — it is the thing you chose, and for a lot of people it has been their face online for years. But the moment you want that character to do something it was not drawn doing — a different angle, a scene, eight seconds of video — something has to generate it again, and that is where it usually stops being the same character.

How far it drifts is measurable. A prompt that merely describes a character in words scores 0.18–0.27 on a face-identity scale where 0.363 is the line between “same person” and “different person”. Not a worse version of your character. A different one.

The obvious fix is the one that fails

Most platforms now offer a character system: upload a handful of images, wait for training, get a reusable character you can drop into any prompt. It looks like exactly the right tool. We measured what it does to the two kinds of character that profile-picture collections actually consist of — non-human creatures, and stylised or illustrated humans.

Non-human: refused at the door

Apes, frogs, cats, robots, blobs. We submitted 20 images of a non-human character to Higgsfield’s Soul training. It was accepted, queued, and then:

status: failed   fail_reason: "face_not_found"

The training step looks for a face. No face, no character. The cinematic variant behaved identically, so this is a platform-level check rather than a quirk of one model. To the platform’s credit, the failure costs nothing — but it also arrives after the submission looked like it had worked, which is why we never treat “queued” as success.

Stylised human: trains perfectly, comes back as someone else

Punks, anime, illustrated characters. This time training completed normally — 25 credits, about half an hour, status: completed, no error. Then we generated a close portrait with it. Hair colour wrong. Eye colour wrong. The star-shaped birthmark gone. Collar colours wrong. A competent picture of a different character.

Same character, same platform Overall appearanceIdentity markers kept
Trained character system 25 credits · 30 min 43.0%0%
One reference image attached no training 63.1%100%

None of the identity markers survived training. All of them survived simply attaching one reference image. The reason holds up: these systems lock the face. For a photorealistic human the face is most of the identity, so locking it is a reasonable trade. For an illustrated character the identity is mostly not on the face at all — it is hair colour, a birthmark, a palette, a collar, an accessory. Lock the face, let everything else drift, and you get a stranger with the right bone structure.

We saw the same pattern on a second, unrelated platform: attaching its character object measured no better than simply supplying the start frame, with the markers holding at 100% either way. Two independent platforms agreeing turned this from a platform quirk into a rule we now apply by default to any platform of that type.

What does work

Reference mechanisms — a written specification, a small set of reference images, and wording adapted per platform. For non-human characters those score 107% (Gemini), 105.5% (Midjourney) and 110.8% (Veo) against the reference set’s own ceiling: the generated frames agree with the character more closely than the reference images, shot across five extreme angles, agree with each other.

This is why the notes that travel with a registered character do not point you at the expensive option. For a non-human character the training route is not offered at all; for a stylised one it warns that training measured worse than not training. The two warnings are worded differently on purpose — “this will fail” and “this will succeed and be worse” are not the same message, and blurring them would leave people thinking illustrated characters cannot be done here.

Limits, as always: five images per group, one character per type, single base model per platform. Read them as magnitudes, not as a ranking with decimal places. The full run is on the measurements page.

A word on rights, since people ask

The token and the licence are two different things, and the chain records only the first. Bored Ape holders were granted broad commercial rights to their ape. Nouns, Cryptoadz and mfers are CC0 — which means the artwork is free for anyone to use, and that includes anyone who is not you. Many other collections grant limited rights or none at all. Holding the token does not tell you which of those you are in; the collection’s terms do. Read them before you build anything commercial on the output. That is a description of how these licences generally work, not advice about your particular collection.

Registering a character here does not change any of it. An ID records that this specification existed here, unaltered, from a given date — evidence, not title. We are specific about that distinction, because a registry that overstates it is writing a cheque its records cannot cash.

Bring your character in →  ·  Argue with the numbers in Discord

← 部落格

兩種角色會讓平台訓練失效,多數 PFP 剛好是其中之一

頭像型 NFT 是一枚指向一張圖的代幣。那張圖沒問題,是你選的,對很多人來說它已經當了好幾年的網路門面。但只要你想讓那隻角色做一件當初沒被畫出來的事,換個角度、換個場景、來一段八秒影片,就得有東西把它重新生成一次;而通常就在這一步,它不再是同一隻角色。

漂走多遠是量得出來的。只用文字描述角色的 prompt,在人臉身分量表上落在 0.18 到 0.27,而 0.363 是「同一人」和「不同人」的分界。那不是你角色的劣化版,是另一隻角色。

最直覺的解法,正是失敗的那個

現在多數平台都有角色系統:上傳幾張圖、等訓練、拿到一個可以丟進任何 prompt 的可重用角色。看起來就是對的工具。我們量了它對頭像系列實際上由哪兩種角色組成的效果:非人生物,以及風格化或插畫風的人類。

非人:在門口就被擋下

猿、蛙、貓、機器人、一團什麼。我們把一隻非人角色的 20 張圖送進 Higgsfield 的 Soul 訓練。被接受、排隊,然後:

status: failed   fail_reason: "face_not_found"

訓練這一步在找一張臉。沒有臉,就沒有角色。電影風的變體行為完全相同,所以這是平台層級的檢查,不是單一模型的怪癖。公道講,這次失敗不花錢,但它出現在提交看起來已經成功之後;這就是為什麼我們從不把「排隊中」當成功。

風格化人類:訓得很順,回來的是另一個人

Punk、動漫、插畫角色。這次訓練正常完成:25 credits、大約半小時、status: completed、沒有錯誤。然後我們用它生成一張近距肖像。髮色錯、眼睛顏色錯、星形胎記不見、衣領顏色錯。一張畫得不錯、但屬於另一隻角色的圖。

同一角色,同一平台 整體外觀保留的身分特徵
訓練式角色系統 25 credits · 30 分鐘 43.0%0%
附一張參考圖 不訓練 63.1%100%

沒有任何一個身分特徵撐過訓練。只是附上一張參考圖,全部撐過。道理說得通:這些系統鎖的是臉。對寫實人像來說,臉就是大部分的身分,鎖它是合理的取捨;對插畫角色來說,身分多半根本不在臉上,而是髮色、胎記、配色、衣領、配件。鎖住臉、放任其他一切漂移,得到的是一個骨架正確的陌生人。

我們在第二個無關的平台上看到同樣的模式:掛上它的角色物件,量出來並沒有比單純給起始影格好,兩種做法的特徵都維持 100%。兩個獨立平台一致,這件事就從平台怪癖變成我們對同類平台預設套用的規則。

什麼有效

參考機制:一份書面規格、一小組參考圖、按平台調整的措辭。對非人角色,這些做法相對於參考圖集自身上限的得分是 107%(Gemini)、105.5%(Midjourney)、110.8%(Veo):生成的影格和角色的一致程度,比五個極端角度拍的參考圖彼此之間還高。

這就是為什麼跟著登記角色走的平台備註,不會把你推向貴的那個選項。對非人角色,訓練路線根本不提供;對風格化角色,它會警告訓練量出來比不訓練更差。兩句警告刻意寫得不一樣:「這會失敗」和「這會成功但更差」不是同一個訊息,混在一起會讓人以為插畫角色在這裡做不了。

限制照例:每組五張圖、每種類型一隻角色、每個平台一個底座模型。當量級讀,不要當帶小數點的排名讀。完整實驗在量測筆記。

關於權利,因為有人問

代幣和授權是兩回事,鏈上只記錄前者。Bored Ape 持有人被授予自己那隻猿的廣泛商業權利;Nouns、Cryptoadz、mfers 是 CC0,意思是圖像任何人都能自由使用,包括不是你的任何人;許多其他系列只授予有限權利,或完全不授予。持有代幣不會告訴你自己屬於哪一種,系列的條款才會。用產出做任何商業用途之前,先讀它。這是對這類授權一般運作方式的描述,不是對你那個系列的意見。

在這裡登記角色不會改變上面任何一點。編號記錄的是這份規格從某一天起就存在、沒改過:是證據,不是權利。這個區分我們講得很明確,因為誇大它的登錄簿,是在開一張紀錄兌不了現的支票。

把你的角色帶進來 →  ·  到 Discord 跟這些數字吵一架