Method note · Open specification v1.3
Every rule below changed a measured score. Each one is stated with the failure that produced it, because a rule without its failure case is just an opinion about prompting.
Apply these by hand if you like — no tool required. The most valuable rule on this page costs nothing: an identity marker you cannot answer with yes or no is not an identity marker.
Where this sits now. OpenStela is a registry and market. The specification below is the format a character's definition takes when it is registered: the silhouette, structure, surface, palette and identity markers that describe it. Nothing on the registry measures works against it — the platform records declarations, it does not grade pictures. The prompt-writing rules come from the measurement work that preceded the registry and are kept here as a method note.
方法筆記 · 開放規格 v1.3
下面每一條規則,都曾經改變過一個量測分數。每一條都附上當初讓它成立的失敗案例,因為沒有失敗案例的規則,只是一個關於 prompt 的意見。
這些規則可以徒手套用,不需要任何工具。這一頁最有價值的那條規則一毛錢都不用花:一個沒辦法用「有」或「沒有」回答的身分特徵,就不是身分特徵。
這一頁現在的位置。OpenStela 是登錄簿與市集。下面這份規格,是角色登記時「定義」欄位的寫法基準:輪廓、結構、表面、配色、身分特徵。登錄簿上沒有任何機制拿作品去對照它打分,平台記錄的是聲明,不評斷畫面。至於各條 prompt 寫法,來自登錄簿之前的量測工作,留在這裡當方法筆記。
A JSON schema for characters is easy and almost worthless. The difficulty is not where the fields go; it is that most of what people naturally write about a character has no effect on a generative model.
"Elegant bearing", "an air of mystery", "futuristic" — these produce different images every time because they have no verifiable content. Meanwhile "a 2 cm pale scar above the outer end of the left eyebrow" survives lighting changes, scene changes and platform changes, because there is nothing to interpret.
So the specification is a set of rules about what to write. The file format is the trivial part.
替角色定一個 JSON schema 很容易,價值也很低。難的從來不是欄位放哪裡,而是人們自然寫下來的角色描述,大多對生成模型沒有作用。
「氣質優雅」「帶點神祕感」「未來風」,每次生出來的圖都不一樣,因為這些句子沒有可以驗證的內容。反過來,「左眉外端上方一道 2 公分的淡疤」撐得過換光線、換場景、換平台,因為沒有任何需要詮釋的空間。
所以這份規格其實是一組「該寫什麼」的規則。檔案格式是其中最不重要的部分。
| Layer | Holds | Fails as |
|---|---|---|
| Silhouette | Build, age range, ethnicity, hair length and shape | Gender and ethnicity drift between images |
| Structure | Face or body geometry that never changes | A similar-looking but different person |
| Surface | Fixed garment, materials, palette | The scene replaces the clothing |
| Markers | Two to four things nobody else has | Nothing to verify identity against |
Silhouette is the one most often skipped and the one that causes the worst failures. In a run where it was left blank, one of five reference images came back as a different ethnicity — the model had nothing to hold on to.
| 層 | 放什麼 | 失敗時的樣子 |
|---|---|---|
| 輪廓 | 體型、年齡區間、族裔、髮長與髮型 | 性別與族裔一張一張漂移 |
| 結構 | 永遠不變的臉部或身體幾何 | 一個長得很像、但不是同一個的人 |
| 表面 | 固定的服裝、材質、配色 | 場景把衣服換掉 |
| 特徵 | 兩到四件別人沒有的東西 | 沒有東西可以核對身分 |
輪廓是最常被跳過、也造成最糟失敗的一層。有一次實驗把它留白,五張參考圖裡有一張回來是另一個族裔,因為模型沒有東西可以抓。
Rules 1–6
Models follow enumerated constraints far more reliably than continuous description. Same content, different structure, different outcome.
✗ RIN, 25-year-old East Asian woman, asymmetric bob, hair #1C1A1F, 2cm scar above outer left eyebrow, amber irises, crimson high-collar jacket. Keep her identity consistent with the references. ✓ MANDATORY IDENTITY MARKERS, all must be visible and unchanged: 1. SCAR: a 2cm thin pale scar above the outer end of her LEFT eyebrow… 2. EYES: amber irises (#A8763E), distinctly lighter than typical… 3. HAIR: asymmetric bob, left side reaching the jawline…
Describing a feature is not the same as requiring it. The scar above disappeared entirely in a stage-lighting scene until the instruction said it must not.
1. SCAR: … This must be clearly visible in every image, never omitted, never covered by hair.
The most common failure of all. Ask for "performing on stage" and the model dresses the character for a stage, discarding the garment that identifies them. A generic "don't change the clothes" does not work — the negative list has to name the substitutes it expects.
4. GARMENT: a crimson (#C8102E) jacket with a STANDING HIGH COLLAR. This exact garment must be worn in every scene regardless of the setting. Do not substitute a stage outfit, sportswear, a hoodie, or any other clothing.
Recency weighting is real. Put the scene at the top and the identity constraints at the bottom, and close with an explicit statement of priority.
SCENE: {scene}
CHARACTER LOCK — the person in this image must be {name}, exactly as
shown in the attached reference images.
{numbered markers}
FORBIDDEN: {negative list}
The scene may change. {name} must not.
#1C1A1F holds. "Almost black with a blue cast" resolves to pure black
or to navy, differently each time.
Good: a 2 cm scar above the outer left eyebrow; nine tentacles of
unequal length; a brass ring at the tip of each tentacle.
Not markers: a cold demeanour; elegant posture; futuristic.
This is the single test that decides whether a specification will work. If two people could disagree about whether the image contains it, it is not a marker.
規則 1–6
模型遵守逐條列出的限制,比遵守一段連續描述可靠得多。同樣的內容、不同的結構、不同的結果。
✗ RIN,25 歲東亞女性,不對稱鮑伯頭,髮色 #1C1A1F, 左眉外端上方 2cm 疤,琥珀色虹膜,深紅高領外套。 請維持她與參考圖一致的身分。 ✓ MANDATORY IDENTITY MARKERS, all must be visible and unchanged: 1. SCAR: a 2cm thin pale scar above the outer end of her LEFT eyebrow… 2. EYES: amber irises (#A8763E), distinctly lighter than typical… 3. HAIR: asymmetric bob, left side reaching the jawline…
描述一個特徵,不等於要求它出現。上面那道疤在舞台燈光的場景裡整個不見,直到指令寫明它不可以消失。
1. SCAR: … This must be clearly visible in every image, never omitted, never covered by hair.
最常見的失敗。要求「在舞台上表演」,模型就替角色換上舞台裝,把用來辨認他的那件衣服丟掉。籠統寫「不要改衣服」沒有用,禁止清單得點名它預期會出現的替代品。
4. GARMENT: a crimson (#C8102E) jacket with a STANDING HIGH COLLAR. This exact garment must be worn in every scene regardless of the setting. Do not substitute a stage outfit, sportswear, a hoodie, or any other clothing.
近因效應是真的。把場景放最上面、身分限制放最下面,最後用一句明確的優先順序收尾。
SCENE: {scene}
CHARACTER LOCK — the person in this image must be {name}, exactly as
shown in the attached reference images.
{numbered markers}
FORBIDDEN: {negative list}
The scene may change. {name} must not.
#1C1A1F 撐得住。「近乎黑、帶一點藍」每次會落成純黑或海軍藍,而且每次不一樣。
算特徵的:左眉外端上方 2 公分的疤、九條長度不一的觸手、每條觸手末端一枚黃銅環。
不算的:冷淡的態度、優雅的姿態、未來風。
這是判斷一份規格能不能用的唯一測試。如果兩個人可能對「圖裡有沒有這個東西」意見不同,它就不是特徵。
Rule 7
The only rule so far that moved a score from failing to passing on its own.
| Anchor version | Face score (threshold 0.363) |
|---|---|
| v1 — appearance only | 0.334 ✗ |
| v2 — structure + specific negative list | 0.394 ✓ |
Only the text changed; input image, model, duration, aspect ratio, resolution and scene text were all held fixed.
The failure. v1 said "matte black high-collar technical jacket with a single asymmetric cyan seam". The model kept the colour and decoration — black, high collar, cyan seam — and lost the structure: the open-front jacket became a pullover, the zip vanished, the stiff standing collar became a soft rolled neck.
1. A matte black technical JACKET that OPENS DOWN THE FRONT with a full-length front zip. This is an open-front jacket, NEVER a pullover, NEVER a sweater, NEVER a hoodie, NEVER a turtleneck. 2. The collar is a STIFF STANDING collar that holds its shape, NEVER a soft rolled neck, NEVER a crew neck.
Two points. Describe how it is constructed — opens at the front, has a zip, holds its shape — rather than how it looks; models hold colour easily and structure poorly. And name the structural substitutes explicitly, because the thing to defend against is not only "replaced by the scene" but "replaced by something the same colour with a different construction".
One more line draws the boundary between what may vary and what may not:
Stage lighting may change the colour of the light, but must not change the garment's cut, structure or fastening.
規則 7
到目前為止,唯一一條單靠自己就把分數從不及格拉到及格的規則。
| 錨句版本 | 人臉分數(門檻 0.363) |
|---|---|
| v1:只寫外觀 | 0.334 ✗ |
| v2:寫結構,加上點名的禁止清單 | 0.394 ✓ |
只有文字改了;輸入圖、模型、時長、長寬比、解析度、場景文字全部固定。
失敗長這樣。v1 寫的是「霧面黑色高領機能外套,單一不對稱青色縫線」。模型保住了顏色和裝飾:黑、高領、青色縫線;卻丟掉了結構:前開的外套變成套頭、拉鍊不見、挺立的立領變成軟塌的翻領。
1. A matte black technical JACKET that OPENS DOWN THE FRONT with a full-length front zip. This is an open-front jacket, NEVER a pullover, NEVER a sweater, NEVER a hoodie, NEVER a turtleneck. 2. The collar is a STIFF STANDING collar that holds its shape, NEVER a soft rolled neck, NEVER a crew neck.
兩個重點。寫它怎麼構成:前開、有拉鍊、能撐住形狀,而不是它長什麼樣,因為模型很會保住顏色、很不會保住結構。另外要點名結構上的替代品,因為要防的不只是「被場景換掉」,還有「被同色但不同構造的東西換掉」。
再加一行,把可變與不可變的界線畫清楚:
Stage lighting may change the colour of the light, but must not change the garment's cut, structure or fastening.
Rule 8
| Anchor version | Mean | Frames passing | Worst frame |
|---|---|---|---|
| v1 baseline | 0.334 ✗ | 2/8 | 0.212 |
| v2 + garment structure | 0.394 ✓ | 7/8 | 0.295 |
| v3 + pose constraint | 0.424 ✓ | 7/8 | 0.346 |
The worst frame rising from 0.212 to 0.346 matters more than the mean. It means the whole clip became usable, rather than a few good frames pulling an average up.
Diagnosing it correctly took two wrong guesses first. The low frames looked like decay over time, then like insufficient light. Frame-by-frame comparison killed both: brightness had no correlation with score — the worst frame was the brightest — and the garment and markers were intact in every frame. The actual cause was head pose. When the character tilted back to look at the ceiling, the score fell from 0.487 to 0.295. The reference set is all level, forward-facing; a steeply upturned face is simply far away in embedding space.
POSE — he nods and sways with the beat and may glance down at the controller, but his head stays roughly level and mostly toward the camera. He never tips his head far back, never looks up at the ceiling, and his face is never turned more than 45 degrees away from the lens.
Three points. Describe the permitted movement first, then the prohibitions — a list of prohibitions alone produces a stiff performance. Use a verifiable quantity (45 degrees) rather than an adjective. And diagnose frame by frame: acting on either of the first two guesses would have cost money and fixed nothing.
Generalisation check. The same anchor was run in a scene with the opposite lighting conditions — bright daylight and hard shadows instead of a dark room with coloured sweeps. Mean 0.414 versus 0.424, worst frame 0.352 versus 0.346. A gap of 0.010: the rules are not accommodating one particular scene.
But the pose wording must be scene-independent. The first version said "may glance down at the controller", which tied it to one setting. A rule that mentions objects only present in one scene has to be rewritten for the next one, and is therefore not a rule.
規則 8
| 錨句版本 | 平均 | 通過影格 | 最差影格 |
|---|---|---|---|
| v1 基本版 | 0.334 ✗ | 2/8 | 0.212 |
| v2 加服裝結構 | 0.394 ✓ | 7/8 | 0.295 |
| v3 加姿勢限制 | 0.424 ✓ | 7/8 | 0.346 |
最差影格從 0.212 升到 0.346,這比平均值更重要。它代表整段影片變得能用,而不是靠幾個好影格把平均拉高。
診斷對之前,先猜錯了兩次。低分影格一開始看起來像隨時間衰退,後來又像光線不足。逐格比對推翻了兩個猜測:亮度和分數沒有相關,最差的影格反而最亮,而且服裝與特徵每一格都完整。真正的原因是頭部姿勢:角色仰頭看天花板時,分數從 0.487 掉到 0.295。參考圖集全是水平、正面的角度,一張大幅仰起的臉在嵌入空間裡本來就很遠。
POSE — he nods and sways with the beat and may glance down at the controller, but his head stays roughly level and mostly toward the camera. He never tips his head far back, never looks up at the ceiling, and his face is never turned more than 45 degrees away from the lens.
三個重點。先寫允許的動作再寫禁止,只寫禁止會生出僵硬的表演。用量得出來的數字(45 度)不用形容詞。以及,要逐格診斷:照前兩個猜測去改,錢花了、問題還在。
泛化檢查。同一份錨句放到光線條件相反的場景跑過一次:明亮日光加硬陰影,取代有色光掃過的暗房。平均 0.414 對 0.424,最差影格 0.352 對 0.346。差 0.010,代表這條規則不是在遷就某一個場景。
但姿勢的寫法必須和場景無關。第一版寫「可以低頭看控制器」,等於把它綁在一個場景上。一條提到只存在於某個場景的物件的規則,換到下一個場景就得重寫,那就不是規則。
Rules 9–10
Images generated with an earlier image as reference cannot be pooled with text-only generations when computing the consistency ceiling. The ceiling is supposed to measure how clearly the specification defines the character; once images are conditioned on each other it measures how well image-conditioning works, which is a different question. Mixed together, every rule change looks effective because the images are doing the work.
If you do chain images, use a star topology — condition everything on one chosen anchor image, never a chain, which accumulates drift.
Generating several candidates under different anchor variants and picking one produces a paired comparison under identical conditions — far stronger than assigning variants at random across separate runs. Two conditions make it valid: allocate variants evenly and then shuffle within the batch, so variant and position are not bound together; and record the position, because people prefer the first option and that bias has to be measurable rather than assumed absent.
Keep the rejected candidates. They were paid for, and negative examples are exactly what rule development lacks.
規則 9–10
以前一張圖為參考生出來的圖,計算一致性上限時不能和純文字生成的圖混在一起。上限本來要量的是「規格把角色定義得多清楚」;一旦圖和圖互為條件,量到的就變成「圖像條件化做得多好」,那是另一個問題。混在一起,每一次規則改動看起來都有效,因為工作其實是圖在做。
真的要串接圖像,用星狀:一切以一張選定的錨點圖為條件。絕不用鏈狀,那會累積漂移。
用不同錨句變體生出多個候選再挑一個,得到的是相同條件下的配對比較,遠比在不同實驗之間隨機分配變體有力。兩個條件讓它成立:平均分配變體之後在批次內洗牌,讓變體和位置脫鉤;同時記錄位置,因為人會偏好第一個選項,這種偏誤要量得出來,不能假設它不存在。
留著被否決的候選。它們是花錢換來的,而負例正是規則開發最缺的東西。
Reference set
On the registry the order is reversed. Your reference set is the pictures you already have: you upload them, and the specification is read out of them and then corrected by you. In the measurement work it ran the other way — the set was produced from a written specification — and that ordering had one advantage worth carrying over. A set collected first contains a hundred incidental facts that nobody wrote down, and those are exactly the ones that drift.
So the step the old ordering gave you for free is now one you take by hand: after the draft specification comes back, look at your own pictures again and write down what is true of them that the draft did not mention.
Five views for a human figure — front, three-quarter, full profile, back of head, extreme close-up. For non-human subjects the axis changes to angle, structural detail, posture and lighting. Upload what you have; the list is what to reach for, not a requirement.
Keep full-body views as a separate set rather than mixed in with head-and-shoulders. In the measurements, mixing framings inflated the spread of the set and made every normalised score look better than it was; in practice it also makes the set harder to hand to another platform as a single reference.
參考圖集
在登錄簿上,順序是反過來的。參考圖集就是你手上已經有的圖:你上傳,規格從圖裡讀出來,再由你改對。當年的量測工作是另一個方向——圖集是從書面規格產出的——那個順序有一個好處值得帶過來:先收集的圖集會含有上百個沒人寫下來的偶然事實,而那些正是會漂移的東西。
所以舊順序白送你的那一步,現在要自己動手:規格草稿回來之後,再看一次自己的圖,把草稿沒提到、但圖上確實成立的事寫下來。
人形有五個視角:正面、四分之三側、全側面、後腦、極近特寫。非人主體的軸換成角度、結構細節、姿態與光線。有什麼就傳什麼——這份清單是可以往哪裡找,不是門檻。
全身視角另存成獨立一組,不要和頭肩構圖混在一起。量測時,混合取景會擴大圖集的離散度、拉高上限,讓每一個正規化分數看起來比實際好;實務上,混在一起的圖集也比較難當成單一份參考交給別的平台。
Boundaries
邊界