角色设定卡(也叫 model sheet、turnaround、三视图)是在一张图里从几个角度展示同一个角色。它是交给下一步的东西 —— 插画师、3D 美术,或者另一个要在多个场景里保持角色一致的 AI 模型。常见的三视图提示词会把三四个小小的全身像并排摆。本页的版式不一样:它把半张画布留给正脸胸像,因为后续生成里最容易跑掉的就是脸。
这套结构不是纸上谈兵。apimodels.app 上有一天用 gpt-image-2.5-sunburst 按这个版式连续生成了 400 多张设定卡,平均每张 56 秒。我们用自己的话重写了一遍,把可变部分单独标出来,并配了三个原创示例角色,让你看清每一块该怎么填。
脸占了一半像素。在 2048x1152 的设定卡上,正脸胸像大约 1024 像素宽,裁下来喂给别的模型时,眼型、皮肤质感、发际线都还在。四个小人并排时,每张脸只有几十像素高。
侧面和背面补上胸像看不到的东西:轮廓、体态、衣服怎么垂、鞋长什么样。要求两个全身像的头顶和脚底在同一条线上,能让它们保持同一比例,身材比例才看得准。
中性灰 #808080 背景和「不要文字、不要标签」这一句,比看上去重要。图像模型很喜欢自作主张加上「FRONT VIEW」这类歪歪扭扭的字;禁掉文字,设定卡才干净,纯灰背景后期也很好抠。
把角色写成可核对的事实,并按固定顺序:外观年龄、身高、体型、脸型、眼睛、鼻子、嘴唇、皮肤、头发、神态。「好看」「酷」这类词,模型没法在几个视角之间保持一致;「大大的青绿色眼睛、齐刘海」就可以。
服装给五个槽位:上装、一件露出来的内搭、下装、鞋、只要一件配饰,最后加一句「No handheld objects, no props」。道具是几个视角之间人物走样最常见的原因,而背面视角会把你没描述到的地方暴露出来。
第一栏的表情要写,别留空,可以轮换。好用的六个:neutral calm、gentle warm、soft thoughtful、focused determined、alert intense、quietly confident。一张卡只用一种表情。
生成场景时,把做好的设定卡当参考图传进去:用 gpt-image-2.5 做图像编辑,或者用 apimodels.app 上任何支持参考图的视频模型,比如 Wan 3.0。在提示词里说明这张参考是角色设定卡、脸、发型和服装都要和它一致;模型主要靠那张大幅正脸来锁定身份。
只有在本人同意的情况下,才为真实人物做写实设定卡。如果是原创角色,其他风格里也建议保留预设 C 里那句「不像任何真实人物」。
计费 = 分辨率档 × 画质档,默认画质是 high。只有成功出图才扣费。尺寸请用 16:9 —— 这个版式就是为横版画布设计的。
| 档位与 16:9 尺寸 | medium | high(默认) |
|---|---|---|
| 1K · 1360x768 | $0.020 | $0.030 |
| 2K · 2048x1152 | $0.025 | $0.040 |
| 4K · 3840x2160 | $0.045 | $0.080 |
Template — replace the four [BRACKETED] parts
Character reference sheet of one consistent character on a seamless neutral grey (#808080) studio background, clean editorial layout, landscape composition read left to right, split into three unequal vertical panels: PANEL 1 takes the entire left half of the frame (50% of the width); PANEL 2 and PANEL 3 share the right half at 25% each. Same identity, outfit, lighting and colour grading in every panel. No text, no labels, no logos, no watermark.
PANEL 1 (left half, by far the largest): chest-up portrait, straight front view, head and upper chest filling the panel, eyes in sharp focus with soft catchlights, [EXPRESSION] expression.
PANEL 2 (25% width, narrow): full-body side profile, standing relaxed and straight, arms hanging naturally, weight evenly balanced, head to toe inside the frame with even margins, showing silhouette, posture, garment volume and shoe profile.
PANEL 3 (25% width, rightmost): full-body back view, same standing pose, head to toe, showing how the hair falls, back posture, how the outfit fits from behind, and the shoe heels.
The two full-body figures share one scale: tops of the heads and soles of the feet sit on the same horizontal lines. Exactly one character, no duplicate figures, no other people, no background objects.
CHARACTER (identical in every panel): [age, height, build, face shape, eyes, nose, lips, skin, hair, overall demeanour]
WARDROBE (identical in every panel): [top, visible under-layer, bottoms, shoes, one accessory]. No handheld objects, no props.
LIGHTING & RENDER: [pick one style preset below]Style preset A — photoreal studio
LIGHTING & RENDER: soft even studio lighting, large diffused key light with gentle fill, soft natural shadows, no harsh highlights, true-to-life skin tones, neutral white balance, polished editorial model-sheet look, full-frame camera with an 85mm lens look, deep focus with the whole figure sharp in every panel, realistic pore and fabric texture.Style preset B — anime / manga
LIGHTING & RENDER: clean anime illustration style, crisp black line art with varied line weight, flat cel shading with soft gradient shadows, limited but vivid colour palette, no photorealism, consistent line quality and colour in every panel.Style preset C — 3D animated feature
LIGHTING & RENDER: stylized 3D animated feature-film character render, appealing exaggerated but coherent proportions, soft subsurface skin, detailed hair groom and cloth, soft global illumination with three-point studio lighting and a gentle rim light, deep focus in every panel. Original character design, not resembling any real person or existing copyrighted character.Example fill — photoreal (preset A, expression: focused determined)
CHARACTER (identical in every panel): Woman, apparent age 29, about 1.70 m, lean athletic build, oval face with a defined jaw, straight nose, medium lips, dark brown eyes, light olive skin with a few freckles across the nose, shoulder-length black hair tied in a low ponytail, alert focused demeanour.
WARDROBE (identical in every panel): Fitted charcoal cycling jacket with reflective piping and a high collar; a mustard merino base layer visible at the neck; black slim technical trousers rolled once at the ankle; white and grey trail running shoes; a small orange crossbody bag on one shoulder. No handheld objects, no props.Example fill — anime (preset B, expression: quietly confident)
CHARACTER (identical in every panel): Girl, apparent age 17, about 1.60 m, slim build, heart-shaped face, large teal eyes, small nose, light skin, long silver hair in a high ponytail with straight bangs, calm self-assured demeanour.
WARDROBE (identical in every panel): Navy short haori jacket with white wave pattern at the hem over a white high-collar blouse; pleated dark grey skirt; black knee socks and brown lace-up boots; a red cord tied around the left wrist. No handheld objects, no props.Example fill — 3D animated (preset C, expression: gentle warm)
CHARACTER (identical in every panel): Man, apparent age 64, about 1.66 m, round sturdy build, round face with full cheeks, big bulbous nose, wide smiling mouth, small twinkling brown eyes, warm tan skin with rosy cheeks, thick grey moustache, short grey hair with a bald crown, cheerful kindly demeanour.
WARDROBE (identical in every panel): Flour-dusted white baker's jacket with two rows of buttons; a striped blue-and-white apron tied at the back; loose brown trousers; black clogs; a small wooden pencil tucked behind one ear. No handheld objects, no props.cURL
curl -X POST https://api.apimodels.app/v1/images/generations \
-H "Authorization: Bearer $APIMODELS_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-image-2.5-sunburst",
"prompt": "Character reference sheet of one consistent character on a seamless neutral grey (#808080) studio background ... (paste the full template with your CHARACTER, WARDROBE and LIGHTING & RENDER blocks)",
"size": "2048x1152",
"response_format": "url"
}'Python
from openai import OpenAI
import os
client = OpenAI(api_key=os.environ["APIMODELS_API_KEY"], base_url="https://api.apimodels.app/v1")
TEMPLATE = open("character_sheet_template.txt").read() # the template above
def character_sheet(character, wardrobe, render, expression="neutral calm"):
prompt = (TEMPLATE
.replace("[EXPRESSION]", expression)
.replace("[age, height, build, face shape, eyes, nose, lips, skin, hair, overall demeanour]", character)
.replace("[top, visible under-layer, bottoms, shoes, one accessory]", wardrobe)
.replace("LIGHTING & RENDER: [pick one style preset below]", render))
img = client.images.generate(
model="gpt-image-2.5-sunburst",
prompt=prompt,
size="2048x1152", # 2K, $0.04 at the default high quality
response_format="url",
)
return img.data[0].url在一张图里从几个角度展示同一个角色 —— 通常是正面、侧面、背面 —— 每个视角的脸、服装和光线都一致。它是让角色在插画、3D 模型或 AI 生成场景里保持一致的参考。本页模板用的是一张大幅正脸胸像加全身侧面和背面。
先点明是设定卡,再用数字把版式定死,最后才写角色。上面的模板先写「character reference sheet」,把画布按 50/25/25 分开,要求正脸胸像、全身侧面、全身背面,让头顶脚底对齐,禁掉文字和道具,然后才填角色、服装和渲染风格。
能听懂长而结构化的提示词、版式渲染得干净的模型。在 apimodels.app 上,这个模板已经在 gpt-image-2.5-sunburst 上以 2048x1152 跑过几百次,每张约一分钟;gpt-image-2.5-flare 是更快更便宜的一档,适合打草稿。
用具体、可核对的特征描述角色,服装限制在五件、只要一件配饰,禁掉道具,并在角色和服装两块都写上「identical in every panel」。描述越具体,模型在几个视角之间能改动的余地就越小。
在 apimodels.app 上用 gpt-image-2.5-sunburst:2K(2048x1152)默认 high 画质一张 $0.04,medium $0.025,4K high $0.08。生成失败不扣费。
可以 —— 用 gpt-image-2.5 的图像编辑端点传入你的照片,沿用同一个模板,把角色那一块换成「参考图里的这个人」。只对你自己或已经同意的人这样做。