Timelines filled up with product films and beat-synced clips labelled "made with Opus 5.5" the week it launched. What is happening underneath matters for anyone trying to reproduce them: the model writes an animation program, and a renderer turns that program into video. That makes the result deterministic and editable — you can change one scene by asking for a code change at a timecode — but it also means you need a rendering step, and there is no sound unless you add it.
One streamed call to claude-opus-5-5 with no max_tokens; the first character of the answer arrived after 110 seconds because the model thinks first. The HTML it returned rendered all 360 frames without a single error, and four checkpoint frames matched the storyboard on the first try: clutter at 2.3 s, one key at 3.5 s, model cards at 7.2 s, end card at 11.3 s.
| Item | Value |
|---|---|
| Model | claude-opus-5-5 via api.apimodels.app |
| First token / total | 110 s / 6 min 37 s |
| Output tokens | 31,786 |
| Billed | $0.368 |
| Render | 360 frames in 14 s, headless Chrome |
| Video | 12 s, 1280x720, 30 fps, 1.8 MB MP4 |
Left: the final frame. Right: four checkpoint frames against the storyboard. All from the untouched HTML the model returned.
The full MP4, the verbatim prompt, the render and encode scripts and the HTML are on the Claude Opus 5.5 model page.
A 12-second promo for "APIMODELS" rendered by window.renderFrame(t) on a 1280x720 canvas, 120 BPM beat plan: clutter of API keys, snap into one glowing key, model cards lighting up per beat, end card "One key. Every model."
Ask for a program, not a video: one canvas, a global renderFrame(t) that draws frame t from scratch with no clock or randomness, and a fixed duration and frame rate. Describe the story in time ranges before the style, and give a tempo so cuts land on beats. Stream the call with a large or no max_tokens and reasoning at the default; expect a couple of minutes before the first character.
Then open the page in a headless browser, call renderFrame(i / fps) for every frame, save each canvas as PNG, and encode with ffmpeg. Add a soundtrack afterwards at the tempo you gave the model. Always look at a few frames before judging: the model does not run its own code.
cURL
curl https://api.apimodels.app/v1/messages \
-H "x-api-key: $APIMODELS_API_KEY" \
-H "anthropic-version: 2023-06-01" \
-H "Content-Type: application/json" \
-d '{"model": "claude-opus-5-5", "max_tokens": 16000,
"messages": [{"role": "user", "content": "Summarise the risks in this pull request: ..."}]}'Not as video files. It writes code that draws each frame of an animation, and a renderer such as headless Chrome plus ffmpeg turns that into an MP4. In our run one call produced a 12-second 1280x720 promo for $0.368.
No. The program only draws pictures. Add music or voice afterwards with ffmpeg; giving the model a tempo in the prompt lets you line the cuts up with the beat.
Our 12-second clip took one call of 6 minutes 37 seconds, with the first character after 110 seconds, and 31,786 output tokens billed at $0.368 on apimodels.app ($2.40 / $12 per million). Rendering 360 frames took 14 seconds on a laptop.
For realistic footage, yes: text-to-video models such as Seedance or Veo produce camera-like clips. Code-rendered video from Opus 5.5 suits motion graphics, product explainers, UI animations and anything with exact text, shapes and timing that must be editable.