Short answer: yes, you can call Space Bunny Alpha through apimodels.app. Send POST https://api.apimodels.app/v1/chat/completions with your apimodels key and "model": "space-bunny-alpha"; the OpenAI SDKs work with only base_url and the key changed. It costs $0.10 per 1M input tokens and $0.40 per 1M output tokens, cached input $0.01, billed on actual usage with reasoning tokens counted as output and failed calls free.
Nobody knows who makes it. The model is listed under a "Stealth" provider with a 1M-token context window, up to 524K output tokens, adjustable reasoning_effort, tool calling, response_format and image and video input. Stealth releases are how labs collect real-world feedback before a launch; they are usually renamed or withdrawn once the maker reveals itself, so treat the id as temporary and keep a fallback model in your code.
Instead of repeating the listing, we tested it. Twenty-eight tasks across 3D graphics, web pages, coding, video explainers, reasoning, writing in four languages, structured output, tool calling, image and video input, a 91K-token document and the reasoning_effort setting, one real API call each, every artifact opened in Chrome or checked by hand. What follows is what we saw, including the parts that did not work.
The listing went up on OpenRouter on 23 September 2026 as stealth/space-bunny-alpha, priced at $0 during the preview. On apimodels.app the same model is space-bunny-alpha at $0.10 / $0.40 per 1M tokens: there is no official list price to compare against, because the price is the gateway, the shared key and the billing, not the model. If you already have an OpenRouter account with credit, the model is free there and the only reasons to call it here are one key for everything else, a single balance and the same OpenAI-compatible endpoint as our other 46 language models.
The provider's stealth terms say prompts and completions may be retained by them but are not used for training. Rate limits belong to the preview, not to us: sustained bursts can return HTTP 429, and a 404 with "no provider" would mean the preview has ended.
| Item | apimodels.app | OpenRouter (preview) |
|---|---|---|
| Model id | space-bunny-alpha | stealth/space-bunny-alpha |
| Endpoint | POST https://api.apimodels.app/v1/chat/completions | POST https://openrouter.ai/api/v1/chat/completions |
| Input / output price per 1M tokens | $0.10 / $0.40, cached input $0.01 | $0 / $0 while the preview lasts |
| Context / max output | 1M tokens / 524K tokens | Same |
| Reasoning | reasoning_effort; trace returned; reasoning tokens billed as output | Same |
| Tools / response_format | Tool calling verified; schema not enforced (see below) | Listed as supported |
| Image / video input | Image works; video arrives as sampled frames, no audio | Listed as supported |
| Failed calls | Not billed | Free |
Every case was one call on 29 September 2026 with a 200,000-token max_tokens budget so the reasoning trace could not starve the answer. Verdicts: pass means it does what was asked and runs or is correct; partial means it runs or answers with defects we could name; fail means it does not run or is wrong. HTML artifacts were opened in Chrome with the console watched; tests were run with Node 24; SQL was run on PGlite; the Remotion composition was type-checked against the real remotion and React packages; numeric answers were checked by hand.
The pattern across groups is consistent: generation is strong and verification is absent. Everything visual it produced worked; every coding task had a small defect the model could have caught by running its own code; it reports word counts, test suites and "exact outputs" that it never executed. Treat what it returns as untested source.
| Group | Cases | Pass | Partial | Fail | What decided it |
|---|---|---|---|---|---|
| 3D / graphics | 4 | 4 | 0 | 0 | Three.js solar system and product showcase, pure-CSS flip cards, SMIL logo intro: all render, no console errors |
| Web / UI | 2 | 2 | 0 | 0 | 58 KB landing page and 64 KB Chart.js dashboard, every control verified in Chrome |
| Coding | 4 | 0 | 4 | 0 | Logic right every time; one missed bug, 3 wrong test assertions, a garbled SQL token, 4 wrong hand-computed quantiles |
| Video / explainers | 3 | 2 | 0 | 1 | Chinese storyboard and Remotion composition pass; canvas explainer is one missing quote from passing |
| Reasoning | 3 | 3 | 0 | 0 | Number theory correct, under-constrained puzzle correctly called non-unique, itinerary meets every constraint (medium effort re-run) |
| Writing | 4 | 4 | 0 | 0 | Metrically correct quatrain, byte-faithful translation, 2,300-word article, native-quality Japanese and Korean copy |
| Structured output | 2 | 1 | 0 | 1 | Parallel tool calls perfect; strict JSON schema silently not enforced |
| Multimodal | 2 | 1 | 1 | 0 | Image understanding excellent; video arrives as a few frames, no audio |
| Long context | 1 | 1 | 0 | 0 | 3 of 3 needles from 91,298 tokens with exact quotes |
| Effort / speed | 3 | 3 | 0 | 0 | Same correct answer at low, default and high; 11 s of reasoning before an 86-word paragraph |
The first pair is the same prompt at reasoning_effort low and medium: both pages are complete and both render, and the difference is 434 versus 9,301 characters of reasoning before the first byte of HTML. The second pair is one artifact twice: as the model returned it, and with a single missing quotation mark restored. The model does not run what it writes; you have to.
Both render correctly. low reached the first token in 3.8 s, medium in 23 s; the default never finished (cut at 303 s after 110,334 characters of reasoning).
Six cards flip in 3D on hover with a sheen sweep and a prefers-reduced-motion block, no JavaScript in either. Output length is almost identical (25,057 vs 24,611 characters); what changes is the reasoning trace, by twenty times.
Write a single-file HTML page with pure CSS (no JS) 3D effects: a grid of six pricing cards that flip on hover in 3D with perspective, a subtle tilt that follows a CSS-only hover, layered depth using transform-style: preserve-3d, a glossy sheen sweep, and reduced-motion support. Mobile responsive. Return only the HTML.
As returned, the page is blank: one closing quote is missing in 1,450 lines, so the browser reports a SyntaxError and nothing runs. With that character restored, five scenes, captions, seek and play/pause all work.
The fixed version still has layout defects the model did not see: the two header strings overlap, and a syringe illustration sits on top of the headline in scene 1. The science in the captions is careful (temporary recipe, DNA unchanged, antigen-presenting cells, memory cells).
Make a single-file HTML "explainer video" about how mRNA vaccines train the immune system, 30 seconds long, 1280x720, rendered on a canvas with requestAnimationFrame: 5 scenes with smooth transitions, simple vector illustrations drawn in code (cell, mRNA strand, ribosome, spike protein, antibodies), animated captions with a timeline, a play/pause button, a progress bar, and an "Export WebM" button that records the canvas with MediaRecorder. Scientifically accurate captions. No external assets. Return only the HTML.
Both ran unmodified. The solar system has orbit controls, labels that follow each planet and a 0.1x to 10x speed slider; the dashboard has a seeded PRNG so its 30-day chart is identical on every load, a sortable and filterable 40-row table, and dark mode persisted in localStorage.
Generation times at medium effort: 73 s for the solar system (10,121 output tokens) and 247 s for the dashboard (18,895 output tokens), the slowest artifact of the set. Neither produced a console error.
Build a single-file HTML page (Three.js from a CDN, no build step) showing an animated solar system: the Sun with an emissive glow, the 8 planets with correct relative orbit order, distinct sizes and colors, orbit rings, Saturn with a ring, Earth with a moon, a starfield background, OrbitControls for mouse rotation/zoom, a floating label that follows each planet, and a speed slider (0.1x–10x). / Create a single-file admin dashboard in HTML/CSS/JS using Chart.js from a CDN: a collapsible sidebar, top bar with search, four KPI tiles with sparklines, a line chart of daily revenue for 30 days and a doughnut of traffic sources (seeded PRNG), a sortable and filterable data table of 40 orders with pagination, and a dark/light toggle persisted in localStorage.
This model reasons before it answers, streams the reasoning first, and bills it inside completion_tokens. At the default effort the reasoning is enormous on anything generative: our three first coding prompts produced 129,262, 58,603 and 110,334 characters of reasoning, reached the first byte of the answer after 284 and 276 seconds (the second never did), and all three streams were cut by the provider around 300 seconds with no finish_reason. The same happened later to a three-day itinerary (94,085 characters, cut at 303 s), and a four-line poem spent 226 seconds reasoning before its 352 characters of text.
Sending reasoning_effort fixes it without visible cost. On the CSS-cards prompt, low used 434 characters of reasoning and medium 9,301, both finished in about 55 seconds, and both pages render identically well. On an easy, well-posed question the setting changes nothing: the number-theory problem got the same correct answer at low, default and high with 415, 499 and 452 characters of reasoning. The knob only bites on open-ended or generation-heavy prompts, which is exactly where you need it.
| Effort (same CSS-cards prompt) | Reasoning chars | First content token | Total | Result |
|---|---|---|---|---|
| low | 434 | 3.8 s | 53 s | renders, no errors |
| medium | 9,301 | 23 s | 55 s | renders, no errors |
| default (not sent) | 110,334 | 276 s | cut at 303 s | unfinished HTML |
Strict JSON schema. We sent response_format with a json_schema, strict: true, additionalProperties: false and six required keys. The reply was valid JSON with the right vendor, invoice number, three line items and total, but date came back as issue_date, total as amount_due plus subtotal, and every line item carried a line_total the schema forbids. Nothing on this route enforced the schema. Validate the response against your schema and retry on failure; tool calling, which returned two parallel calls with exact arguments in 2 seconds, is the reliable way to get structure out of this model today.
Video input. We sent a 2.02-second, 60-frame clip with an AAC soundtrack. The model described the opening overlook and the closing water-level view correctly, but it received timestamps that ended at 0.2 seconds and no audio, so it estimated the clip at 0.2 to 0.3 seconds and could not say whether there was sound. Video reaches the model as a few sampled frames. Ask it what is in a clip; do not ask it about duration, motion or sound.
Self-verification. Two of thirteen code artifacts were broken by exactly one character (a stray token in a SQL join predicate, a missing quote in 1,450 lines of JavaScript); three of the thirteen tests it wrote for a correct cache assert the opposite of its own code; the "exact table" it claimed a pandas script prints had four wrong quantiles out of twelve; a 2,300-word article was labelled 2,500 words. Nothing it produced was executed on its side. Where we ran things ourselves, the underlying logic held up every time.
Use the OpenAI-compatible chat endpoint with your apimodels key; the samples below are what we actually send. Two settings: pass reasoning_effort ("low" or "medium" for code, pages, plans and long writing; leave the default only for short questions), and do not cap max_tokens. The gateway forwards max_tokens only when you send it, so omitting it lets the upstream apply the model's own 524K limit; a small value returns empty content because the reasoning trace consumes it first. Stream long generations: non-streaming requests that run past a few minutes come back empty.
Throughput is roughly 90 to 140 tokens per second including the reasoning stream, so a 16K-token page takes about 75 seconds at medium. The response carries the reasoning as a separate field alongside content; usage reports reasoning inside completion_tokens (reasoning_tokens is reported as 0), which is what you are billed on.
cURL
curl https://api.apimodels.app/v1/chat/completions \
-H "Authorization: Bearer $APIMODELS_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "space-bunny-alpha",
"reasoning_effort": "medium",
"stream": true,
"messages": [
{"role": "user", "content": "Build a single-file HTML page with a pure-CSS 3D flip-card grid. Return only the HTML."}
]
}'
# No max_tokens: the gateway forwards nothing and the model's own 524K limit applies.
# Reasoning streams first (delta.reasoning), then the answer (delta.content).Python
from openai import OpenAI
client = OpenAI(base_url="https://api.apimodels.app/v1", api_key="YOUR_APIMODELS_KEY")
stream = client.chat.completions.create(
model="space-bunny-alpha",
messages=[{"role": "user", "content": "Build a single-file HTML page with a pure-CSS 3D flip-card grid. Return only the HTML."}],
stream=True,
extra_body={"reasoning_effort": "medium"}, # low or medium for code, pages, plans
# max_tokens omitted on purpose: the reasoning trace is billed inside completion_tokens
)
for chunk in stream:
delta = chunk.choices[0].delta
if delta.content:
print(delta.content, end="", flush=True)Yes. Send POST https://api.apimodels.app/v1/chat/completions with your apimodels key and "model": "space-bunny-alpha". The request and response are OpenAI Chat Completions, so the OpenAI SDKs work with only base_url and the key changed. It costs $0.10 / $0.40 per 1M tokens, cached input $0.01, and failed calls are free.
Nobody has said. It is listed under an anonymous "Stealth" provider on OpenRouter since 23 September 2026, a common way for labs to gather feedback before a launch. The response carries provider: "Stealth". Expect the model to be renamed or withdrawn when its maker reveals it, and keep a fallback in your code.
Almost always max_tokens. The reasoning trace streams before the answer and is billed inside completion_tokens, so a cap of a few thousand tokens is consumed entirely by reasoning and content stays empty with finish_reason length. Omit max_tokens or set it large, and send reasoning_effort low or medium so the trace stays short. A second cause is a non-streaming request that runs past a few minutes: use stream: true for long generations.
Good at producing, poor at checking. All six 3D and web artifacts ran unmodified in Chrome with no console errors, and a Remotion composition type-checked under strict TypeScript against the real packages. But all four pure coding tasks were partial: it missed one bug, wrote three tests that contradict its own correct code, left a garbled token in a SQL query and mis-computed four quantiles it claimed a script would print. Run, lint and type-check everything it returns.
response_format is accepted but a strict schema was not enforced in our test: required keys were renamed and forbidden extra fields appeared, although the JSON itself was valid and the facts were right. Validate on your side and retry, or use tool calling, which returned two parallel calls with exact arguments in 2 seconds.
It depends on reasoning_effort far more than on the prompt. At low or medium, first content token arrived in 4 to 40 seconds across our coding and page tasks and a 16K-token page took about 75 seconds end to end; throughput is roughly 90 to 140 tokens per second including reasoning. At the default effort, generation-heavy prompts reasoned for 4 to 5 minutes before the first character and 4 of 28 calls were cut upstream around 300 seconds. Short questions answer in a few seconds either way.