Space Bunny Alpha Is Probably MiniMax M3.1-Flash
MiniMax-M3.1-Flash-Preview at roughly 85% confidence: strong, converging, not confirmed.On 23 September 2026 a model called Space Bunny Alpha appeared on OpenRouter with no lab attached to it. The provider field just said Stealth. It was free, it had a million-token context window, it accepted video as input, and it would not tell you who built it.
Its listing carries expiration_date: "2026-10-05" — today. So this is not a "try it before it's gone" piece. It is the identification and the post-mortem: what the model was, what the evidence says about who shipped it, how strong that evidence is, and how to run the analysis yourself next time. That last part matters more than the answer. We did this in August with a stealth model called Ox Alpha and concluded from its parameter surface that it was Z.ai's GLM-5.3 generation. It was. The method travels; the specific answer does not.
What exactly is space-bunny-alpha?
Every figure below was pulled from OpenRouter's live model API on 5 October 2026 — https://openrouter.ai/api/v1/models and the per-model /endpoints route.
| Field | Value |
|---|---|
| Model ID | stealth/space-bunny-alpha |
| Provider | Stealth — the only model listed under it |
| Created | 2026-09-23, 14:48 UTC |
| Price | $0 input / $0 output |
| Context window | 1,000,000 tokens exactly |
| Max output | 524,288 tokens |
| Input modalities | text, image, video |
| Output modalities | text |
| Reasoning | Mandatory — cannot be switched off |
| Reasoning efforts | max, xhigh, high, medium, low — default max |
| Tokenizer | "Other" (i.e. not one OpenRouter recognises) |
| Moderation | Unmoderated |
| Tool choice | auto only — none, required and function all unsupported |
| Quantization | Unknown (not disclosed) |
| Uptime, trailing 24h | 99.87% |
| Listed expiry | 2026-10-05 |
Two of those numbers do almost all the identification work: the 524,288 output cap, and the context length written as a round decimal 1,000,000 rather than the more common 1048576 (1024 × 1024). Labs that think in binary report 1,048,576; labs that think in round marketing numbers report 1,000,000. Mixing both conventions in one listing is itself a habit worth noticing.
One correction to the record: although the listing expires today, the endpoint was still serving requests at zero cost when this article was written. An expiry date on an OpenRouter listing is a scheduled delisting, not a cutoff that has already happened.
Who is behind space-bunny-alpha?
Nobody has said. There is no statement from any lab and none from OpenRouter, whose page still describes a provider that "has chosen to remain anonymous during this preview." Anyone telling you it is confirmed is overselling. What exists is a set of converging fingerprints, and they point at MiniMax — specifically at MiniMax-M3.1-Flash-Preview, not the already-public M3.
Here is each line of evidence with its strength labelled honestly.
Evidence 1: the 524,288 output cap is unique — strong
We pulled all 466 models on OpenRouter and checked every max_completion_tokens value. Exactly one reports 524,288: space-bunny-alpha. No other entry in the catalogue shares it. The nearest neighbours are 512,000 (MiniMax M3) and 943,718 (the GLM-5.3-Flash and DeepSeek V4 families).
Evidence 2: MiniMax's own documentation describes a model with this exact surface — strongest
This is the piece that moves the argument, and it is first-party. MiniMax's developer documentation at platform.minimax.io documents a model called MiniMax-M3.1-Flash-Preview that is not on OpenRouter, not on Hugging Face, and not in MiniMax's own release notes. It was shipped quietly into a gated tier: "MiniMax-M3.1-Flash-Preview is available only through M Plan and MiniMax Code for now."
Its published specification matches space-bunny-alpha on five independent fields:
| Field | MiniMax docs, M3.1-Flash-Preview | space-bunny-alpha |
|---|---|---|
| Context window | 1,000,000 | 1,000,000 ✓ |
| Thinking | Mandatory — effort: "none" returns HTTP 400, "requires adaptive thinking" | Mandatory ✓ |
| Effort ladder | low, medium, high, xhigh, max | Identical ✓ |
| Default effort | max when omitted | max ✓ |
| Input modalities | "text, images, and video" | text, image, video ✓ |
What makes this more than a loose resemblance: we checked all 466 OpenRouter models for that combination — mandatory reasoning, exactly the five-rung ladder, and a max default. space-bunny-alpha is the only one. Twenty-eight models share the ladder; every OpenAI and Anthropic entry among them defaults to medium or high. MiniMax's docs are the only place we found max documented as a default.
MiniMax does not publish a maximum output figure or pricing for M3.1-Flash-Preview, so the 524,288 cap cannot be checked against it directly.
Evidence 3: an abnormally large fixed prompt overhead shared with MiniMax — strong, and reproducible
This test we ran ourselves, and it is the most interesting result here. Send a two-character message — just hi — through OpenRouter and read back usage.prompt_tokens. Most models return a single-digit or low-double-digit number, because almost nothing is being added to your prompt. A large number means the serving stack is wrapping your message in a big hidden template.
| Model | prompt_tokens for "hi" |
|---|---|
| minimax/minimax-01 | 712 |
| minimax/minimax-m1 | 447 |
| minimax/minimax-m3 | 164 |
| stealth/space-bunny-alpha | 157 |
| moonshotai/kimi-k3 | 86 |
| qwen/qwen3.8-flash | 62 |
| minimax/minimax-m2.7 | 42 |
| bytedance-seed/seed-2.0-lite | 34 |
| z-ai/glm-5.3-flash | 13 |
| anthropic/claude-opus-5.5 | 10 |
| xiaomi/mimo-v2.6-pro | 8 |
| openai/gpt-6-sol | 7 |
| deepseek/deepseek-v4-flash | 5 |
| google/gemini-3.8-flash | 1 |
Every model above 100 tokens of overhead in this sweep is a MiniMax model — except space-bunny-alpha, which lands seven tokens from MiniMax M3. Every other lab tested sits at 86 or below, most in single digits. We ran each measurement twice and got identical numbers, so this is not noise.
Note what it also rules out. GLM-5.3-Flash comes back at 13. If this were a repeat of the Ox Alpha situation — Z.ai testing another checkpoint anonymously — the overhead would look like Z.ai's. It does not. Nor does it look like OpenAI (7), Gemini (1) or Claude (10).
Two independent testers published token-count tables pointing the same way, in which MiniMax M2 and M3 reproduced Space Bunny's per-probe token deltas exactly across 7 and 24 test strings. Both noted their counts came from a gateway rather than the provider directly, and that the test cannot distinguish M2 from M3 — so treat it as supporting, not decisive.
Evidence 4: the tool-choice surface matches MiniMax and nothing else — moderate
space-bunny-alpha supports only tool_choice: auto, rejecting none, required and function. Checking that four-way pattern across candidate families, the only exact match was MiniMax M2.7. GLM-5.3-Flash, Qwen3.8-Max-Prime, Kimi K3 and GPT-6-Luna all support none; MiniMax M3 supports all four. So the stealth model's tool plumbing looks like the MiniMax M2 line rather than M3 — consistent with a new config in the family, inconsistent with it being M3 under a codename.
What is the counter-evidence, and could this be wrong?
Yes, it could. Five things genuinely cut against the conclusion.
No lab has confirmed anything. Every fingerprint above is circumstantial. Parameter surfaces are configuration choices, and configurations can be copied, coincidental, or set by OpenRouter rather than by the lab.
The model appears to have changed at least twice during its run, and not only in one direction. Testers documented a capability jump around 29 September and a separate regression around 1 October, with users reporting worse writing, more guessing, and a model that stopped continuing long tasks unprompted — one noting it changed mid-session. No provider statement exists about any of it. If multiple checkpoints ran under one name, "what was space-bunny-alpha" may have no single answer, and any benchmark you read is pinned to an unstated date.
It is definitively not MiniMax M3 itself. M3 reports mandatory reasoning as false, a 512,000 output cap, a 1,048,576 context and four-way tool-choice support. All four differ. That is why the hypothesis has to be the gated M3.1 preview — a model whose surface mostly cannot be independently inspected.
The rival hypothesis — a GPT-based runtime — is stronger than expected and deserves a fair hearing. A dedicated research microsite, which labels its own verdict "Unconfirmed," reports probing the model into emitting Valid channels: analysis, commentary, final, summary alongside a # Juice marker — artefacts of OpenAI's Harmony response format and its internal reasoning-budget field, not things a MiniMax model should produce. Its system string, "You are an AI assistant accessed via API", was independently reported by an unrelated tester, so that artefact is genuinely corroborated even though the two groups draw opposite conclusions from it. The site now favours a router with a GPT-6.1-class core, while conceding that switching "has not been directly observed."
Two things stop us following it. Its own tokenizer comparison across 24 phrases scored MiniMax M3 at 24/24 exact matches and GPT-6.1 Sol at 7/24 — its strongest quantitative test points away from its own conclusion. And the modality field: zero of the 102 openai/* entries on OpenRouter accept video input. A stealth OpenAI model advertising native video ingestion would be a first.
One widely-cited piece of evidence we could not stand behind. It is claimed that Space Bunny and MiniMax M3 returned the same 32-character hex request-ID prefix, implying a shared serving stack. We could not reproduce it: OpenRouter normalises response IDs into its own gen-<timestamp>-<random> format, so the upstream provider's ID never reaches the response body. The argument may hold from raw provider headers, but it is not checkable from the public API.
Net position: roughly 85% that this is a MiniMax M3.1-generation model, and we may be wrong. That is a strong converging inference, not a reveal.
Why is asking the model who made it completely worthless?
Because people did, and it answered "OpenAI", "ChatGPT" and "Claude Opus 4.8" — same model, different sessions, mutually exclusive answers.
A model has no privileged introspective access to its own training run. Asked "who made you", it is not consulting a manifest; it is predicting a plausible continuation. Its training data is saturated with text in which assistants identify themselves as ChatGPT, because an enormous share of public assistant-dialogue online is ChatGPT transcripts — so any model trained on scraped or synthetic dialogue inherits that claim regardless of who trained it.
Here there is also a documented mechanism: testers found earlier turns were fed back as compressed developer messages, and those summaries dropped the "Space Bunny" identification within a few turns. The model loses its own codename mid-conversation. So "it told me it was ChatGPT" is evidence about the harness's context compression, not the weights. Self-identification is weak evidence about the training corpus and none at all about the lab.
Why do AI labs ship anonymous models at all?
This is not a conspiracy theory — it is a documented programme with published terms, and the Ox Alpha case is the cleanest proof. Z.ai wrote it down on their own blog:
"Before release, we tested GLM-5.3-Flash anonymously as ox-alpha on OpenCode and OpenRouter to gather user feedback. It quickly became the most popular model of the week — with all of this traffic served on Chinese AI chips."Note the stated motive: to gather user feedback. Put a known brand on a model and every rating is contaminated by expectation; an unbranded model gets judged on output alone. That is the one thing a lab cannot buy at any price.
OpenRouter's Stealth Program EULA spells the bargain out more plainly than any lab would. Providers "offer Models anonymously through our Service free of charge for a limited period of time," and the exchange is stated outright: "In consideration for the provision of your User Content to each Stealth Provider, access to the Stealth Models is provided to you free of charge." So the trade is free inference in exchange for your prompts. Retention terms vary per model, so read the specific listing — the umbrella EULA permits collecting content "for use in Stealth Model training and improvement," while space-bunny-alpha's page promised prompts "may be retained by the provider but are not used for training."
The anonymity is contractual rather than incidental: "upon request from each Stealth Provider, OpenRouter may not, in certain instances, disclose the name or origin of Stealth Providers to you." Nobody is going to slip up and tell you.
Beyond clean feedback, free public traffic buys adversarial load-testing that synthetic evals never produce, and it tells a lab where to price before launch. The cost is a block of GPU time — a cheap trade, which is why the practice keeps spreading, particularly across the open-weight and Chinese-lab ecosystem.
How can you fingerprint a stealth model yourself?
This is the transferable skill — the next unmarked model will arrive within weeks, and none of this requires special access.
Step 1 — pull the parameter surface. This one call reveals more than any amount of chatting with the model:
curl -s https://openrouter.ai/api/v1/models \
| python3 -c "
import json,sys
for m in json.load(sys.stdin)['data']:
if 'bunny' in m['id'].lower():
print(json.dumps(m, indent=2))
"Step 2 — hunt for a unique value. Take each numeric field — context length, max output — and count how many other models in the catalogue share it. A value held by exactly one model is a fingerprint; a value held by forty is noise. The 524,288 cap was the whole break in this case.
curl -s https://openrouter.ai/api/v1/models \
| python3 -c "
import json,sys,collections
d=json.load(sys.stdin)['data']
c=collections.Counter((m.get('top_provider') or {}).get('max_completion_tokens') for m in d)
for m in d:
v=(m.get('top_provider') or {}).get('max_completion_tokens')
if c[v]==1: print(c[v], v, m['id'])
"Step 3 — measure the fixed prompt overhead. Send hi to the stealth model and to every candidate family, then compare usage.prompt_tokens. As the table above shows, this separates serving stacks sharply, and it costs a handful of tokens.
Step 4 — check number conventions. Binary (1,048,576 / 131,072) versus decimal (1,000,000 / 128,000) is a habit that runs consistently through a lab's listings. Mixed conventions narrow the field fast.
Step 5 — compare capability flags, not prose. Tool-choice support, moderation status, modality list, tokenizer label, implicit caching. Plumbing decisions nobody bothers to disguise.
Step 6 — ignore what the model says about itself. It is the one input guaranteed to mislead you.
One caution from our own run: we also tried the token-delta method — measuring how many extra tokens a tricky multilingual probe adds, to isolate tokenizer efficiency from fixed overhead — and it did not discriminate. Deltas clustered between 53 and 59 tokens across MiniMax, DeepSeek, StepFun, Qwen and GLM alike; only Claude stood out, at 94. An attribution resting entirely on token deltas is on thinner ice than it looks.
What happens next, and where will confirmation show up?
There are two specific places, and both are checkable in under a minute.
1. OpenRouter back-fills the reveal into the stealth model's own page. This is the single highest-signal place to look, and it is how several previous cases resolved. The description text gets rewritten from "chosen to remain anonymous" into a named attribution. Verified examples live on OpenRouter right now:
| Codename | Back-filled description says |
|---|---|
union-alpha | "was a stealth model, revealed to be Pareto by Unbiased" |
bert-nebulon-alpha | "revealed on December 2nd as an early testing version of Mistral Large 3" |
polaris-alpha | "was an early snapshot of GPT-5.1 with reasoning effort set to minimal" |
andromeda-alpha | "has been revealed as NVIDIA Nemotron Nano 2 VL" |
space-bunny-alpha | "chosen to remain anonymous during this preview" — unchanged as of today |
So watch https://openrouter.ai/stealth/space-bunny-alpha for that sentence to change. The back-fill is not guaranteed, though: Ox Alpha's page still reads "chosen to remain anonymous" even though Z.ai confirmed it on their own blog, and codenames like horizon-alpha and cypher-alpha were never revealed at all.
2. An M3.1-Flash listing appearing on OpenRouter or Hugging Face. The model is already documented on MiniMax's own platform, but it exists nowhere else: no Hugging Face repo under MiniMaxAI, no entry in MiniMax's release notes (whose newest item is from July 2026), no mention on their news page, and no OpenRouter listing — the newest MiniMax entry there is still minimax-m3. If an M3.1 listing appears carrying a 524,288 output cap and mandatory reasoning defaulting to max, the case is effectively closed:
curl -s https://openrouter.ai/api/v1/models \
| python3 -c "
import json,sys
for m in json.load(sys.stdin)['data']:
if 'minimax' in m['id']:
print(m['id'], (m.get('top_provider') or {}).get('max_completion_tokens'))
"For the cost picture in advance: MiniMax M3 lists at $0.30 per million input tokens and $1.20 per million output. A Flash-tier sibling would normally land at or below that — which would make a million-token context with a half-million-token output ceiling genuinely cheap, assuming the attribution holds.
What should you take from this?
The decision rule is simple. Do not build on a stealth model. A listing with an expiry date, no lab behind it, no published pricing and a checkpoint that may change under you is something to evaluate, not a dependency. Run your prompts through it while it is free, write down what it was good at, and wait for the named release before wiring it into anything.
It is also worth knowing what a real confirmation looks like, because this is not one. With Ox Alpha, Z.ai published the admission on its own blog, Bloomberg wrote it up, and the story drew hundreds of points on Hacker News within days. Space Bunny Alpha has none of that: no lab statement, no press coverage, a handful of low-engagement forum posts, and an OpenRouter page that still says "anonymous." The absence of that signature is how you tell inference from fact.
What is worth keeping is the method. In August the same six steps correctly placed GLM-5.3-Flash behind the Ox Alpha codename before Z.ai confirmed it. This time they point at MiniMax's M3.1 preview, with five fingerprints converging, one serious dissenting theory, and no confirmation. Eighty-five percent, laid out so you can check the reasoning and disagree with it, beats a confident guess — and if a different lab claims it, the method is still what produced the shortlist.
FAQ
What is space-bunny-alpha?
space-bunny-alpha is an anonymous large language model served on OpenRouter under the provider name "Stealth" from 23 September 2026, with a listed expiry of 5 October 2026. It was free to use, had a 1,000,000-token context window and a 524,288-token maximum output, accepted text, image and video input, and used mandatory reasoning with five effort levels defaulting to max. No lab has publicly claimed it.
Who made space-bunny-alpha?
No lab has claimed it, and OpenRouter still says the provider "has chosen to remain anonymous." The evidence points to MiniMax, specifically MiniMax-M3.1-Flash-Preview, at roughly 85% confidence: MiniMax documents that model with a 1M context, mandatory thinking, an effort ladder of exactly low/medium/high/xhigh/max defaulting to max, and video input — and space-bunny-alpha is the only one of 466 OpenRouter models matching that signature. Not confirmed.
Is space-bunny-alpha still free?
It was free throughout its run — $0 input, $0 output — and was still serving at zero cost on 5 October 2026, the expiry date on its OpenRouter listing. Expect delisting imminently. When it resurfaces under a real name it will almost certainly be paid; MiniMax M3, for reference, lists at $0.30 per million input tokens and $1.20 per million output.
Is space-bunny-alpha MiniMax?
Probably, but unconfirmed — and it is definitely not the public MiniMax M3, which differs on four fields: reasoning is not mandatory, the output cap is 512,000, the context is 1,048,576, and it supports all four tool-choice modes where space-bunny-alpha supports only auto. The specific hypothesis is MiniMax-M3.1-Flash-Preview, documented on MiniMax's developer platform but gated to their M Plan and MiniMax Code tiers.
Why do AI labs release anonymous models?
To get evaluation feedback that brand expectations cannot contaminate — an unbranded model is judged on output alone. Z.ai documented doing exactly this: "before release, we tested GLM-5.3-Flash anonymously as ox-alpha on OpenCode and OpenRouter to gather user feedback." OpenRouter's Stealth Program terms state the exchange plainly — access is free "in consideration for the provision of your User Content to each Stealth Provider." The lab also gets adversarial load-testing at scale.
How can you tell which lab made a stealth model?
Compare parameter surfaces, not prose. Pull https://openrouter.ai/api/v1/models, then: find numeric values the model shares with exactly one other family; measure fixed prompt overhead by sending a two-character message and reading usage.prompt_tokens; check whether sizes use binary or decimal conventions; and compare capability flags like tool-choice support and tokenizer label. Never ask the model who made it — space-bunny-alpha variously claimed to be OpenAI, ChatGPT and Claude Opus 4.8.
Did space-bunny-alpha change during its run?
Apparently more than once. Testers documented a capability improvement around 29 September 2026 and a separate regression around 1 October, with users reporting degraded writing and a model that stopped continuing long tasks unprompted. None of it is confirmed by OpenRouter or any lab, and all of it is user perception rather than measurement — but if several checkpoints ran under one name, any benchmark for "space-bunny-alpha" needs a date attached.
Was space-bunny-alpha any good?
Its measurable traits were unusual: most 1M-context models cap output at 65,536 or 131,072 tokens, while space-bunny-alpha allowed 524,288 — beaten at that context size only by Grok 4.3 — and it held 99.87% uptime serving free traffic. Beyond that, be careful: reasoning was mandatory and could not be disabled, inflating latency on simple requests, and the reported mid-run changes mean published impressions may not describe one consistent model.