LIVE MODEL INTELLIGENCE · UPDATED AS THINGS HAPPEN

Ox Alpha Tracker

Ox Alpha is an anonymous model on OpenRouter built around coding, long-running agents, and a 1,048,576-token context window. This site helps developers decide whether to try it, what the public evidence actually says, and what remains unknown.

Current public status · checked August 23, 2026
Identity
Not officially confirmed
OpenRouter price
$0 input · $0 output
Context
1,048,576 tokens
Benchmark posture
Community tests only

LATEST UPDATES

What changed

  1. Bloomberg coverage brings the mystery to a wider audience

    Bloomberg reported on the model's fast developer adoption and unknown creator. A syndicated copy inThe Straits Times repeats the core public facts: a stealth OpenRouter listing, roughly one million tokens of context, and text, image, and video input. Media attention does not identify the developer.

  2. A reproducible LiveCodeBench run is published

    An independent community repository reports 49 correct answers from 175 LiveCodeBench v6 problems, or 28.0% pass@1, using one attempt and temperature zero. Therepository and methodology make this more inspectable than a screenshot alone, but it is still a community run rather than an official model evaluation.

  3. A small DeepSWE subset starts the benchmark debate

    A public10-task community test on X reports 80% for Ox Alpha, 65% for Fable, and 52% for gpt-5.6-sol. The author explicitly warns that ten tasks can have high variance. We track the result as an early signal, not proof that Ox Alpha wins a complete benchmark.

FOUR DECISIONS

Find the answer you need

If the comparison is the main question, start withOx Alpha vs Claude Fable 5 orOx Alpha vs GPT-5. Both pages keep vendor specifications separate from Ox Alpha's early community results.

THE SHORT VERSION

Known, reported, and still unknown

Confirmed: OpenRouter lists `stealth/ox-alpha` as a reasoning model for coding and sustained agentic work. Its public metadata exposes a 1,048,576-token context window, up to 131,072 output tokens, text/image/video input, tool calling, and a current price of zero for prompt and completion tokens. TheStealth provider page says OpenRouter routes requests but is not the developer, owner, or provider. It also says the anonymous provider retains prompts and completions without using them for training.

Reported: Developers are publishing mixed results. Some describe strong repository reading and long-running work; others report slowness, capacity errors, or repeated coding mistakes. The two numerical tests we can currently inspect use very different setups. The honest conclusion is that the model is interesting enough to test, but the evidence is not mature enough to rank it cleanly.

Unknown: the creator, training details, parameter count, safety evaluation, service commitment, future price, permanent availability, and broad commercial terms. Community fingerprinting points toward the GLM family while other posts speculate about Google DeepMind. Neither theory is official. Theidentity tracker records clues without turning them into a reveal.

RECOMMENDATION

Treat it as a preview, not a dependency

Ox Alpha is easiest to justify as a controlled experiment: choose a non-sensitive repository task, define the success criteria before the run, record latency and retries, inspect every change, and compare it with the model you already use. Free tokens can make broad testing affordable, but they do not remove review time, reliability risk, or data-governance questions. Do not send credentials, client code, or confidential material to an unnamed provider simply because the API price is zero.

The next practical step is theOpenRouter setup guide. Before a larger test, check the price and availability recordbecause a stealth preview can change or disappear without the release process developers expect from a named production model.

Keep one named model as a baseline and preserve failed Ox Alpha runs. A failure with its prompt, model route, tool log, and test output is actionable evidence; a hand-picked success screenshot is not. The goal is to learn which tasks the preview completes reliably, what review it demands, and whether those gains survive a repeated run.

Last updated: 2026-08-23