umans/status
Live · updated just now

All systems operational

Live status and speed of the Umans Code gateway and its models, refreshed every 30 seconds. For each model we show its output speed and median time to first token.

5 / 5production models operational
98.89%gateway uptime · 90 days
1model in testing
0active issues
Gateway

API endpoint

the front door every model is served through
API gateway Operational
api.code.umans.ai
100.00%
service uptime · 24h
1.16s
median TTFT · all models
90 days agotoday
Production

Models in production

5 listed

tok/s = output tokens per second · TTFT = time to first token · p50 = median over the last 5 minutes. On each gauge the midpoint is that model's target; the marker sits further right when it's beating target (faster TTFT, higher throughput).

umans-glm-5.2 · GLM
Operational
81.3tok/s
throughput · p50 · last 5 min
708ms
TTFT · p50 · last 5 min
100.00%
uptime · 24h

GLM 5.2 is our best model for coding right now, with a 400K context window for large codebases. Vision is available on the Anthropic Messages API (`/v1/messages`) only, through a server-side handoff (GLM 5.2 generates the text, Kimi preprocesses the image); that handoff will be retired soon in favour of more efficient client-side image handling.

90-day speed trends & events →
90 days agoin production since Jun 21, 2026today
Context
406K
Max output
131K
Recommended
131K
Vision
Via handoff
Tools
Yes
Reasoning
Toggle · none/high/max
Weights
umans-kimi-k2.7 · Kimi K2.7-Code · Moonshot
successor to Kimi K2.6
also served as umans-coder
Operational
76.6tok/s
throughput · p50 · last 5 min
6.35s
TTFT · p50 · last 5 min
100.00%
uptime · 24h

Kimi K2.7-Code via Umans Code - Moonshot's strongest coding model and the successor to Kimi K2.6. Built for complex, tool-heavy agentic coding; it reasons more efficiently than K2.6, so agent sessions run faster at the same depth. Deprecated: use `umans-kimi-k3` instead (sunset 2026-08-10).

90-day speed trends & events →
90 days agoin production since Jun 12, 2026today
Context
262K
Max output
262K
Recommended
33K
Vision
Yes
Tools
Yes
Reasoning
Always on
Umans Flash Fastest
umans-flash · Qwen3.6-35B-A3B · Qwen
also served as umans-qwen3.6-35b-a3b
Operational
523.7tok/s
throughput · p50 · last 5 min
319ms
TTFT · p50 · last 5 min
100.00%
uptime · 24h

Our fastest model: a light workflow complement, not a standalone coder. Think Haiku next to Opus: not everything needs a frontier model, and Flash's speed (200+ tokens per second) compounds on the roles around umans-coder: gathering context, scout subagents, research, summaries, documentation, and quick edits.

90-day speed trends & events →
90 days agoin production since May 3, 2026today
Context
262K
Max output
262K
Recommended
33K
Vision
Yes
Tools
Yes
Reasoning
Toggle · none/low/medium/high
Weights
umans-kimi-k3 · Kimi K3 · Moonshot
Degraded performance
47.2tok/s
throughput · p50 · last 5 min
39.09s
TTFT · p50 · last 5 min
100.00%
uptime · 24h

Kimi K3: Moonshot's most capable model and the first open 3T-class release - 2.8T parameters, a 1M-token context window, and native vision, built for repository-scale understanding and long agentic runs at closed-frontier quality (92.4% vs Claude Fable 5's 92.6% across ~1,030 agentic tasks in Fireworks' independent study). It thinks by default at maximum reasoning effort; select none, low, high, or max to trade depth for speed.

90-day speed trends & events →
90 days agoin production since Jul 31, 2026today
Context
1049K
Max output
131K
Recommended
131K
Vision
Yes
Tools
Yes
Reasoning
Toggle · none/low/high/max
umans-deepseek-v4-flash-0731 · DeepSeek-V4-Flash · DeepSeek
Operational
329.8tok/s
throughput · p50 · last 5 min
1.31s
TTFT · p50 · last 5 min
100.00%
uptime · 24h

DeepSeek V4 Flash: DeepSeek's fast agentic coding MoE (284B total, 13B active), served from the official 0731 release on a 1M-token context. The cheapest production model in the lineup for real agentic work. Reasoning has four modes: non-think (none), think low (low, the default), think high (high) and think max (max). Served on our own GPU infrastructure with high availability.

90-day speed trends & events →
90 days agoin production since Aug 3, 2026today
Context
1049K
Max output
393K
Recommended
393K
Vision
No
Tools
Yes
Reasoning
Toggle · none/low/high/max
Testing

In the playground

1 · try before it's promoted
What "playground" meansPlayground models are short, experimental test runs: not permanent, offered at low capacity, and expected to be flaky under load. They're here so you can push them and tell us what you find. Uptime and speed are measured the same way as production, but don't build on them. For real work, use the production model each one is based on.
umans-deepseek-v4-flash-0731-lab · DeepSeek-V4-Flash · DeepSeek
In testing
379.4tok/s
throughput · p50 · last 5 min
1.14s
TTFT · p50 · last 5 min
100.00%
in testing

DeepSeek V4 Flash as a Labs experiment, open for a short test window: temporary, not a permanent id. DeepSeek's fast agentic coding MoE (284B total, 13B active), served from the official 0731 release, on a 1M-token context. Reasoning has four modes: non-think (none), think low (low, the default), think high (high) and think max (max). Access is seat-gated through the Labs page while an experiment is live. It is offered at limited capacity and availability, so expect it to be flaky and to go down under load: crash it, give it a moment, and try again. When the window ends, the model keeps serving as the pay-per-token umans-deepseek-v4-flash-0731.

90-day speed trends & events →
Stage
Playground
Context
1049K
Max output
393K
Recommended
393K
Vision
No
Tools
Yes
Reasoning
Toggle · none/low/high/max
History

Gateway uptime

90-day window · daily worst status
90 days ago98.89% operationaltoday
Changelog

Recent events & model lifecycle

releases, retirements, and playground changes
Aug 32026
Released pay-per-token: Umans DeepSeek V4 Flash Released
umans-deepseek-v4-flash-0731 joins the lineup as the cheapest way we serve real agentic work: $0.14 / $0.28 / $0.028 per 1M (input / output / cache read), a 1M context window, thinking at low effort by default (dial up high or max when a task deserves more). It is the new default for new chats and CLI setups. Founding users pay the 10x cheaper cache rate until Monday, August 10, 2026 (see /pricing). Served on our own GPU infrastructure with high availability.
Aug 32026
The V4 Flash lab continues as umans-deepseek-v4-flash-0731-lab Testing
The V4 Flash pilot closed at the pay-per-token release: the production id now bills per token, so the seat-gated pilot on it ended rather than charge anyone by surprise. The lab reopens on the new umans-deepseek-v4-flash-0731-lab id with a smaller cohort - free, seat-gated, same experimental capacity as before. The model keeps serving as umans-deepseek-v4-flash-0731 regardless: the model stays.
Aug 12026
Generally available: Umans Kimi K3 Released
The Kimi K3 Labs pilot closed and umans-kimi-k3 is generally available: no Labs seat required anymore - plan and pay-per-token keys alike get the 1M context window, native vision, and max-effort reasoning at $3.00 / $15.00 / $0.30 per 1M (input / output / cache read).
Jul 312026
Released pay-per-token: Umans Kimi K3 Released
umans-kimi-k3 opened to pay-per-token (wallet and service-account keys): Moonshot's largest open model, a 1M context window, native vision, and max reasoning effort by default, billed at $3.00 / $15.00 / $0.30 per 1M (input / output / cache read).
Jul 262026
Resolved: umans-flash (Qwen3.6-35B-A3B-FP8) Resolved
The incident is resolved. Throughout the outage, some requests continued to be served, and a subset of users were affected. Only umans-flash (Qwen3.6-35B-A3B-FP8) model was affected. Incident duration 25min.
Jul 262026
Major outage: umans-flash (Qwen3.6-35B-A3B-FP8) Incident
We are currently experiencing a major outage affecting umans-flash (Qwen3.6-35B-A3B-FP8). Other models are unaffected. Our team is fully mobilized and working to restore normal operation as quickly as possible. We will post updates here as the situation evolves.
Jul 202026
Planned maintenance: Umans GLM 5.2 Maintenance
We're making a change to GLM 5.2 on 2026-07-20, from 11:15 to 12:00 CET. During the window, around 30 to 40 minutes, GLM 5.2 responses may noticeably be slower than usual. Requests still go through, and other models are unaffected. The work is aimed at making GLM 5.2 start responding faster once it lands.
Jul 162026
Playground closed: Umans DeepSeek V4 Pro DSpark Testing
The DSpark test window closed after two days. What we took from it: DSpark speculative decoding lets us serve more users at once while keeping each session fast enough, the DeepSeek V4 architecture is now mature enough to serve at scale, and the model itself is solid. V4 Pro is not joining the lineup though: it is still a preview build, and the issues testers hit (DSML leaks, language bleed, long-context artifacts) are model-side. DeepSeek confirmed a better version is coming this month, so we would rather roll the learnings into that. Thanks to everyone who tested.
Jul 142026
Playground opened: Umans DeepSeek V4 Pro DSpark Testing
umans-deepseek-v4-pro-dspark entered the playground for a short, seat-gated test window: DeepSeek V4 Pro served from the original weights with DSpark speculative decoding. Experimental and temporary; not for production.
Jul 102026
Resolved: Umans GLM 5.2 Resolved
The incident is resolved. Throughout the outage, most requests continued to be served, and a subset of users were affected. Umans GLM 5.2 is still running at reduced capacity, so responses may be slower than usual at peak times. We will post an update once full capacity is restored.
Jul 102026
Major outage: Umans GLM 5.2 Incident
We are currently experiencing a major outage affecting Umans GLM 5.2. Other models are unaffected. Our team is fully mobilized and working to restore normal operation as quickly as possible. We will post updates here as the situation evolves.
Jul 92026
GLM 5.2 performance stabilized, still on reduced capacity Resolved
GLM 5.2 performance has stabilized after deploying the prefill kernel update and selective prioritization for interactive requests. We’re still operating today in reduced-capacity mode while capacity is being restored, but the service is holding and requests are flowing normally. We’ll keep monitoring closely and update here if anything changes.
Jul 92026
Degraded performance on Umans GLM 5.2 Incident
We’re currently seeing degraded performance on Umans GLM 5.2, mainly higher time-to-first-token and uneven streaming speed. Requests are still going through, but the experience can feel slower than usual. We’re investigating and applying mitigations. We’ll update here once things are back to target.
Jul 62026
Resolved: back to full capacity Resolved
Hardware capacity was restored and both umans-glm-5.2 and umans-kimi-k2.7 are back to normal speed. During the outage the service ran in a reduced-capacity mode that favoured continuity over speed: requests kept flowing, at the cost of an uneven experience. The affected window is shaded on each model's speed trends.
Jul 22026
Incident: hardware outage, running at reduced capacity Incident
A hardware failure took part of our GPU fleet offline. We failed over to reduced capacity to keep the service available: umans-glm-5.2 and umans-kimi-k2.7 stayed up, but slower than usual and with an uneven experience under load. Live updates were posted on Discord throughout.
Jul 22026
Playground closed: Umans GLM 5.2 NVFP4 Testing
The short NVFP4 test window ended after four days. Thanks to everyone who pushed it and shared findings.
Jun 292026
Playground opened: Umans GLM 5.2 NVFP4 Testing
umans-glm-5.2-nvfp4 entered the playground for a short, low-capacity test window. Experimental and temporary; not for production.
Jun 242026
Retired: Umans GLM 5.1 Retired
umans-glm-5.1 was retired in favour of GLM 5.2. Requests to the old id now return a clear deprecation error pointing to umans-glm-5.2.
Jun 212026
Released to production: Umans GLM 5.2 Released
umans-glm-5.2 was released as the long-context model, with a 405K context window, after a pre-release period that started Jun 16.
Jun 182026
Retired: Umans Kimi K2.6 Code Retired
umans-kimi-k2.6 was retired and superseded by K2.7. Requests to the old id now return a clear deprecation error pointing to umans-kimi-k2.7.
Jun 122026
Released to production: Umans Kimi K2.7 Code Released
umans-kimi-k2.7 was released as the recommended coding model (also served as umans-coder).
May 132026
Retired: Umans Kimi K2.5 Retired
umans-kimi-k2.5 was retired and superseded by Kimi K2.6.
May 92026
Retired: Umans MiniMax M2.5 Retired
umans-minimax-m2.5 was retired without a direct replacement.
Past models 6 retired
umans-deepseek-v4-pro-dspark · DeepSeek
retired Jul 16, 2026
umans-glm-5.2-nvfp4 · GLM
retired Jul 2, 2026
umans-glm-5.2
umans-glm-5.1 · GLM
retired Jun 24, 2026
umans-glm-5.2
umans-kimi-k2.6 · Moonshot
retired Jun 18, 2026
umans-kimi-k2.7
umans-kimi-k2.5 · Moonshot
retired May 13, 2026
umans-kimi-k2.6
umans-minimax-m2.5 · MiniMax
retired May 9, 2026