umans/status/umans-flash
Live · refreshes every 30s
← all models
Umans Flash Fastest
umans-flash · Qwen3.6-35B-A3B · Qwen
also served as umans-qwen3.6-35b-a3b
Operational
523.7tok/s
throughput · p50 · last 5 min
318ms
TTFT · p50 · last 5 min
99.83%
uptime · 24h

Our fastest model: a light workflow complement, not a standalone coder. Think Haiku next to Opus: not everything needs a frontier model, and Flash's speed (200+ tokens per second) compounds on the roles around umans-coder: gathering context, scout subagents, research, summaries, documentation, and quick edits.

90 days agoin production since May 3, 2026today
Context
262K
Max output
262K
Recommended
33K
Vision
Yes
Tools
Yes
Reasoning
Toggle · none/low/medium/high
Weights
Trends

Speed over the last 90 days

daily medians · dashed line = target
throughput p50 · output tokens per second, higher is better
now 519.5 tok/s
90 days agopre-release before May 3, 2026today
TTFT p50 · time to first token, lower is better
best 342ms · Aug 3now 516ms
90 days agopre-release before May 3, 2026today
Changelog

Events for Umans Flash

incl. gateway-wide announcements
Jul 262026
Resolved: umans-flash (Qwen3.6-35B-A3B-FP8) Resolved
The incident is resolved. Throughout the outage, some requests continued to be served, and a subset of users were affected. Only umans-flash (Qwen3.6-35B-A3B-FP8) model was affected. Incident duration 25min.
Jul 262026
Major outage: umans-flash (Qwen3.6-35B-A3B-FP8) Incident
We are currently experiencing a major outage affecting umans-flash (Qwen3.6-35B-A3B-FP8). Other models are unaffected. Our team is fully mobilized and working to restore normal operation as quickly as possible. We will post updates here as the situation evolves.
Jul 62026
Resolved: back to full capacity Resolved
Hardware capacity was restored and both umans-glm-5.2 and umans-kimi-k2.7 are back to normal speed. During the outage the service ran in a reduced-capacity mode that favoured continuity over speed: requests kept flowing, at the cost of an uneven experience. The affected window is shaded on each model's speed trends.
Jul 22026
Incident: hardware outage, running at reduced capacity Incident
A hardware failure took part of our GPU fleet offline. We failed over to reduced capacity to keep the service available: umans-glm-5.2 and umans-kimi-k2.7 stayed up, but slower than usual and with an uneven experience under load. Live updates were posted on Discord throughout.