Umans Flash Fastest
also served as umans-qwen3.6-35b-a3b
523.7tok/s
throughput · p50 · last 5 min
318ms
TTFT · p50 · last 5 min
99.83%
uptime · 24h
Our fastest model: a light workflow complement, not a standalone coder. Think Haiku next to Opus: not everything needs a frontier model, and Flash's speed (200+ tokens per second) compounds on the roles around umans-coder: gathering context, scout subagents, research, summaries, documentation, and quick edits.
Context
262K
Max output
262K
Recommended
33K
Vision
Yes
Tools
Yes
Reasoning
Toggle · none/low/medium/high
Weights
Trends
Speed over the last 90 days
now 519.5 tok/s
90 days agopre-release before May 3, 2026today
best 342ms · Aug 3now 516ms
90 days agopre-release before May 3, 2026today
Changelog
Events for Umans Flash
Jul 262026
Resolved: umans-flash (Qwen3.6-35B-A3B-FP8) Resolved
The incident is resolved. Throughout the outage, some requests continued to be served, and a subset of users were affected.
Only umans-flash (Qwen3.6-35B-A3B-FP8) model was affected.
Incident duration 25min.
Jul 262026
Major outage: umans-flash (Qwen3.6-35B-A3B-FP8) Incident
We are currently experiencing a major outage affecting umans-flash (Qwen3.6-35B-A3B-FP8). Other models are unaffected. Our team is fully mobilized and working to restore normal operation as quickly as possible. We will post updates here as the situation evolves.
Jul 62026
Resolved: back to full capacity Resolved
Hardware capacity was restored and both umans-glm-5.2 and umans-kimi-k2.7 are back to normal speed. During the outage the service ran in a reduced-capacity mode that favoured continuity over speed: requests kept flowing, at the cost of an uneven experience. The affected window is shaded on each model's speed trends.
Jul 22026
Incident: hardware outage, running at reduced capacity Incident
A hardware failure took part of our GPU fleet offline. We failed over to reduced capacity to keep the service available: umans-glm-5.2 and umans-kimi-k2.7 stayed up, but slower than usual and with an uneven experience under load. Live updates were posted on Discord throughout.