Umans DeepSeek V4 Flash
300.9tok/s
throughput · p50 · last 5 min
2.08s
TTFT · p50 · last 5 min
100.00%
uptime · 24h
DeepSeek V4 Flash: DeepSeek's fast agentic coding MoE (284B total, 13B active), served from the official 0731 release on a 1M-token context. The cheapest production model in the lineup for real agentic work. Reasoning has four modes: non-think (none), think low (low, the default), think high (high) and think max (max). Served on our own GPU infrastructure with high availability.
Context
1049K
Max output
393K
Recommended
393K
Vision
No
Tools
Yes
Reasoning
Toggle · none/low/high/max
Weights
Trends
Speed over the last 90 days
peak 317.7 tok/s · Aug 1now 317.7 tok/s
90 days agopre-release before Aug 3, 2026today
best 1.50s · Aug 1now 1.70s
90 days agopre-release before Aug 3, 2026today
Changelog
Events for Umans DeepSeek V4 Flash
Aug 32026
Released pay-per-token: Umans DeepSeek V4 Flash Released
umans-deepseek-v4-flash-0731 joins the lineup as the cheapest way we serve real agentic work: $0.14 / $0.28 / $0.028 per 1M (input / output / cache read), a 1M context window, thinking at low effort by default (dial up high or max when a task deserves more). It is the new default for new chats and CLI setups. Founding users pay the 10x cheaper cache rate until Monday, August 10, 2026 (see /pricing). Served on our own GPU infrastructure with high availability.
Jul 62026
Resolved: back to full capacity Resolved
Hardware capacity was restored and both umans-glm-5.2 and umans-kimi-k2.7 are back to normal speed. During the outage the service ran in a reduced-capacity mode that favoured continuity over speed: requests kept flowing, at the cost of an uneven experience. The affected window is shaded on each model's speed trends.
Jul 22026
Incident: hardware outage, running at reduced capacity Incident
A hardware failure took part of our GPU fleet offline. We failed over to reduced capacity to keep the service available: umans-glm-5.2 and umans-kimi-k2.7 stayed up, but slower than usual and with an uneven experience under load. Live updates were posted on Discord throughout.