Umans DeepSeek V4 Flash Recommended
327.0tok/s
throughput · p50 · last 5 min
796ms
TTFT · p50 · last 5 min
100.00%
uptime · 24h
DeepSeek V4 Flash: DeepSeek's fast agentic coding MoE (284B total, 13B active), served from the official 0731 release on a 1M-token context. The cheapest production model in the lineup for real agentic work. Reasoning has four modes: non-think (none), think low (low, the default), think high (high) and think max (max). Served on our own GPU infrastructure with high availability.
Context
1049K
Max output
393K
Recommended
393K
Vision
No
Tools
Yes
Reasoning
Toggle · none/low/high/max
Weights
Trends
Speed over the last 90 days
90 days agopre-release before Aug 3, 2026today
90 days agopre-release before Aug 3, 2026today
Changelog
Events for Umans DeepSeek V4 Flash
Aug 252026
Completed: maintenance on umans-deepseek-v4-flash-0731
The scheduled maintenance work on umans-deepseek-v4-flash-0731 is complete. The service remained available throughout, with no noticeable impact. No further action is needed.
Aug 242026
Scheduled maintenance on umans-deepseek-v4-flash-0731
We're performing maintenance work on umans-deepseek-v4-flash-0731 over the next couple of hours. The service will remain available throughout, with no interruption expected. We'll post an update here once the work is complete.
Aug 222026
Completed: maintenance on umans-deepseek-v4-flash-0731
The scheduled maintenance work on umans-deepseek-v4-flash-0731 is complete. The service remained available throughout, with no noticeable impact. No further action is needed.
Aug 222026
Scheduled maintenance on umans-deepseek-v4-flash-0731
We're performing maintenance work on umans-deepseek-v4-flash-0731 over the next couple of hours. The service will remain available throughout, with no interruption expected. We'll post an update here once the work is complete.
Older events 1
Aug 32026
Released pay-per-token: Umans DeepSeek V4 Flash Released
umans-deepseek-v4-flash-0731 joins the lineup as the cheapest way we serve real agentic work: $0.14 / $0.28 / $0.028 per 1M (input / output / cache read), a 1M context window, thinking at low effort by default (dial up high or max when a task deserves more). It is the new default for new chats and CLI setups. Founding users pay the 10x cheaper cache rate until Monday, August 10, 2026 (see /pricing). Served on our own GPU infrastructure with high availability.