Umans DeepSeek V4 Pro Retired
no longer available; use umans-deepseek-v4.1-flash
retired Sep 14, 2026 · replaced by umans-deepseek-v4.1-flash · weights ↗
115.5tok/s
throughput · p50 · whole period
1.28s
TTFT · p50 · whole period
100.00%
uptime · whole period
DeepSeek V4 Pro, served from the official 0813 release: DeepSeek's flagship coding and reasoning MoE, built for long-horizon agentic coding and demanding tool-heavy workloads on a 1M-token context window. The 0813 release supersedes the April preview with substantially stronger agentic performance. Reasoning has three modes: non-think (none), think high (high, the default) and think max (max). Billed per token ($1.32 / $3.96 / $0.044 per 1M; input / output / cache read). Served on our own GPU infrastructure with high availability. Deprecated: use `umans-deepseek-v4.1-flash` instead (sunset 2026-09-14).
Trends
Speed over its final 90 days
Jun 18, 2026retired Sep 14, 2026Sep 16, 2026
Jun 18, 2026retired Sep 14, 2026Sep 16, 2026
Changelog
Events for Umans DeepSeek V4 Pro
Sep 142026
Retired: Umans DeepSeek V4 Pro Retired
umans-deepseek-v4-pro-0813 reached its sunset and was retired in favour of DeepSeek V4.1 Flash: it leaves the catalog and the model pickers, and its history stays on the past-models list. Requests still pinned to the old id keep routing for a short grace tail - migrate to umans-deepseek-v4.1-flash (the DeepSeek lineage's successor: 1M context, native vision) now; a final cutoff date will be announced ahead of the hard stop.
Older events 1
Aug 152026
Released pay-per-token: Umans DeepSeek V4 Pro Released
umans-deepseek-v4-pro-0813 joins the lineup as the long-context coding flagship: DeepSeek's official 0813 release of V4 Pro, the checkpoint that served here as the seat-gated pre-release lab since August 13, with a 1M context window and thinking at high effort by default (dial to max when a task deserves more). Billed per token: $1.32 / $3.96 / $0.044 per 1M (input / output / cache read). It succeeds GLM 5.2, which is deprecated and sunsets on August 23, 2026. The pre-release window's metrics stay on the model's status page as its pre-release period (before Aug 15). Served on our own GPU infrastructure with high availability.