XT.PT The API → This story
News The API

DeepSeek backs off rerouting deepseek-v4-pro to V4.1-Flash, and its own announcement still says otherwise

Two DeepSeek documents disagree about which model answers the deepseek-v4-pro ID, and the legacy V4-Flash IDs are already served by a different model.

Filed15 Sep 2026, 14:00 UTC Length3 min · 594 words ReportingPrelo
DeepSeek

DeepSeek released DeepSeek-V4.1-Flash on September 10 and, in the same announcement, told API users that its flagship model ID would soon stop meaning its flagship model. Within days the plan had changed. As of September 15, DeepSeek's own documentation gives two different answers to the question every integrator cares about: what runs when you send "model": "deepseek-v4-pro"?

What the launch said

The English announcement page, dated 2026/09/10, is blunt about V4-Pro's future:

Tests by multiple parties put V4.1-Flash ahead of V4-Pro on performance, cost, speed & total runtime. We're phasing out V4-Pro.

DeepSeek, V4.1-Flash announcement

The next line set the date: "Starting at 04:00 UTC on Sept 14, 2026, all deepseek-v4-pro requests will route to V4.1-Flash at V4.1-Flash rates. This will continue until V4.1-Pro launches." The Chinese version gives the same cutover as 12:00 Beijing time on September 14, which is the same instant. That was four days' notice for a silent model swap behind an unchanged ID, 32 days after DeepSeek's changelog announced V4-Pro's general availability on August 13.

What the changelog says now

The September 10 changelog entry, in English and Chinese, now carries a different paragraph: "In response to user demand, we have decided to continue providing API services for DeepSeek V4 Pro after September 14, 2026, with the billing method remaining unchanged. We will provide further notice should there be any changes."

The Models & Pricing page repeats that sentence as a footnote and lists deepseek-v4-pro as its own column, with model version DeepSeek-V4-Pro-0813, its own prices and a concurrency limit of 500 (against 2,500 for deepseek-flash). The announcement page, English and Chinese, still describes the reroute. Neither document is dated for the change, so XT.PT cannot say when the reversal was posted, only that both versions were live when we read them on September 15.

The IDs that did change

The reversal covers V4-Pro only. The Flash IDs were swapped on September 10 as announced. The changelog says V4 Flash and V4 Flash Vision Exp "have been retired," and the pricing page is specific about what that means for callers:

deepseek-v4-flash              -> served by DeepSeek-V4.1-Flash, billed at Flash price
deepseek-v4-flash-vision-exp   -> served by DeepSeek-V4.1-Flash, billed at Flash price
deepseek-flash                 -> DeepSeek-V4.1-Flash (the new canonical ID)

The docs call that routing temporary. V4-Flash-Vision-Exp was announced in the changelog on August 21, so it lasted 20 days as a distinct model.

The price gap explains why the V4-Pro decision matters to someone's invoice. Per million tokens at peak rates, the pricing page lists V4-Pro at $1.32 for cache-miss input and $3.96 for output, against $0.30 and $1.20 for deepseek-flash. Off-peak rates are half; peak hours are 01:00-04:00 and 06:00-10:00 UTC, Monday through Friday. Had the reroute gone ahead, V4-Pro callers would have paid less for a different model, one DeepSeek describes as a 552B-parameter MoE with "just 8B active parameters for input, 16B for output."

What to do with a model ID

If you pin deepseek-v4-pro for evaluation reproducibility, the current pricing page says you are still getting V4-Pro-0813, but DeepSeek has now shown it is willing to announce a same-ID model swap on four days' notice, and its promise is only to "provide further notice." If you pin deepseek-v4-flash, you are already on V4.1-Flash and should move to deepseek-flash so your config says what it actually runs.

Primary sources: DeepSeek API Change Log, DeepSeek-V4.1-Flash announcement, DeepSeek Models & Pricing, Chinese changelog, Chinese announcement, read 2026-09-15.

Corrections and source documents: contact the desk
Read next →
Read next
Pricing · 4 min

GPT-6 Sol costs what GPT-5.6 Terra did, and exactly what Claude Sonnet 5.5 does

The API · 4 min

Claude Sonnet 5.5 makes thinking: disabled a 400, one of five breaking changes