News

DeepSeek’s latest drop forces another round of price cuts

An open-weight model that is good enough is now a pricing event, not a research paper. The closed labs are answering the only way they still know how.

Candlestick chart on a dark trading screen. Photo by Nicholas Cappello on Unsplash.
Photo by Nicholas Cappello on Unsplash

DeepSeek released a new open-weight model this week that is not, on the public benches, a crown jewel. It is something more inconvenient: good enough, cheap to run, and already in the hands of people who do not send a purchase order to San Francisco.

Within thirty-six hours, two U.S. labs trimmed API list prices on mid-tier reasoning SKUs. A third “clarified” that a previously quiet discount tier was now generally available. Nobody called it a response. Everyone treated it as one.

The product is no longer the model. The product is the willingness to keep cutting until the open weights look expensive.

What actually shipped

The weights are public. The training recipe is not a full confession, but it is more than a teaser. Independent evals put the model in the same band as last quarter’s closed systems on code and multilingual short-form, and a step behind on long-horizon tool use. That gap is real. It is also the gap a fine-tune and a decent harness keep closing.

More important than the leaderboard is the serving story. Early vLLM and SGLang numbers circulating among infra teams show tokens per dollar that make last year’s “we will lose money on inference to win the platform” memos look quaint. If you already own GPUs, the argument for paying a closed API for generic work just got thinner again.

The price cut is the press release

Frontier labs have a limited set of public moves. They can ship a better model, which takes time they may not have. They can restrict the open-weight story with licensing and export theater, which invites a political fight they keep losing in the developer community. Or they can cut prices and call it generosity.

Price cuts are not generosity. They are an admission that the moat was never the architecture diagram. It was distribution, latency, and the habit of paying. Habits break when the substitute is free to download and merely annoying to operate.

What to watch

Three things, none of them the next demo video:

  1. Whether enterprise procurement treats “open weight plus vendor support” as a real line item, or keeps buying closed APIs out of fear of the audit.
  2. Whether the labs answer with features that do not clone — data residency, evals, and tool-use reliability — or with another 20 percent off.
  3. Whether regulators notice that “open” is doing a lot of work for models trained on piles nobody can name.

The news is not that a Chinese lab can train a strong model. That news is old. The news is that the American price list is now a lagging indicator of someone else’s release notes.