Founders NorthFounders North
Back to Home

Ollama Cuts DeepSeek-V4 Cloud Prices by 50% During Off-Peak Hours

Ollama has introduced a 50 percent token price reduction for DeepSeek-V4-Flash and DeepSeek-V4-Pro during off-peak hours across its US and Europe cloud infrastructure.

Monday, September 7, 2026

Key Takeaways

  • Ollama has slashed token prices by 50% for DeepSeek-V4-Flash and DeepSeek-V4-Pro on its US and Europe cloud infrastructure.
  • Off-peak hours include weekdays before 12:00 UTC or after 18:00 UTC, as well as all day on weekends.
  • Peak pricing remains strictly active between 12:00 and 18:00 UTC from Monday to Friday.
  • Founders and builders can leverage this pricing structure to significantly lower inference costs for asynchronous or batch workloads.

The artificial intelligence infrastructure landscape is shifting toward aggressive cost optimization for developers. Ollama announced a significant 50 percent reduction in token pricing for the DeepSeek-V4-Flash and DeepSeek-V4-Pro models. According to the company's recent newsletter update, these lower rates apply across Ollama's cloud infrastructure in the United States and Europe during designated off-peak hours.

Under the new pricing structure, off-peak hours are defined as before 12:00 UTC or after 18:00 UTC on weekdays, alongside full-day coverage on weekends. Peak pricing continues to apply between 12:00 and 18:00 UTC from Monday to Friday. This time-based pricing model mirrors traditional cloud computing strategies pioneered by providers like AWS, bringing utility-style pricing dynamics directly to large language model execution.

For founders and engineering leaders, this pricing shift offers a clear tactical advantage. High-performance LLMs often strain early-stage engineering budgets, particularly when running continuous integration pipelines, batch data processing, or large-scale evaluation tasks. By shifting non-urgent workloads to off-peak windows, startups can effectively cut their inference overhead in half without sacrificing model capability.

The introduction of off-peak discounting for frontier open models like DeepSeek-V4 highlights the maturing economics of AI deployment. As infrastructure providers compete for developer mindshare, competing primarily on raw capability is no longer sufficient. Cost efficiency, predictable billing, and flexible scheduling tools are becoming critical differentiators for cloud platforms hosting open models.

Business leaders must evaluate their internal workflows to capitalize on these savings. Automated testing routines, nightly batch jobs, and asynchronous data generation are prime candidates for off-peak execution. As infrastructure costs become more variable and manageable, engineering teams that adapt their operational schedules will capture a distinct margin advantage.

Sources & References

Web Sources

Newsletter Sources

Ollama - New off-peak rates: 50% lower prices for DeepSeek-V4

Share this intelligence briefing

Pass insights along to your team and network.

Related Stories

AI & Machine Learning

Hardware Bottlenecks and Model Trade-offs: Inside the Local LLM Community's Qwen3.8 Evaluation

2 min read

AI & Machine Learning

Ollama Cuts DeepSeek-V4 Cloud Prices by 50 Percent for Off-Peak Hours

3 min read

AI & Machine Learning

Unlocking GPU Efficiency: Why Continuous Batching is the New Frontier for LLM Infrastructure

3 min read