Founders NorthFounders North
Back to Home

Ollama Cuts DeepSeek-V4 Cloud Pricing by 50 Percent During Off-Peak Hours

Ollama has introduced significant off-peak price reductions for its DeepSeek-V4 cloud models, offering developers cheaper AI compute alongside strict data privacy guarantees.

Monday, September 7, 2026

Key Takeaways

  • Ollama has slashed token prices by 50 percent for DeepSeek-V4-Pro and DeepSeek-V4-Flash models on its cloud platform.
  • The discounted rates apply during off-peak hours, defined as before 12:00 or after 18:00 UTC on weekdays, and all day on weekends.
  • The cost-effective compute options maintain Ollama's Zero Data Retention standards, ensuring privacy for enterprise and startup workloads.
  • Founders and engineering teams can leverage these pricing windows to optimize costs for batch processing and automated background AI tasks.

The battle for developer mindshare and cost efficiency in artificial intelligence infrastructure has entered a new phase. Ollama, a platform widely known for making local model execution seamless, has announced a 50 percent price reduction for token usage on its cloud platform, specifically targeting the DeepSeek-V4-Flash and DeepSeek-V4-Pro models. According to updates from Ollama, the steep discount applies during designated off-peak windows, specifically before 12:00 UTC and after 18:00 UTC on weekdays, as well as the entirety of weekends.

This pricing strategy signals a broader maturity in how AI cloud providers manage compute capacity. By incentivizing developers to shift their heavy workloads, batch processing, and non-urgent inference tasks to off-peak hours, Ollama can optimize its underlying GPU utilization while passing substantial savings directly to builders. Crucially, these lower rates come paired with the platform's Zero Data Retention policy, ensuring that companies do not have to compromise on data privacy to secure economical compute.

For founders and engineering leaders, the introduction of time-based discounting for high-performance open models like DeepSeek-V4 opens up fresh architectural possibilities. Historically, running advanced AI workloads meant absorbing flat, predictable, but often prohibitive API costs. With off-peak pricing models, startups can orchestrate automated background tasks, heavy data parsing, and model evaluations during cheaper windows, effectively cutting their operational expenditure in half for batch operations.

As competition in the foundational model and cloud hosting layer intensifies, pricing innovation is becoming just as critical as raw model performance. Ollama's move indicates that infrastructure providers are looking for creative ways to capture developer loyalty, moving beyond simple subscription tiers to dynamic pricing strategies that reflect the realities of global compute demand.

Sources & References

Web Sources

Newsletter Sources

Ollama - New off-peak rates: 50% lower prices for DeepSeek-V4

Share this intelligence briefing

Pass insights along to your team and network.

Related Stories

AI & Machine Learning

Hardware Bottlenecks and Model Trade-offs: Inside the Local LLM Community's Qwen3.8 Evaluation

2 min read

AI & Machine Learning

Ollama Cuts DeepSeek-V4 Cloud Prices by 50 Percent for Off-Peak Hours

3 min read

AI & Machine Learning

Unlocking GPU Efficiency: Why Continuous Batching is the New Frontier for LLM Infrastructure

3 min read