Founders NorthFounders North
Back to Home

Ollama Cuts DeepSeek-V4 Cloud Prices by 50 Percent for Off-Peak Hours

Ollama has introduced a 50 percent price reduction for DeepSeek-V4-Flash and DeepSeek-V4-Pro token usage during off-peak hours, signaling a new wave of cost optimization for cloud AI builders.

Monday, September 7, 2026

Key Takeaways

  • Ollama has slashed token prices for DeepSeek-V4-Flash and DeepSeek-V4-Pro by 50 percent during designated off-peak hours.
  • Off-peak hours include weekdays before 12:00 UTC and after 18:00 UTC, as well as all day on weekends.
  • Models are hosted in the US and Europe and feature Zero Data Retention guarantees, addressing key enterprise security requirements.
  • Founders and engineering teams can leverage this pricing structure by routing non-urgent batch jobs and agent workflows into off-peak windows to reduce operational costs.

The economics of running advanced artificial intelligence models in the cloud are shifting rapidly. According to recent updates from Ollama, the platform has rolled out a significant 50 percent price reduction for DeepSeek-V4-Flash and DeepSeek-V4-Pro token usage. This discount applies during off-peak hours, defined as before 12:00 or after 18:00 UTC on weekdays, and covers all day on weekends.

For founders and engineering leaders managing heavy AI workloads, this pricing adjustment offers a tangible lever for margin expansion. By shifting non-urgent batch processing, background tasks, and asynchronous agent workflows into these discounted windows, organizations can effectively cut their token operational expenditures in half without sacrificing model capability.

The Mechanics of Off-Peak AI Infrastructure

Ollama hosts these models across infrastructure in the US and Europe, backing them with Zero Data Retention guarantees. This security posture is critical for enterprise builders who must comply with strict data privacy regulations while leveraging third-party cloud infrastructure. By pairing high-performance open models like DeepSeek-V4 with predictable geographic hosting and strict privacy guarantees, Ollama is positioning its cloud tier as a viable alternative to hyperscale proprietary APIs.

The introduction of time-based pricing models mirrors strategies long used in traditional cloud computing and telecommunications. Compute power and network bandwidth have always fluctuated in demand, but the artificial intelligence sector has historically maintained flat, around-the-clock pricing for API calls. Ollama's move suggests that AI infrastructure is maturing, adopting utility-style pricing models that incentivize load balancing across global data centers.

Strategic Implications for Founders and Builders

For early-stage startups and scaling enterprises alike, infrastructure costs can quickly become a bottleneck. When scaling agentic workflows or high-frequency inference applications, token costs accumulate aggressively. The 50 percent discount on DeepSeek-V4 models during weekends and weekday off-peak hours allows technical teams to rethink their architecture.

Engineers can design queuing systems that hold non-interactive tasks, such as content summarization, dataset processing, or large-scale evaluation runs, and execute them automatically during off-peak windows. This operational discipline requires minor architectural adjustments but yields major cost savings. As competition among AI infrastructure providers intensifies, founders should expect more platforms to experiment with dynamic pricing models to capture utilization during troughs in global demand.

Ultimately, Ollama's latest update demonstrates that the commoditization of foundational intelligence is driving providers to compete not just on model quality, but on the granular economics of delivery. For builders willing to optimize their workflows around these new time-based rates, the barrier to deploying capable AI systems continues to drop.

Sources & References

Web Sources

Newsletter Sources

Ollama - New off-peak rates: 50% lower prices for DeepSeek-V4

Share this intelligence briefing

Pass insights along to your team and network.

Related Stories

AI & Machine Learning

Hardware Bottlenecks and Model Trade-offs: Inside the Local LLM Community's Qwen3.8 Evaluation

2 min read

AI & Machine Learning

Unlocking GPU Efficiency: Why Continuous Batching is the New Frontier for LLM Infrastructure

3 min read

AI & Machine Learning

Ollama Introduces 50 Percent Off-Peak Discounts for DeepSeek-V4 Models

2 min read