Founders NorthFounders North
Back to Home

Ollama Introduces 50 Percent Off-Peak Discounts for DeepSeek-V4 Models

Ollama slashes token pricing for DeepSeek-V4-Flash and DeepSeek-V4-Pro by 50 percent during off-peak hours on its cloud platform.

Monday, September 7, 2026

Key Takeaways

  • Ollama has cut token prices by 50 percent for DeepSeek-V4-Pro and DeepSeek-V4-Flash models on its cloud platform.
  • Off-peak discounts apply before 12:00 UTC and after 18:00 UTC on weekdays, as well as all day on weekends.
  • Peak pricing remains active only between 12:00 and 18:00 UTC, Monday through Friday.
  • The pricing structure allows engineering teams to optimize operational costs by shifting batch workflows and non-urgent API requests to off-peak hours.

The battle for developer mindshare and cost efficiency in artificial intelligence infrastructure has taken a new turn. According to recent announcements from Ollama, the platform has introduced a 50 percent reduction in token pricing for DeepSeek-V4-Flash and DeepSeek-V4-Pro models on its cloud service. This aggressive pricing adjustment applies during designated off-peak hours, providing a strategic financial advantage for development teams and enterprise builders relying on open models.

Breaking Down the Off-Peak Structure

Under the new pricing schedule announced by Ollama, token prices for the DeepSeek-V4-Pro and DeepSeek-V4-Flash models are slashed by half during specific temporal windows. Specifically, these lower rates apply before 12:00 UTC or after 18:00 UTC on weekdays, encompassing a significant portion of the global workday. Furthermore, the discount extends to cover all day on Saturdays and Sundays. Peak pricing, by contrast, is strictly confined to the six-hour window between 12:00 and 18:00 UTC, Monday through Friday.

This time-based pricing model reflects broader trends in cloud computing and utility-style resource management, encouraging developers to shift batch processing, training verification, and automated workflows to off-peak hours in exchange for substantial cost savings.

Strategic Implications for Founders and Builders

For startup founders and engineering leaders, infrastructure costs represent a continuous operational hurdle as AI application usage scales. The integration of DeepSeek-V4 models into Ollama's cloud ecosystem, combined with this new off-peak discount structure, offers a compelling pathway to optimize unit economics. By scheduling non-urgent API calls, heavy data processing, and background agent execution during weekends or off-peak weekday hours, teams can effectively double their compute budget efficiency.

Moreover, Ollama continues to position its platform as a flexible environment for open model deployment. With tiered offerings ranging from free local execution to structured Team and Enterprise plans featuring centralized billing, budget controls, and shared usage credits, the addition of discounted DeepSeek tokens gives organizations greater flexibility in balancing performance against operational expenditure.

Sources & References

Web Sources

Newsletter Sources

Ollama - New off-peak rates: 50% lower prices for DeepSeek-V4

Share this intelligence briefing

Pass insights along to your team and network.

Related Stories

AI & Machine Learning

Hardware Bottlenecks and Model Trade-offs: Inside the Local LLM Community's Qwen3.8 Evaluation

2 min read

AI & Machine Learning

Ollama Cuts DeepSeek-V4 Cloud Prices by 50 Percent for Off-Peak Hours

3 min read

AI & Machine Learning

Unlocking GPU Efficiency: Why Continuous Batching is the New Frontier for LLM Infrastructure

3 min read