NVIDIA DGX Spark performance
We ran performance tests on release day firmware and an updated Ollama version to see how Ollama performs.
We ran performance tests on release day firmware and an updated Ollama version to see how Ollama performs.
GLM-4.6 and Qwen3-coder-480B are available on Ollama’s cloud service with easy integrations to the tools you are familiar with. Qwen3-Coder-30B has been updated for faster, more…
The latest NVIDIA DGX Spark is here! Ollama has partnered with NVIDIA to ensure it runs fast and efficiently out-of-the-box.
A new web search API is now available in Ollama. Ollama provides a generous free tier of web searches for individuals to use, and higher rate…
Ollama now includes a significantly improved model scheduling system, reducing crashes due to out of memory issues, maximizing GPU utilization and performance, especially on multi-GPU systems.
Cloud models are now in preview, letting you run larger models with fast, datacenter-grade hardware. You can keep using your local tools while running larger models…
Ollama partners with OpenAI to bring gpt-oss to Ollama and its community.
Ollama's new app is now available for macOS and Windows.
Secure Minions is a secure protocol built by Stanford's Hazy Research lab to allow encrypted local-remote communication.
Ollama now has the ability to enable or disable thinking. This gives users the flexibility to choose the model’s thinking behavior for different applications and use…