Articles Tagged “Ollama”
9 articles found
Local LLM Servers Compared: Ollama vs vLLM vs llama.cpp
Ollama, llama.cpp, vLLM and LM Studio all serve local models. Here's which one fits your hardware, from a 16GB laptop to a four-GPU workstation.
Gemma 4 Runs 90% Faster on Apple Silicon in Ollama
Ollama v0.32.0 makes Google's Gemma 4 nearly 90% faster on Apple Silicon via multi-token prediction — local AI on a laptop just got a lot snappier.
Best Mini PC for Local LLMs in 2026: A Buyer's Guide
The best mini PC for local LLMs comes down to unified memory and bandwidth. Our 2026 buyer's guide compares top picks from budget to 128GB powerhouses.
Ollama Raises $65M to Power Local Open-Source AI
Ollama closed a $65M Series B led by Theory Ventures on July 9, growing to 8.9M monthly developers and a presence in 85% of the Fortune 500.
Ollama v0.31.1 Boosts Local AI Performance on Apple Silicon
Ollama v0.31.1 makes Gemma 4 about 90% faster on Apple Silicon via multi-token prediction, advancing local AI performance and privacy.
Ollama 0.30.8 Widens Local AI Hardware Support and Speeds Up Apple Silicon
Ollama 0.30.8, released June 12, broadens GGUF hardware support through llama.cpp and upgrades its Apple Silicon MLX engine for faster, private local AI.
Gemma 4 QAT Lands in Ollama, Cutting Local AI Memory by ~72%
Quantization-aware-trained Gemma 4 weights are now runnable in Ollama, cutting VRAM roughly 72% so a 26B model fits on a 16GB laptop for self-hosted AI.
Ollama v0.24 Lands With Qwen 3.6 Support — Local AI Just Got a Major Upgrade for Self-Hosted LLM Builders
Ollama released v0.24.0 on May 14, 2026 with first-class support for Qwen 3.6 — bringing Alibaba's 35B-A3B mixture-of-experts model to anyone running local LLMs on their own hardware.
Qwen3.6 Arrives on Ollama: Run a 35B Agentic Coding AI Locally With 256K Context
Alibaba's Qwen3.6 is now on Ollama — a 35B open-weight model with 256K context, vision support, and thinking preservation built for agentic coding workflows you can run on your own hardware.






