Articles Tagged “Gemma 4”
8 articles found

Gemma Downloads Top 900 Million Across Open Models
Google's Gemma open models have now passed 900 million downloads, with Gemma 4 alone contributing over 300 million since its April 2026 launch.

Gemma Translator: Offline AI Translation on a Pi 5
A Raspberry Pi 5 runs a full speech-to-speech interpreter offline using Gemma 4 E2B and LiteRT, at 9 tokens per second and 1,432 MB peak memory.
Ollama v0.31.1 Boosts Local AI Performance on Apple Silicon
Ollama v0.31.1 makes Gemma 4 about 90% faster on Apple Silicon via multi-token prediction, advancing local AI performance and privacy.

Gemma 4 QAT Lands in Ollama, Cutting Local AI Memory by ~72%
Quantization-aware-trained Gemma 4 weights are now runnable in Ollama, cutting VRAM roughly 72% so a 26B model fits on a 16GB laptop for self-hosted AI.

Gemma 4 12B Brings Full Multimodal AI to a 16GB Laptop — Free Under Apache 2.0
Google DeepMind released Gemma 4 12B on June 3, 2026 — an open-weight, encoder-free multimodal model with native audio that runs locally on a 16GB consumer laptop.

Google Drops Multi-Token Prediction Drafters for Gemma 4 — Up to 3x Faster Local LLM Inference With Zero Quality Loss
On May 5, 2026 Google released open Multi-Token Prediction drafters for the Gemma 4 family, delivering up to 3x faster local LLM inference without any quality loss — Apache 2.0 licensed.

Google Gemma 4 Comes to Android: On-Device AI in 140+ Languages, No Cloud Required
Google's AICore Developer Preview brings Gemma 4 natively to Android devices — offline, privacy-preserving AI inference in over 140 languages that upgrades automatically to Gemini Nano 4.

Google Gemma 4 Launches With Four Sizes, Apache 2.0 License, and a Top-3 Open Model Ranking
Google's Gemma 4 arrives with model sizes from 2B to 31B, a permissive Apache 2.0 license, native multimodal support across all sizes, and the #3 spot on the global open model leaderboard.
