Articles Tagged “Quantization”
7 articles found

Transformers Now Runs GGUF Quants on Mac at llama.cpp Speed
Hugging Face Transformers can now run GGUF quants packed on Apple Silicon. A Qwen3.5-4B Q4_K_M model shrinks from 8.42GB to 2.74GB. Here's how it works.

How Much RAM Do You Need to Run a Local LLM in 2026?
A practical sizing guide: 8GB runs an 8B model at 4-bit, 24GB handles a 27B, and 240GB is the floor for a trillion-parameter model at 1-bit.

Qwen3.8-Flash-Next Fits in 75GB With No GPU Needed
Qwen's 125B Flash-Next MoE runs locally in 75GB of RAM with no GPU VRAM required, scoring 62.5 on SWE-bench Pro with just 6B active parameters.

GLM-5.3-Flash 3-Bit Quant Runs on 128GB of Local RAM
Z.ai's 320B GLM-5.3-Flash now runs at 3-bit on 128GB of RAM via Unsloth GGUFs, retaining 82% of top-1 accuracy at under a fifth of its 650GB size.

On-Device Piano Model Autocompletes Music on iPhone
A 125M-parameter transformer generates about 108 piano notes per second on an iPhone 15, trained on 300 million note events with no cloud call.

LLM Quantization Guide: GGUF vs AWQ vs MLX in 2026
A practical guide to LLM quantization formats — GGUF, AWQ, GPTQ and MLX — with VRAM math, quality trade-offs and picks for every kind of machine.

Gemma 4 QAT Lands in Ollama, Cutting Local AI Memory by ~72%
Quantization-aware-trained Gemma 4 weights are now runnable in Ollama, cutting VRAM roughly 72% so a 26B model fits on a 16GB laptop for self-hosted AI.
