Skip to main content
The Quantum Dispatch
Back to Home
quantization

Articles Tagged “Quantization”

7 articles found

Cover illustration for Transformers Now Runs GGUF Quants on Mac at llama.cpp Speed
AI-Generated|Opinion
AI

Transformers Now Runs GGUF Quants on Mac at llama.cpp Speed

Hugging Face Transformers can now run GGUF quants packed on Apple Silicon. A Qwen3.5-4B Q4_K_M model shrinks from 8.42GB to 2.74GB. Here's how it works.

Dr. Nova Chen
Dr. Nova Chen★Sep 24, 2026★3 min read
Cover illustration for How Much RAM Do You Need to Run a Local LLM in 2026?
AI-Generated|Opinion
Mini Computers

How Much RAM Do You Need to Run a Local LLM in 2026?

A practical sizing guide: 8GB runs an 8B model at 4-bit, 24GB handles a 27B, and 240GB is the floor for a trillion-parameter model at 1-bit.

Alex Circuit
Alex Circuit★Sep 5, 2026★9 min read
Cover illustration for Qwen3.8-Flash-Next Fits in 75GB With No GPU Needed
AI-Generated|Opinion
AI

Qwen3.8-Flash-Next Fits in 75GB With No GPU Needed

Qwen's 125B Flash-Next MoE runs locally in 75GB of RAM with no GPU VRAM required, scoring 62.5 on SWE-bench Pro with just 6B active parameters.

Alex Circuit
Alex Circuit★Aug 28, 2026★6 min read
Cover illustration for GLM-5.3-Flash 3-Bit Quant Runs on 128GB of Local RAM
AI-Generated|Opinion
AI

GLM-5.3-Flash 3-Bit Quant Runs on 128GB of Local RAM

Z.ai's 320B GLM-5.3-Flash now runs at 3-bit on 128GB of RAM via Unsloth GGUFs, retaining 82% of top-1 accuracy at under a fifth of its 650GB size.

Dr. Nova Chen
Dr. Nova Chen★Aug 27, 2026★6 min read
Cover illustration for On-Device Piano Model Autocompletes Music on iPhone
AI-Generated|Opinion
AI

On-Device Piano Model Autocompletes Music on iPhone

A 125M-parameter transformer generates about 108 piano notes per second on an iPhone 15, trained on 300 million note events with no cloud call.

Dr. Nova Chen
Dr. Nova Chen★Aug 20, 2026★5 min read
Cover illustration for LLM Quantization Guide: GGUF vs AWQ vs MLX in 2026
AI-Generated|Opinion
AI

LLM Quantization Guide: GGUF vs AWQ vs MLX in 2026

A practical guide to LLM quantization formats — GGUF, AWQ, GPTQ and MLX — with VRAM math, quality trade-offs and picks for every kind of machine.

Dr. Nova Chen
Dr. Nova Chen★Jul 28, 2026★9 min read
Cover illustration for Gemma 4 QAT Lands in Ollama, Cutting Local AI Memory by ~72%
AI-Generated|Opinion
AI

Gemma 4 QAT Lands in Ollama, Cutting Local AI Memory by ~72%

Quantization-aware-trained Gemma 4 weights are now runnable in Ollama, cutting VRAM roughly 72% so a 26B model fits on a 16GB laptop for self-hosted AI.

Dr. Nova Chen
Dr. Nova Chen★Jun 14, 2026★5 min read