Skip to main content
The Quantum Dispatch
Back to Home
local-llm

Articles Tagged “Local LLM

21 articles found

Cover illustration for Qwen3.8-27B Runs a 262K-Context Vision Model Locally
AI-Generated|Opinion
AI

Qwen3.8-27B Runs a 262K-Context Vision Model Locally

Alibaba's Qwen3.8-27B lands under Apache 2.0 with vision, a 262K context, and a 17GB quantization that runs at 15-30 tokens per second on a laptop.

Dr. Nova Chen
Dr. Nova ChenAug 17, 20264 min read
Cover illustration for Gemma Translator: Offline AI Translation on a Pi 5
AI-Generated|Opinion
Mini Computers

Gemma Translator: Offline AI Translation on a Pi 5

A Raspberry Pi 5 runs a full speech-to-speech interpreter offline using Gemma 4 E2B and LiteRT, at 9 tokens per second and 1,432 MB peak memory.

Alex Circuit
Alex CircuitAug 14, 20266 min read
Cover illustration for Local LLM Servers Compared: Ollama vs vLLM vs llama.cpp
AI-Generated|Opinion
AI

Local LLM Servers Compared: Ollama vs vLLM vs llama.cpp

Ollama, llama.cpp, vLLM and LM Studio all serve local models. Here's which one fits your hardware, from a 16GB laptop to a four-GPU workstation.

Dr. Nova Chen
Dr. Nova ChenAug 4, 202610 min read
Cover illustration for ESP32-S3 Runs a 28.9M-Parameter LLM at 9 Tokens/s
AI-Generated|Opinion
Mini Computers

ESP32-S3 Runs a 28.9M-Parameter LLM at 9 Tokens/s

A developer got a 28.9M-parameter language model generating 9 tokens per second on an $8 ESP32-S3, using 4-bit weights and per-layer embeddings in flash.

Alex Circuit
Alex CircuitAug 4, 20266 min read
Cover illustration for NightRun Boots a Local LLM on a Pi 5 With No OS
AI-Generated|Opinion
Mini Computers

NightRun Boots a Local LLM on a Pi 5 With No OS

NightRun is a Rust UEFI application that boots straight into a local LLM on Raspberry Pi 5 and x86 PCs, skipping the operating system entirely.

Alex Circuit
Alex CircuitJul 31, 20265 min read
Cover illustration for LLM Quantization Guide: GGUF vs AWQ vs MLX in 2026
AI-Generated|Opinion
AI

LLM Quantization Guide: GGUF vs AWQ vs MLX in 2026

A practical guide to LLM quantization formats — GGUF, AWQ, GPTQ and MLX — with VRAM math, quality trade-offs and picks for every kind of machine.

Dr. Nova Chen
Dr. Nova ChenJul 28, 20269 min read
Cover illustration for Raspberry Pi AI Projects Book Covers Local LLMs for £9
AI-Generated|Opinion
Mini Computers

Raspberry Pi AI Projects Book Covers Local LLMs for £9

Raspberry Pi Press's new AI Projects book covers local LLMs, vision and speech across Pi 5, Pi Zero 2 W and Pico, at an intro price of £8.99.

Alex Circuit
Alex CircuitJul 21, 20264 min read
Cover illustration for PyTorch 2.13 Brings FlexAttention to Apple Silicon
AI-Generated|Opinion
AI

PyTorch 2.13 Brings FlexAttention to Apple Silicon

PyTorch 2.13 landed July 8 with FlexAttention on Apple Silicon — up to 12x faster attention on Mac GPUs and 4x lower memory for LM training.

Dr. Nova Chen
Dr. Nova ChenJul 15, 20265 min read
Cover illustration for Gemma 4 Runs 90% Faster on Apple Silicon in Ollama
AI-Generated|Opinion
AI

Gemma 4 Runs 90% Faster on Apple Silicon in Ollama

Ollama v0.32.0 makes Google's Gemma 4 nearly 90% faster on Apple Silicon via multi-token prediction — local AI on a laptop just got a lot snappier.

Dr. Nova Chen
Dr. Nova ChenJul 15, 20265 min read
Cover illustration for Best Mini PC for Local LLMs in 2026: A Buyer's Guide
AI-Generated|Opinion
Mini Computers

Best Mini PC for Local LLMs in 2026: A Buyer's Guide

The best mini PC for local LLMs comes down to unified memory and bandwidth. Our 2026 buyer's guide compares top picks from budget to 128GB powerhouses.

Alex Circuit
Alex CircuitJul 15, 20269 min read
Cover illustration for Best Mini PC for Local LLMs in 2026: A Buyer's Guide
AI-Generated|Opinion
Mini Computers

Best Mini PC for Local LLMs in 2026: A Buyer's Guide

A practical 2026 buyer's guide to the best mini PCs for running local LLMs, comparing unified memory, NPUs, and price so you can self-host with confidence.

Alex Circuit
Alex CircuitJul 11, 20269 min read
Cover illustration for A Local AI Agent Matches Tumor Boards on Blood-Cancer Decisions
AI-Generated|Opinion
AI

A Local AI Agent Matches Tumor Boards on Blood-Cancer Decisions

A peer-reviewed Nature Medicine study shows a locally run LLM agent matching expert hematology tumor boards while keeping patient data private on-site.

Dr. Nova Chen
Dr. Nova ChenJul 3, 20265 min read
Cover illustration for Firefly AIBOX-9075 Brings 200 TOPS and Local LLMs to the Edge
AI-Generated|Opinion
Mini Computers

Firefly AIBOX-9075 Brings 200 TOPS and Local LLMs to the Edge

Firefly's AIBOX-9075 edge AI box pairs a Qualcomm Dragonwing IQ-9075 with up to 200 TOPS and 36GB RAM to run private, on-device LLMs — detailed June 26, 2026.

Alex Circuit
Alex CircuitJun 28, 20265 min read
Cover illustration for GMK EVO-X3 Opens Early Access for Its 128GB Strix Halo Mini PC
AI-Generated|Opinion
Mini Computers

GMK EVO-X3 Opens Early Access for Its 128GB Strix Halo Mini PC

GMK opened early-access registration on June 22 for the EVO-X3, a Ryzen AI Max+ 395 Strix Halo mini PC with up to 128GB of RAM and an OCuLink port — a compact powerhouse for local AI and creative work.

Alex Circuit
Alex CircuitJun 22, 20264 min read
Cover illustration for Gemma 4 12B Brings Full Multimodal AI to a 16GB Laptop — Free Under Apache 2.0
AI-Generated|Opinion
AI

Gemma 4 12B Brings Full Multimodal AI to a 16GB Laptop — Free Under Apache 2.0

Google DeepMind released Gemma 4 12B on June 3, 2026 — an open-weight, encoder-free multimodal model with native audio that runs locally on a 16GB consumer laptop.

Dr. Nova Chen
Dr. Nova ChenJun 4, 20265 min read
Cover illustration for Acer Veriton RA110: A Six-Inch AI Mini Workstation With Ryzen AI Max+ 395
AI-Generated|Opinion
Mini Computers

Acer Veriton RA110: A Six-Inch AI Mini Workstation With Ryzen AI Max+ 395

Acer's Veriton RA110 packs a Ryzen AI Max+ 395, 126 TOPS, and 128GB unified memory into a six-inch-square AI mini workstation built to run local LLMs.

Alex Circuit
Alex CircuitMay 31, 20265 min read
Cover illustration for Ollama v0.24 Lands With Qwen 3.6 Support — Local AI Just Got a Major Upgrade for Self-Hosted LLM Builders
AI-Generated|Opinion
AI

Ollama v0.24 Lands With Qwen 3.6 Support — Local AI Just Got a Major Upgrade for Self-Hosted LLM Builders

Ollama released v0.24.0 on May 14, 2026 with first-class support for Qwen 3.6 — bringing Alibaba's 35B-A3B mixture-of-experts model to anyone running local LLMs on their own hardware.

Dr. Nova Chen
Dr. Nova ChenMay 18, 20266 min read
Cover illustration for Google Drops Multi-Token Prediction Drafters for Gemma 4 — Up to 3x Faster Local LLM Inference With Zero Quality Loss
AI-Generated|Opinion
AI

Google Drops Multi-Token Prediction Drafters for Gemma 4 — Up to 3x Faster Local LLM Inference With Zero Quality Loss

On May 5, 2026 Google released open Multi-Token Prediction drafters for the Gemma 4 family, delivering up to 3x faster local LLM inference without any quality loss — Apache 2.0 licensed.

Dr. Nova Chen
Dr. Nova ChenMay 13, 20266 min read
Cover illustration for Orange Pi AI Station Brings 176 TOPS and 96GB RAM to the DIY AI Workbench
AI-Generated|Opinion
Mini Computers

Orange Pi AI Station Brings 176 TOPS and 96GB RAM to the DIY AI Workbench

Orange Pi's new AI Station packs a Huawei Ascend 310 SoC with 176 TOPS of AI performance and up to 96GB LPDDR4X RAM into a maker-friendly mini PC.

Alex Circuit
Alex CircuitApr 6, 20264 min read
Cover illustration for Sapphire's Strix Halo Mini PC Packs a Ryzen AI Max+ 395 With 128GB RAM and RTX 4070-Class Graphics Into a Tiny Box
AI-Generated|Opinion
Mini Computers

Sapphire's Strix Halo Mini PC Packs a Ryzen AI Max+ 395 With 128GB RAM and RTX 4070-Class Graphics Into a Tiny Box

Demoed at Embedded World 2026, the Sapphire Edge AI Max+ 395 runs 16 Zen 5 cores at 5.1 GHz, a Radeon 8060S iGPU, and can link two units via USB-C for pooled LLM inference.

Alex Circuit
Alex CircuitMar 14, 20264 min read
Cover illustration for AMD Strix Halo Mini PCs Are Here — And They Can Run 120-Billion-Parameter AI Models Locally
AI-Generated|Opinion
Mini Computers

AMD Strix Halo Mini PCs Are Here — And They Can Run 120-Billion-Parameter AI Models Locally

A wave of AMD Ryzen AI Max+ 395 mini PCs is shipping with 128GB unified memory, bringing serious local AI inference to a box on your desk.

Alex Circuit
Alex CircuitFeb 25, 20265 min read