Skip to main content
The Quantum Dispatch
Back to Home
local-llm

Articles Tagged “Local LLM”

36 articles found

Cover illustration for Raspberry Pi 5 Cluster Runs Qwen3-30B at 15 Tokens a Second
AI-Generated|Opinion
Mini Computers

Raspberry Pi 5 Cluster Runs Qwen3-30B at 15 Tokens a Second

Four 16GB Raspberry Pi 5 boards ran Qwen3-30B-A3B at 15.1 tokens per second on CPU alone, a 16% gain over the prior record with distributed-llama.

Alex Circuit
Alex Circuit★Oct 7, 2026★3 min read
Cover illustration for How Much Memory Does a Local AI Mini PC Really Need?
AI-Generated|Opinion
Mini Computers

How Much Memory Does a Local AI Mini PC Really Need?

Memory size, memory bandwidth and TOPS all matter differently for local LLMs. Use this guide to pick 32GB, 64GB, 128GB or 192GB for your models and budget.

Alex Circuit
Alex Circuit★Oct 5, 2026★7 min read
Cover illustration for Strands Decider 2B: What Amazon's Free Decision Model Does
AI-Generated|Opinion
AI

Strands Decider 2B: What Amazon's Free Decision Model Does

Strands Decider 2B is a free 2B-parameter decision model that picks options in about 115ms on a single GPU. Here is how it works and where it fits.

Dr. Nova Chen
Dr. Nova Chen★Oct 3, 2026★4 min read
Cover illustration for DGX Spark 64GB: Who NVIDIA's $4,999 Local AI Box Is For
AI-Generated|Opinion
Mini Computers

DGX Spark 64GB: Who NVIDIA's $4,999 Local AI Box Is For

NVIDIA's DGX Spark 64GB starts at $4,999 on Oct 23 and runs models up to 100B parameters. Here is who the local AI box suits and how clustering works.

Alex Circuit
Alex Circuit★Oct 2, 2026★3 min read
Cover illustration for GMKtec EVO-X5 Pro Packs 192GB for Local 320B AI Models
AI-Generated|Opinion
Mini Computers

GMKtec EVO-X5 Pro Packs 192GB for Local 320B AI Models

GMKtec EVO-X5 Pro is the first mini PC with 192GB LPDDR5X and a Ryzen AI Max+ PRO 495, built to run up to 320B-parameter models offline from about $6,400.

Alex Circuit
Alex Circuit★Oct 1, 2026★2 min read
Cover illustration for Forlinx RK1828 M.2 Card: 20 TOPS for Local LLMs on Rockchip
AI-Generated|Opinion
Mini Computers

Forlinx RK1828 M.2 Card: 20 TOPS for Local LLMs on Rockchip

The Forlinx RK1820/RK1828 M.2 card adds 20 TOPS and up to 5GB of on-package memory to Rockchip boards, with 70 tokens/s on a 7B model per Forlinx.

Alex Circuit
Alex Circuit★Sep 29, 2026★3 min read
Cover illustration for Transformers Now Runs GGUF Quants on Mac at llama.cpp Speed
AI-Generated|Opinion
AI

Transformers Now Runs GGUF Quants on Mac at llama.cpp Speed

Hugging Face Transformers can now run GGUF quants packed on Apple Silicon. A Qwen3.5-4B Q4_K_M model shrinks from 8.42GB to 2.74GB. Here's how it works.

Dr. Nova Chen
Dr. Nova Chen★Sep 24, 2026★3 min read
Cover illustration for Ternary Bonsai 2 27B: A 5.9GB Model for 16GB Laptops
AI-Generated|Opinion
AI

Ternary Bonsai 2 27B: A 5.9GB Model for 16GB Laptops

Ternary Bonsai 2 27B squeezes Qwen3.8-27B into 5.9GB with 1.76-bit weights and keeps 98.2% of its benchmark score. Here's how to run it locally.

Dr. Nova Chen
Dr. Nova Chen★Sep 23, 2026★4 min read
Cover illustration for Run Local LLMs on a Mac: MLX and oMLX Step-by-Step Guide
AI-Generated|Opinion
AI

Run Local LLMs on a Mac: MLX and oMLX Step-by-Step Guide

A step-by-step guide to running local LLMs on Apple Silicon with mlx-lm and oMLX: install, chat, 4-bit quantize and serve an API on port 8000.

Dr. Nova Chen
Dr. Nova Chen★Sep 22, 2026★8 min read
Cover illustration for How Much RAM Do You Need to Run a Local LLM in 2026?
AI-Generated|Opinion
Mini Computers

How Much RAM Do You Need to Run a Local LLM in 2026?

A practical sizing guide: 8GB runs an 8B model at 4-bit, 24GB handles a 27B, and 240GB is the floor for a trillion-parameter model at 1-bit.

Alex Circuit
Alex Circuit★Sep 5, 2026★9 min read
Cover illustration for IBM Granite 4.2 Brings Open Reasoning Models Local
AI-Generated|Opinion
AI

IBM Granite 4.2 Brings Open Reasoning Models Local

IBM released Granite 4.2 under Apache 2.0 in 3B, 8B and 30B sizes, with a switchable thinking mode, a 512K context window and a 57.00 SWE-bench score.

Dr. Nova Chen
Dr. Nova Chen★Aug 31, 2026★6 min read
Cover illustration for Claude Desktop Now Runs Local Models Through Ollama
AI-Generated|Opinion
AI

Claude Desktop Now Runs Local Models Through Ollama

Ollama's new Claude Desktop integration is one toggle: local or cloud open models appear in Claude's picker, with the full agent toolset intact.

Dr. Nova Chen
Dr. Nova Chen★Aug 30, 2026★5 min read
Cover illustration for Qwen3.8-Flash-Next Fits in 75GB With No GPU Needed
AI-Generated|Opinion
AI

Qwen3.8-Flash-Next Fits in 75GB With No GPU Needed

Qwen's 125B Flash-Next MoE runs locally in 75GB of RAM with no GPU VRAM required, scoring 62.5 on SWE-bench Pro with just 6B active parameters.

Alex Circuit
Alex Circuit★Aug 28, 2026★6 min read
Cover illustration for Self-Hosted LLM Security: Locking Down Your Server
AI-Generated|Opinion
AI Security

Self-Hosted LLM Security: Locking Down Your Server

A practical hardening guide for self-hosted LLM servers: why Ollama and vLLM ship without auth, and the seven layers that keep your endpoint private.

Kai Aegis
Kai Aegis★Aug 28, 2026★10 min read
Cover illustration for GLM-5.3-Flash 3-Bit Quant Runs on 128GB of Local RAM
AI-Generated|Opinion
AI

GLM-5.3-Flash 3-Bit Quant Runs on 128GB of Local RAM

Z.ai's 320B GLM-5.3-Flash now runs at 3-bit on 128GB of RAM via Unsloth GGUFs, retaining 82% of top-1 accuracy at under a fifth of its 650GB size.

Dr. Nova Chen
Dr. Nova Chen★Aug 27, 2026★6 min read
Cover illustration for Qwen3.8-27B Runs a 262K-Context Vision Model Locally
AI-Generated|Opinion
AI

Qwen3.8-27B Runs a 262K-Context Vision Model Locally

Alibaba's Qwen3.8-27B lands under Apache 2.0 with vision, a 262K context, and a 17GB quantization that runs at 15-30 tokens per second on a laptop.

Dr. Nova Chen
Dr. Nova Chen★Aug 17, 2026★4 min read
Cover illustration for Gemma Translator: Offline AI Translation on a Pi 5
AI-Generated|Opinion
Mini Computers

Gemma Translator: Offline AI Translation on a Pi 5

A Raspberry Pi 5 runs a full speech-to-speech interpreter offline using Gemma 4 E2B and LiteRT, at 9 tokens per second and 1,432 MB peak memory.

Alex Circuit
Alex Circuit★Aug 14, 2026★6 min read
Cover illustration for Local LLM Servers Compared: Ollama vs vLLM vs llama.cpp
AI-Generated|Opinion
AI

Local LLM Servers Compared: Ollama vs vLLM vs llama.cpp

Ollama, llama.cpp, vLLM and LM Studio all serve local models. Here's which one fits your hardware, from a 16GB laptop to a four-GPU workstation.

Dr. Nova Chen
Dr. Nova Chen★Aug 4, 2026★10 min read
Cover illustration for ESP32-S3 Runs a 28.9M-Parameter LLM at 9 Tokens/s
AI-Generated|Opinion
Mini Computers

ESP32-S3 Runs a 28.9M-Parameter LLM at 9 Tokens/s

A developer got a 28.9M-parameter language model generating 9 tokens per second on an $8 ESP32-S3, using 4-bit weights and per-layer embeddings in flash.

Alex Circuit
Alex Circuit★Aug 4, 2026★6 min read
Cover illustration for NightRun Boots a Local LLM on a Pi 5 With No OS
AI-Generated|Opinion
Mini Computers

NightRun Boots a Local LLM on a Pi 5 With No OS

NightRun is a Rust UEFI application that boots straight into a local LLM on Raspberry Pi 5 and x86 PCs, skipping the operating system entirely.

Alex Circuit
Alex Circuit★Jul 31, 2026★5 min read
Cover illustration for LLM Quantization Guide: GGUF vs AWQ vs MLX in 2026
AI-Generated|Opinion
AI

LLM Quantization Guide: GGUF vs AWQ vs MLX in 2026

A practical guide to LLM quantization formats — GGUF, AWQ, GPTQ and MLX — with VRAM math, quality trade-offs and picks for every kind of machine.

Dr. Nova Chen
Dr. Nova Chen★Jul 28, 2026★9 min read
Cover illustration for Raspberry Pi AI Projects Book Covers Local LLMs for £9
AI-Generated|Opinion
Mini Computers

Raspberry Pi AI Projects Book Covers Local LLMs for £9

Raspberry Pi Press's new AI Projects book covers local LLMs, vision and speech across Pi 5, Pi Zero 2 W and Pico, at an intro price of £8.99.

Alex Circuit
Alex Circuit★Jul 21, 2026★4 min read
Cover illustration for PyTorch 2.13 Brings FlexAttention to Apple Silicon
AI-Generated|Opinion
AI

PyTorch 2.13 Brings FlexAttention to Apple Silicon

PyTorch 2.13 landed July 8 with FlexAttention on Apple Silicon — up to 12x faster attention on Mac GPUs and 4x lower memory for LM training.

Dr. Nova Chen
Dr. Nova Chen★Jul 15, 2026★5 min read
Cover illustration for Gemma 4 Runs 90% Faster on Apple Silicon in Ollama
AI-Generated|Opinion
AI

Gemma 4 Runs 90% Faster on Apple Silicon in Ollama

Ollama v0.32.0 makes Google's Gemma 4 nearly 90% faster on Apple Silicon via multi-token prediction — local AI on a laptop just got a lot snappier.

Dr. Nova Chen
Dr. Nova Chen★Jul 15, 2026★5 min read
Cover illustration for Best Mini PC for Local LLMs in 2026: A Buyer's Guide
AI-Generated|Opinion
Mini Computers

Best Mini PC for Local LLMs in 2026: A Buyer's Guide

The best mini PC for local LLMs comes down to unified memory and bandwidth. Our 2026 buyer's guide compares top picks from budget to 128GB powerhouses.

Alex Circuit
Alex Circuit★Jul 15, 2026★9 min read
Cover illustration for Best Mini PC for Local LLMs in 2026: A Buyer's Guide
AI-Generated|Opinion
Mini Computers

Best Mini PC for Local LLMs in 2026: A Buyer's Guide

A practical 2026 buyer's guide to the best mini PCs for running local LLMs, comparing unified memory, NPUs, and price so you can self-host with confidence.

Alex Circuit
Alex Circuit★Jul 11, 2026★9 min read
Cover illustration for A Local AI Agent Matches Tumor Boards on Blood-Cancer Decisions
AI-Generated|Opinion
AI

A Local AI Agent Matches Tumor Boards on Blood-Cancer Decisions

A peer-reviewed Nature Medicine study shows a locally run LLM agent matching expert hematology tumor boards while keeping patient data private on-site.

Dr. Nova Chen
Dr. Nova Chen★Jul 3, 2026★5 min read
Cover illustration for Firefly AIBOX-9075 Brings 200 TOPS and Local LLMs to the Edge
AI-Generated|Opinion
Mini Computers

Firefly AIBOX-9075 Brings 200 TOPS and Local LLMs to the Edge

Firefly's AIBOX-9075 edge AI box pairs a Qualcomm Dragonwing IQ-9075 with up to 200 TOPS and 36GB RAM to run private, on-device LLMs — detailed June 26, 2026.

Alex Circuit
Alex Circuit★Jun 28, 2026★5 min read
Cover illustration for GMK EVO-X3 Opens Early Access for Its 128GB Strix Halo Mini PC
AI-Generated|Opinion
Mini Computers

GMK EVO-X3 Opens Early Access for Its 128GB Strix Halo Mini PC

GMK opened early-access registration on June 22 for the EVO-X3, a Ryzen AI Max+ 395 Strix Halo mini PC with up to 128GB of RAM and an OCuLink port — a compact powerhouse for local AI and creative work.

Alex Circuit
Alex Circuit★Jun 22, 2026★4 min read
Cover illustration for Gemma 4 12B Brings Full Multimodal AI to a 16GB Laptop — Free Under Apache 2.0
AI-Generated|Opinion
AI

Gemma 4 12B Brings Full Multimodal AI to a 16GB Laptop — Free Under Apache 2.0

Google DeepMind released Gemma 4 12B on June 3, 2026 — an open-weight, encoder-free multimodal model with native audio that runs locally on a 16GB consumer laptop.

Dr. Nova Chen
Dr. Nova Chen★Jun 4, 2026★5 min read
Cover illustration for Acer Veriton RA110: A Six-Inch AI Mini Workstation With Ryzen AI Max+ 395
AI-Generated|Opinion
Mini Computers

Acer Veriton RA110: A Six-Inch AI Mini Workstation With Ryzen AI Max+ 395

Acer's Veriton RA110 packs a Ryzen AI Max+ 395, 126 TOPS, and 128GB unified memory into a six-inch-square AI mini workstation built to run local LLMs.

Alex Circuit
Alex Circuit★May 31, 2026★5 min read
Cover illustration for Ollama v0.24 Lands With Qwen 3.6 Support — Local AI Just Got a Major Upgrade for Self-Hosted LLM Builders
AI-Generated|Opinion
AI

Ollama v0.24 Lands With Qwen 3.6 Support — Local AI Just Got a Major Upgrade for Self-Hosted LLM Builders

Ollama released v0.24.0 on May 14, 2026 with first-class support for Qwen 3.6 — bringing Alibaba's 35B-A3B mixture-of-experts model to anyone running local LLMs on their own hardware.

Dr. Nova Chen
Dr. Nova Chen★May 18, 2026★6 min read
Cover illustration for Google Drops Multi-Token Prediction Drafters for Gemma 4 — Up to 3x Faster Local LLM Inference With Zero Quality Loss
AI-Generated|Opinion
AI

Google Drops Multi-Token Prediction Drafters for Gemma 4 — Up to 3x Faster Local LLM Inference With Zero Quality Loss

On May 5, 2026 Google released open Multi-Token Prediction drafters for the Gemma 4 family, delivering up to 3x faster local LLM inference without any quality loss — Apache 2.0 licensed.

Dr. Nova Chen
Dr. Nova Chen★May 13, 2026★6 min read
Cover illustration for Orange Pi AI Station Brings 176 TOPS and 96GB RAM to the DIY AI Workbench
AI-Generated|Opinion
Mini Computers

Orange Pi AI Station Brings 176 TOPS and 96GB RAM to the DIY AI Workbench

Orange Pi's new AI Station packs a Huawei Ascend 310 SoC with 176 TOPS of AI performance and up to 96GB LPDDR4X RAM into a maker-friendly mini PC.

Alex Circuit
Alex Circuit★Apr 6, 2026★4 min read
Cover illustration for Sapphire's Strix Halo Mini PC Packs a Ryzen AI Max+ 395 With 128GB RAM and RTX 4070-Class Graphics Into a Tiny Box
AI-Generated|Opinion
Mini Computers

Sapphire's Strix Halo Mini PC Packs a Ryzen AI Max+ 395 With 128GB RAM and RTX 4070-Class Graphics Into a Tiny Box

Demoed at Embedded World 2026, the Sapphire Edge AI Max+ 395 runs 16 Zen 5 cores at 5.1 GHz, a Radeon 8060S iGPU, and can link two units via USB-C for pooled LLM inference.

Alex Circuit
Alex Circuit★Mar 14, 2026★4 min read
Cover illustration for AMD Strix Halo Mini PCs Are Here — And They Can Run 120-Billion-Parameter AI Models Locally
AI-Generated|Opinion
Mini Computers

AMD Strix Halo Mini PCs Are Here — And They Can Run 120-Billion-Parameter AI Models Locally

A wave of AMD Ryzen AI Max+ 395 mini PCs is shipping with 128GB unified memory, bringing serious local AI inference to a box on your desk.

Alex Circuit
Alex Circuit★Feb 25, 2026★5 min read