Articles Tagged “Local LLM”
36 articles found

Raspberry Pi 5 Cluster Runs Qwen3-30B at 15 Tokens a Second
Four 16GB Raspberry Pi 5 boards ran Qwen3-30B-A3B at 15.1 tokens per second on CPU alone, a 16% gain over the prior record with distributed-llama.

How Much Memory Does a Local AI Mini PC Really Need?
Memory size, memory bandwidth and TOPS all matter differently for local LLMs. Use this guide to pick 32GB, 64GB, 128GB or 192GB for your models and budget.

Strands Decider 2B: What Amazon's Free Decision Model Does
Strands Decider 2B is a free 2B-parameter decision model that picks options in about 115ms on a single GPU. Here is how it works and where it fits.

DGX Spark 64GB: Who NVIDIA's $4,999 Local AI Box Is For
NVIDIA's DGX Spark 64GB starts at $4,999 on Oct 23 and runs models up to 100B parameters. Here is who the local AI box suits and how clustering works.

GMKtec EVO-X5 Pro Packs 192GB for Local 320B AI Models
GMKtec EVO-X5 Pro is the first mini PC with 192GB LPDDR5X and a Ryzen AI Max+ PRO 495, built to run up to 320B-parameter models offline from about $6,400.

Forlinx RK1828 M.2 Card: 20 TOPS for Local LLMs on Rockchip
The Forlinx RK1820/RK1828 M.2 card adds 20 TOPS and up to 5GB of on-package memory to Rockchip boards, with 70 tokens/s on a 7B model per Forlinx.

Transformers Now Runs GGUF Quants on Mac at llama.cpp Speed
Hugging Face Transformers can now run GGUF quants packed on Apple Silicon. A Qwen3.5-4B Q4_K_M model shrinks from 8.42GB to 2.74GB. Here's how it works.

Ternary Bonsai 2 27B: A 5.9GB Model for 16GB Laptops
Ternary Bonsai 2 27B squeezes Qwen3.8-27B into 5.9GB with 1.76-bit weights and keeps 98.2% of its benchmark score. Here's how to run it locally.

Run Local LLMs on a Mac: MLX and oMLX Step-by-Step Guide
A step-by-step guide to running local LLMs on Apple Silicon with mlx-lm and oMLX: install, chat, 4-bit quantize and serve an API on port 8000.

How Much RAM Do You Need to Run a Local LLM in 2026?
A practical sizing guide: 8GB runs an 8B model at 4-bit, 24GB handles a 27B, and 240GB is the floor for a trillion-parameter model at 1-bit.

IBM Granite 4.2 Brings Open Reasoning Models Local
IBM released Granite 4.2 under Apache 2.0 in 3B, 8B and 30B sizes, with a switchable thinking mode, a 512K context window and a 57.00 SWE-bench score.

Claude Desktop Now Runs Local Models Through Ollama
Ollama's new Claude Desktop integration is one toggle: local or cloud open models appear in Claude's picker, with the full agent toolset intact.

Qwen3.8-Flash-Next Fits in 75GB With No GPU Needed
Qwen's 125B Flash-Next MoE runs locally in 75GB of RAM with no GPU VRAM required, scoring 62.5 on SWE-bench Pro with just 6B active parameters.

Self-Hosted LLM Security: Locking Down Your Server
A practical hardening guide for self-hosted LLM servers: why Ollama and vLLM ship without auth, and the seven layers that keep your endpoint private.

GLM-5.3-Flash 3-Bit Quant Runs on 128GB of Local RAM
Z.ai's 320B GLM-5.3-Flash now runs at 3-bit on 128GB of RAM via Unsloth GGUFs, retaining 82% of top-1 accuracy at under a fifth of its 650GB size.

Qwen3.8-27B Runs a 262K-Context Vision Model Locally
Alibaba's Qwen3.8-27B lands under Apache 2.0 with vision, a 262K context, and a 17GB quantization that runs at 15-30 tokens per second on a laptop.

Gemma Translator: Offline AI Translation on a Pi 5
A Raspberry Pi 5 runs a full speech-to-speech interpreter offline using Gemma 4 E2B and LiteRT, at 9 tokens per second and 1,432 MB peak memory.

Local LLM Servers Compared: Ollama vs vLLM vs llama.cpp
Ollama, llama.cpp, vLLM and LM Studio all serve local models. Here's which one fits your hardware, from a 16GB laptop to a four-GPU workstation.

ESP32-S3 Runs a 28.9M-Parameter LLM at 9 Tokens/s
A developer got a 28.9M-parameter language model generating 9 tokens per second on an $8 ESP32-S3, using 4-bit weights and per-layer embeddings in flash.

NightRun Boots a Local LLM on a Pi 5 With No OS
NightRun is a Rust UEFI application that boots straight into a local LLM on Raspberry Pi 5 and x86 PCs, skipping the operating system entirely.

LLM Quantization Guide: GGUF vs AWQ vs MLX in 2026
A practical guide to LLM quantization formats — GGUF, AWQ, GPTQ and MLX — with VRAM math, quality trade-offs and picks for every kind of machine.

Raspberry Pi AI Projects Book Covers Local LLMs for £9
Raspberry Pi Press's new AI Projects book covers local LLMs, vision and speech across Pi 5, Pi Zero 2 W and Pico, at an intro price of £8.99.
PyTorch 2.13 Brings FlexAttention to Apple Silicon
PyTorch 2.13 landed July 8 with FlexAttention on Apple Silicon — up to 12x faster attention on Mac GPUs and 4x lower memory for LM training.
Gemma 4 Runs 90% Faster on Apple Silicon in Ollama
Ollama v0.32.0 makes Google's Gemma 4 nearly 90% faster on Apple Silicon via multi-token prediction — local AI on a laptop just got a lot snappier.

Best Mini PC for Local LLMs in 2026: A Buyer's Guide
The best mini PC for local LLMs comes down to unified memory and bandwidth. Our 2026 buyer's guide compares top picks from budget to 128GB powerhouses.

Best Mini PC for Local LLMs in 2026: A Buyer's Guide
A practical 2026 buyer's guide to the best mini PCs for running local LLMs, comparing unified memory, NPUs, and price so you can self-host with confidence.

A Local AI Agent Matches Tumor Boards on Blood-Cancer Decisions
A peer-reviewed Nature Medicine study shows a locally run LLM agent matching expert hematology tumor boards while keeping patient data private on-site.

Firefly AIBOX-9075 Brings 200 TOPS and Local LLMs to the Edge
Firefly's AIBOX-9075 edge AI box pairs a Qualcomm Dragonwing IQ-9075 with up to 200 TOPS and 36GB RAM to run private, on-device LLMs — detailed June 26, 2026.

GMK EVO-X3 Opens Early Access for Its 128GB Strix Halo Mini PC
GMK opened early-access registration on June 22 for the EVO-X3, a Ryzen AI Max+ 395 Strix Halo mini PC with up to 128GB of RAM and an OCuLink port — a compact powerhouse for local AI and creative work.

Gemma 4 12B Brings Full Multimodal AI to a 16GB Laptop — Free Under Apache 2.0
Google DeepMind released Gemma 4 12B on June 3, 2026 — an open-weight, encoder-free multimodal model with native audio that runs locally on a 16GB consumer laptop.

Acer Veriton RA110: A Six-Inch AI Mini Workstation With Ryzen AI Max+ 395
Acer's Veriton RA110 packs a Ryzen AI Max+ 395, 126 TOPS, and 128GB unified memory into a six-inch-square AI mini workstation built to run local LLMs.

Ollama v0.24 Lands With Qwen 3.6 Support — Local AI Just Got a Major Upgrade for Self-Hosted LLM Builders
Ollama released v0.24.0 on May 14, 2026 with first-class support for Qwen 3.6 — bringing Alibaba's 35B-A3B mixture-of-experts model to anyone running local LLMs on their own hardware.

Google Drops Multi-Token Prediction Drafters for Gemma 4 — Up to 3x Faster Local LLM Inference With Zero Quality Loss
On May 5, 2026 Google released open Multi-Token Prediction drafters for the Gemma 4 family, delivering up to 3x faster local LLM inference without any quality loss — Apache 2.0 licensed.

Orange Pi AI Station Brings 176 TOPS and 96GB RAM to the DIY AI Workbench
Orange Pi's new AI Station packs a Huawei Ascend 310 SoC with 176 TOPS of AI performance and up to 96GB LPDDR4X RAM into a maker-friendly mini PC.

Sapphire's Strix Halo Mini PC Packs a Ryzen AI Max+ 395 With 128GB RAM and RTX 4070-Class Graphics Into a Tiny Box
Demoed at Embedded World 2026, the Sapphire Edge AI Max+ 395 runs 16 Zen 5 cores at 5.1 GHz, a Radeon 8060S iGPU, and can link two units via USB-C for pooled LLM inference.

AMD Strix Halo Mini PCs Are Here — And They Can Run 120-Billion-Parameter AI Models Locally
A wave of AMD Ryzen AI Max+ 395 mini PCs is shipping with 128GB unified memory, bringing serious local AI inference to a box on your desk.
