Articles Tagged “Local LLM”
21 articles found

Qwen3.8-27B Runs a 262K-Context Vision Model Locally
Alibaba's Qwen3.8-27B lands under Apache 2.0 with vision, a 262K context, and a 17GB quantization that runs at 15-30 tokens per second on a laptop.

Gemma Translator: Offline AI Translation on a Pi 5
A Raspberry Pi 5 runs a full speech-to-speech interpreter offline using Gemma 4 E2B and LiteRT, at 9 tokens per second and 1,432 MB peak memory.

Local LLM Servers Compared: Ollama vs vLLM vs llama.cpp
Ollama, llama.cpp, vLLM and LM Studio all serve local models. Here's which one fits your hardware, from a 16GB laptop to a four-GPU workstation.

ESP32-S3 Runs a 28.9M-Parameter LLM at 9 Tokens/s
A developer got a 28.9M-parameter language model generating 9 tokens per second on an $8 ESP32-S3, using 4-bit weights and per-layer embeddings in flash.

NightRun Boots a Local LLM on a Pi 5 With No OS
NightRun is a Rust UEFI application that boots straight into a local LLM on Raspberry Pi 5 and x86 PCs, skipping the operating system entirely.

LLM Quantization Guide: GGUF vs AWQ vs MLX in 2026
A practical guide to LLM quantization formats — GGUF, AWQ, GPTQ and MLX — with VRAM math, quality trade-offs and picks for every kind of machine.

Raspberry Pi AI Projects Book Covers Local LLMs for £9
Raspberry Pi Press's new AI Projects book covers local LLMs, vision and speech across Pi 5, Pi Zero 2 W and Pico, at an intro price of £8.99.
PyTorch 2.13 Brings FlexAttention to Apple Silicon
PyTorch 2.13 landed July 8 with FlexAttention on Apple Silicon — up to 12x faster attention on Mac GPUs and 4x lower memory for LM training.
Gemma 4 Runs 90% Faster on Apple Silicon in Ollama
Ollama v0.32.0 makes Google's Gemma 4 nearly 90% faster on Apple Silicon via multi-token prediction — local AI on a laptop just got a lot snappier.

Best Mini PC for Local LLMs in 2026: A Buyer's Guide
The best mini PC for local LLMs comes down to unified memory and bandwidth. Our 2026 buyer's guide compares top picks from budget to 128GB powerhouses.

Best Mini PC for Local LLMs in 2026: A Buyer's Guide
A practical 2026 buyer's guide to the best mini PCs for running local LLMs, comparing unified memory, NPUs, and price so you can self-host with confidence.

A Local AI Agent Matches Tumor Boards on Blood-Cancer Decisions
A peer-reviewed Nature Medicine study shows a locally run LLM agent matching expert hematology tumor boards while keeping patient data private on-site.

Firefly AIBOX-9075 Brings 200 TOPS and Local LLMs to the Edge
Firefly's AIBOX-9075 edge AI box pairs a Qualcomm Dragonwing IQ-9075 with up to 200 TOPS and 36GB RAM to run private, on-device LLMs — detailed June 26, 2026.

GMK EVO-X3 Opens Early Access for Its 128GB Strix Halo Mini PC
GMK opened early-access registration on June 22 for the EVO-X3, a Ryzen AI Max+ 395 Strix Halo mini PC with up to 128GB of RAM and an OCuLink port — a compact powerhouse for local AI and creative work.

Gemma 4 12B Brings Full Multimodal AI to a 16GB Laptop — Free Under Apache 2.0
Google DeepMind released Gemma 4 12B on June 3, 2026 — an open-weight, encoder-free multimodal model with native audio that runs locally on a 16GB consumer laptop.

Acer Veriton RA110: A Six-Inch AI Mini Workstation With Ryzen AI Max+ 395
Acer's Veriton RA110 packs a Ryzen AI Max+ 395, 126 TOPS, and 128GB unified memory into a six-inch-square AI mini workstation built to run local LLMs.

Ollama v0.24 Lands With Qwen 3.6 Support — Local AI Just Got a Major Upgrade for Self-Hosted LLM Builders
Ollama released v0.24.0 on May 14, 2026 with first-class support for Qwen 3.6 — bringing Alibaba's 35B-A3B mixture-of-experts model to anyone running local LLMs on their own hardware.

Google Drops Multi-Token Prediction Drafters for Gemma 4 — Up to 3x Faster Local LLM Inference With Zero Quality Loss
On May 5, 2026 Google released open Multi-Token Prediction drafters for the Gemma 4 family, delivering up to 3x faster local LLM inference without any quality loss — Apache 2.0 licensed.

Orange Pi AI Station Brings 176 TOPS and 96GB RAM to the DIY AI Workbench
Orange Pi's new AI Station packs a Huawei Ascend 310 SoC with 176 TOPS of AI performance and up to 96GB LPDDR4X RAM into a maker-friendly mini PC.

Sapphire's Strix Halo Mini PC Packs a Ryzen AI Max+ 395 With 128GB RAM and RTX 4070-Class Graphics Into a Tiny Box
Demoed at Embedded World 2026, the Sapphire Edge AI Max+ 395 runs 16 Zen 5 cores at 5.1 GHz, a Radeon 8060S iGPU, and can link two units via USB-C for pooled LLM inference.

AMD Strix Halo Mini PCs Are Here — And They Can Run 120-Billion-Parameter AI Models Locally
A wave of AMD Ryzen AI Max+ 395 mini PCs is shipping with 128GB unified memory, bringing serious local AI inference to a box on your desk.
