Articles Tagged “Local AI”
26 articles found

Bosgame M5 Brings 128GB Local AI to a Mini Desktop
ServeTheHome's Bosgame M5 review finds a Ryzen AI Max+ 395 box with 128GB of LPDDR5X that undercuts comparable local AI desktops by $500 or more.

LFM2.5-DSpark Speeds Local AI Inference Up to 3.2x
Liquid AI released ~300M-parameter draft models for the LFM2.5 family, delivering up to 3.18x GPU throughput and 57% lower function-calling latency.

Raspberry Pi CM5 Mini PC Runs OpenClaw Agents Locally
EDATEC's ED-CLAWBOX packs a Raspberry Pi CM5 into a 100mm passively cooled aluminum box for $267.19, running local OpenClaw AI agents on just 12W.

Gemma Downloads Top 900 Million Across Open Models
Google's Gemma open models have now passed 900 million downloads, with Gemma 4 alone contributing over 300 million since its April 2026 launch.

Muse Glimmer Runs a 30B Agentic Model on One Consumer GPU
Meta's Muse Glimmer is a 30B open-weight agentic model that compresses under 20GB at 4-bit, so a single 24GB consumer GPU can run it locally.

ZecTrix Note 4 Puts an AI Voice Notepad on E-Paper
ZecTrix Note 4 pairs a 4.2-inch E-Ink display with an ESP32-S3, NFC, and a 2,000 mAh battery for $54.99 — and a four-color variant lands near $30.

ASUS UGen300 Puts 40 TOPS of AI on a USB-C Stick
The $299.99 ASUS UGen300 pairs a Hailo-10H chip with 8GB LPDDR4 to deliver 40 TOPS over USB-C at 2.5W, adding local AI to Windows, Linux or Android hosts.

AI Accelerator Sticks Compared: Hailo, Coral, Jetson
A practical guide to bolting AI acceleration onto hardware you already own — Hailo-8L at $70, Coral USB at 4 TOPS, and when a Jetson is the better buy.

Liquid AI Encoders Hit 8K Context on CPU 3.7x Faster
Liquid AI's new LFM2.5-Encoders run 8,192-token inputs on a plain CPU roughly 3.7x faster than ModernBERT-base, from just 230M parameters and open weights.

Raspberry Pi AI Projects Book Covers Local LLMs for £9
Raspberry Pi Press's new AI Projects book covers local LLMs, vision and speech across Pi 5, Pi Zero 2 W and Pico, at an intro price of £8.99.

reCamera Pro Runs LLMs and Vision on a $300 AI Camera
Seeed Studio's reCamera Pro packs a 3 TOPS NPU to run vision, LLMs, speech-to-text, and TTS fully on-device with no cloud, from $299.90.

Ollama Raises $65M to Power Local Open-Source AI
Ollama closed a $65M Series B led by Theory Ventures on July 9, growing to 8.9M monthly developers and a presence in 85% of the Fortune 500.

Exo Labs' local.ai Helps You Run Frontier AI on Your Own Hardware
Exo Labs launched local.ai, a free platform that matches the best AI model to your own hardware and shows when running locally beats paying per API token.

Poolside's Laguna XS 2.1 Puts a Free Open-Weight Coding Model on Your Machine
Poolside released Laguna XS 2.1 on July 2, 2026 — a free, permissively licensed open-weight coding model that scores 70.9% on SWE-bench Verified and runs locally.
Ollama v0.31.1 Boosts Local AI Performance on Apple Silicon
Ollama v0.31.1 makes Gemma 4 about 90% faster on Apple Silicon via multi-token prediction, advancing local AI performance and privacy.
Ollama 0.30.8 Widens Local AI Hardware Support and Speeds Up Apple Silicon
Ollama 0.30.8, released June 12, broadens GGUF hardware support through llama.cpp and upgrades its Apple Silicon MLX engine for faster, private local AI.

DiffusionGemma Generates Text 4x Faster With Open Diffusion-Based Decoding
Google DeepMind released DiffusionGemma, an open 26B model that generates text via parallel diffusion decoding, reaching up to 2,000 tokens per second and running locally.

Gemma 4 QAT Lands in Ollama, Cutting Local AI Memory by ~72%
Quantization-aware-trained Gemma 4 weights are now runnable in Ollama, cutting VRAM roughly 72% so a 26B model fits on a 16GB laptop for self-hosted AI.

Cohere's North Mini Code Runs an Open Coding Agent on One GPU
Cohere's North Mini Code is a 30B open-weight coding model under Apache 2.0 that runs on a single H100, scoring 83.2% on SWE-Bench Verified with a 256K context.

Google's DiffusionGemma Brings 4x-Faster Text Diffusion to Local AI
Google DeepMind's DiffusionGemma is a 26B open-weight model that writes text in parallel — topping 1,000 tokens/sec and running locally in just 18GB of VRAM.

Holo3.1 Brings Fast, Private Computer-Use AI Agents to Your Own Machine
H Company's open-weight Holo3.1 agents automate desktop and mobile tasks locally, with sizes from 0.8B to 35B and quantized builds that run on consumer hardware.

ASUS Ascent QN10: First Snapdragon X2 Elite Mini PC Hits 80 TOPS
The ASUS Ascent QN10 is the first Snapdragon X2 Elite mini PC, packing an 18-core Oryon CPU and an 80 TOPS NPU for local AI in a 0.7-liter box.

Qwen3.6 Arrives on Ollama: Run a 35B Agentic Coding AI Locally With 256K Context
Alibaba's Qwen3.6 is now on Ollama — a 35B open-weight model with 256K context, vision support, and thinking preservation built for agentic coding workflows you can run on your own hardware.

Beelink's SER10 MAX Ships With AMD's Ryzen AI 9 HX 470 — Bringing Serious AI Compute to Your Desk
The Beelink SER10 MAX pairs AMD's latest Ryzen AI 9 HX 470 processor with a compact desktop form factor, targeting local AI inference and creative workloads.

Minix Unveils Six New Mini PCs With Up to 180 TOPS of AI Compute
From 36-watt office boxes to 180 TOPS AI workstations, Minix’s 2026 mini PC lineup spans every tier of the local AI computing spectrum.

Hugging Face Acquires ggml.ai, Giving llama.cpp a Permanent Open-Source Home
Hugging Face acquires ggml.ai, bringing llama.cpp and the GGUF model format under its umbrella while keeping everything MIT-licensed and open-source for local AI inference.
