Skip to main content
The Quantum Dispatch
Back to Home
on-device-ai

Articles Tagged “On Device AI”

25 articles found

Cover illustration for Liquid AI d1 Models: Open Decision AI That Answers in 8 ms
AI-Generated|Opinion
AI

Liquid AI d1 Models: Open Decision AI That Answers in 8 ms

Liquid AI's open d1-3B and d1-omni-600M decision models skip token generation and answer in one pass, as fast as 8 ms on an RTX 4090 and 50 ms on Jetson.

Dr. Nova Chen
Dr. Nova Chen★Oct 9, 2026★4 min read
Cover illustration for EmbeddingGemma 2: On-Device Search for Text, Images, Audio
AI-Generated|Opinion
AI

EmbeddingGemma 2: On-Device Search for Text, Images, Audio

EmbeddingGemma 2 is a 740M open embedding model that searches text, images, audio and video on a phone using as little as 191MB of RAM. See how it works.

Dr. Nova Chen
Dr. Nova Chen★Oct 6, 2026★3 min read
Cover illustration for Ternary Bonsai 2 27B: A 5.9GB Model for 16GB Laptops
AI-Generated|Opinion
AI

Ternary Bonsai 2 27B: A 5.9GB Model for 16GB Laptops

Ternary Bonsai 2 27B squeezes Qwen3.8-27B into 5.9GB with 1.76-bit weights and keeps 98.2% of its benchmark score. Here's how to run it locally.

Dr. Nova Chen
Dr. Nova Chen★Sep 23, 2026★4 min read
Cover illustration for Needle 2 Brings 14MB Function-Calling AI to Raspberry Pi 5
AI-Generated|Opinion
Mini Computers

Needle 2 Brings 14MB Function-Calling AI to Raspberry Pi 5

Needle 2 is a 14MB function-calling model that turns plain English into Raspberry Pi 5 actions in about 80ms on the CPU alone, with no AI HAT needed.

Alex Circuit
Alex Circuit★Sep 22, 2026★4 min read
Cover illustration for Perplexity Portable Computer Runs Local AI on Windows
AI-Generated|Opinion
AI

Perplexity Portable Computer Runs Local AI on Windows

Perplexity's on-device agent now runs on Windows PCs with 24GB+ RTX GPUs, keeping the model, harness, orchestrator and scheduler off the cloud.

Dr. Nova Chen
Dr. Nova Chen★Sep 16, 2026★5 min read
Cover illustration for iOS 27 Siri: What the New Assistant Can Actually Do
AI-Generated|Opinion
AI

iOS 27 Siri: What the New Assistant Can Actually Do

Apple shipped iOS 27 on September 14 with a rebuilt Siri, a new System Orchestrator and 250+ changes. Here is what the assistant actually does now.

Dr. Nova Chen
Dr. Nova Chen★Sep 15, 2026★5 min read
Cover illustration for Mac mini M6 Speeds Local LLM Prompts by 4.8x for $899
AI-Generated|Opinion
Mini Computers

Mac mini M6 Speeds Local LLM Prompts by 4.8x for $899

Apple's new Mac mini starts at $899 with the M6 chip, 16GB of unified memory and a claimed 4.8x faster LLM prompt processing in LM Studio than M4.

Alex Circuit
Alex Circuit★Aug 27, 2026★6 min read
Cover illustration for LFM2.5-DSpark Speeds Local AI Inference Up to 3.2x
AI-Generated|Opinion
AI

LFM2.5-DSpark Speeds Local AI Inference Up to 3.2x

Liquid AI released ~300M-parameter draft models for the LFM2.5 family, delivering up to 3.18x GPU throughput and 57% lower function-calling latency.

Dr. Nova Chen
Dr. Nova Chen★Aug 22, 2026★3 min read
Cover illustration for On-Device Piano Model Autocompletes Music on iPhone
AI-Generated|Opinion
AI

On-Device Piano Model Autocompletes Music on iPhone

A 125M-parameter transformer generates about 108 piano notes per second on an iPhone 15, trained on 300 million note events with no cloud call.

Dr. Nova Chen
Dr. Nova Chen★Aug 20, 2026★5 min read
Cover illustration for North Micro Vision Packs Document AI Into 2.4B Params
AI-Generated|Opinion
AI

North Micro Vision Packs Document AI Into 2.4B Params

Cohere Labs released North Micro Vision, a 2.4B Apache 2.0 vision model that reads full-resolution A4 pages and scores 0.921 on DocVQA on local hardware.

Dr. Nova Chen
Dr. Nova Chen★Aug 17, 2026★4 min read
Cover illustration for On-Device Vision AI Reads Screens in 3GB of Memory
AI-Generated|Opinion
AI

On-Device Vision AI Reads Screens in 3GB of Memory

Liquid AI's LFM2.5-VL-3B scores 69.4% average across vision benchmarks and decodes 228 tokens/second on an M5 Max, all inside roughly 3GB of memory.

Dr. Nova Chen
Dr. Nova Chen★Aug 13, 2026★5 min read
Cover illustration for Snapdragon C Runs 67% Faster on Battery Than Intel N250
AI-Generated|Opinion
Mini Computers

Snapdragon C Runs 67% Faster on Battery Than Intel N250

Qualcomm's Snapdragon C beat Intel's N250 by up to 67% in unplugged Cinebench multi-core tests, with up to 2.1x better battery power efficiency.

Alex Circuit
Alex Circuit★Aug 13, 2026★5 min read
Cover illustration for Google Sign Language AI Ships in Gboard on Pixel 11
AI-Generated|Opinion
AI

Google Sign Language AI Ships in Gboard on Pixel 11

Google DeepMind's SL2T model brings sign-language-to-text to Gboard and Live Transcribe on Pixel 11, trained on 100,000+ hours across 50 sign languages.

Dr. Nova Chen
Dr. Nova Chen★Aug 12, 2026★6 min read
Cover illustration for LFM2.5-2.6B Runs Tool-Calling AI Agents in 2.5GB of RAM
AI-Generated|Opinion
AI

LFM2.5-2.6B Runs Tool-Calling AI Agents in 2.5GB of RAM

Liquid AI's LFM2.5-2.6B runs full tool-calling AI agents on a phone or Raspberry Pi, hitting 220 tokens per second in under 2.5GB of memory.

Dr. Nova Chen
Dr. Nova Chen★Aug 9, 2026★5 min read
Cover illustration for Liquid AI Encoders Hit 8K Context on CPU 3.7x Faster
AI-Generated|Opinion
AI

Liquid AI Encoders Hit 8K Context on CPU 3.7x Faster

Liquid AI's new LFM2.5-Encoders run 8,192-token inputs on a plain CPU roughly 3.7x faster than ModernBERT-base, from just 230M parameters and open weights.

Dr. Nova Chen
Dr. Nova Chen★Jul 28, 2026★5 min read
Cover illustration for Moonshine Puts Offline Voice AI on a Pico 2 Chip
AI-Generated|Opinion
Mini Computers

Moonshine Puts Offline Voice AI on a Pico 2 Chip

Moonshine AI fit a full offline voice pipeline — detection, speech-to-text, and neural TTS — onto a Raspberry Pi Pico 2, using just 3.6 MiB of flash.

Alex Circuit
Alex Circuit★Jul 26, 2026★5 min read
Cover illustration for Raspberry Pi AI Projects Book Covers Local LLMs for £9
AI-Generated|Opinion
Mini Computers

Raspberry Pi AI Projects Book Covers Local LLMs for £9

Raspberry Pi Press's new AI Projects book covers local LLMs, vision and speech across Pi 5, Pi Zero 2 W and Pico, at an intro price of £8.99.

Alex Circuit
Alex Circuit★Jul 21, 2026★4 min read
Cover illustration for NVIDIA Cosmos 3 Edge Puts World Models Inside Robots
AI-Generated|Opinion
AI

NVIDIA Cosmos 3 Edge Puts World Models Inside Robots

NVIDIA Cosmos 3 Edge is a 4-billion-parameter world model doing spatial reasoning on Jetson and RTX hardware, adaptable to a robot in about a day.

Dr. Nova Chen
Dr. Nova Chen★Jul 20, 2026★4 min read
Cover illustration for Edge AI Dev Boards With NPUs: A 2026 Buyer's Guide
AI-Generated|Opinion
Mini Computers

Edge AI Dev Boards With NPUs: A 2026 Buyer's Guide

Compare five edge AI dev boards from 4 to 67 TOPS and $70 to $249, and see why NPU toolchain maturity, not the TOPS number, decides what runs.

Alex Circuit
Alex Circuit★Jul 20, 2026★8 min read
Cover illustration for Google's Gemma 4 12B Brings Multimodal AI to a 16GB Laptop
AI-Generated|Opinion
AI

Google's Gemma 4 12B Brings Multimodal AI to a 16GB Laptop

Google DeepMind released Gemma 4 12B on June 3, 2026 — an open multimodal model that reads images and audio and runs on a 16GB laptop, free under Apache 2.0.

Dr. Nova Chen
Dr. Nova Chen★Jun 9, 2026★5 min read
Cover illustration for ASUS Ascent QN10 — First Snapdragon X2 Elite Mini PC Hits 80 TOPS
AI-Generated|Opinion
Mini Computers

ASUS Ascent QN10 — First Snapdragon X2 Elite Mini PC Hits 80 TOPS

ASUS unveiled the Ascent QN10, the first mini PC built on Qualcomm's 18-core Snapdragon X2 Elite — 80 TOPS of on-device AI, up to 32GB LPDDR5, and four-display output.

Alex Circuit
Alex Circuit★Jun 9, 2026★4 min read
Cover illustration for Gemma 4 12B Brings Full Multimodal AI to a 16GB Laptop — Free Under Apache 2.0
AI-Generated|Opinion
AI

Gemma 4 12B Brings Full Multimodal AI to a 16GB Laptop — Free Under Apache 2.0

Google DeepMind released Gemma 4 12B on June 3, 2026 — an open-weight, encoder-free multimodal model with native audio that runs locally on a 16GB consumer laptop.

Dr. Nova Chen
Dr. Nova Chen★Jun 4, 2026★5 min read
Cover illustration for Google Gemma 4 Comes to Android: On-Device AI in 140+ Languages, No Cloud Required
AI-Generated|Opinion
AI

Google Gemma 4 Comes to Android: On-Device AI in 140+ Languages, No Cloud Required

Google's AICore Developer Preview brings Gemma 4 natively to Android devices — offline, privacy-preserving AI inference in over 140 languages that upgrades automatically to Gemini Nano 4.

Dr. Nova Chen
Dr. Nova Chen★Apr 9, 2026★4 min read
Cover illustration for Apple's M5 Pro and Max Chips Fuse Neural Accelerators Into Every GPU Core — Delivering 4x the AI Compute
AI-Generated|Opinion
AI

Apple's M5 Pro and Max Chips Fuse Neural Accelerators Into Every GPU Core — Delivering 4x the AI Compute

Apple's new Fusion Architecture bonds two 3nm dies into a single SoC, embedding dedicated neural accelerators in every GPU core for massive on-device AI gains.

Dr. Nova Chen
Dr. Nova Chen★Mar 6, 2026★5 min read
Cover illustration for Apple Is Rebuilding Siri From the Ground Up With LLM-Powered Conversational Intelligence in iOS 26.4
AI-Generated|Opinion
AI

Apple Is Rebuilding Siri From the Ground Up With LLM-Powered Conversational Intelligence in iOS 26.4

Apple's Siri overhaul replaces the command-based architecture with on-device large language models, bringing natural conversation and app-aware context to one billion iPhones.

Dr. Nova Chen
Dr. Nova Chen★Mar 3, 2026★5 min read