Skip to main content
The Quantum Dispatch
Back to Home
llm-inference

Articles Tagged “LLM Inference”

12 articles found

Cover illustration for IBM Granite 4.2 Brings Open Reasoning Models Local
AI-Generated|Opinion
AI

IBM Granite 4.2 Brings Open Reasoning Models Local

IBM released Granite 4.2 under Apache 2.0 in 3B, 8B and 30B sizes, with a switchable thinking mode, a 512K context window and a 57.00 SWE-bench score.

Dr. Nova Chen
Dr. Nova Chen★Aug 31, 2026★6 min read
Cover illustration for OpenAI Jalapeño Benchmarks: 1.9x Work per Kilowatt
AI-Generated|Opinion
AI

OpenAI Jalapeño Benchmarks: 1.9x Work per Kilowatt

OpenAI published the first Jalapeño inference benchmarks: a 700W part claiming up to 1.9x throughput per kilowatt and 3.6x lower end-to-end latency.

Dr. Nova Chen
Dr. Nova Chen★Aug 28, 2026★6 min read
Cover illustration for Nvidia KV Cache Transfer Skips 7-Second Re-Prefills
AI-Generated|Opinion
AI

Nvidia KV Cache Transfer Skips 7-Second Re-Prefills

Nvidia researchers moved a 32,768-token KV cache between model sizes in 278 milliseconds, replacing a 7-second re-prefill with closed-form linear math.

Dr. Nova Chen
Dr. Nova Chen★Aug 22, 2026★3 min read
Cover illustration for LFM2.5-DSpark Speeds Local AI Inference Up to 3.2x
AI-Generated|Opinion
AI

LFM2.5-DSpark Speeds Local AI Inference Up to 3.2x

Liquid AI released ~300M-parameter draft models for the LFM2.5 family, delivering up to 3.18x GPU throughput and 57% lower function-calling latency.

Dr. Nova Chen
Dr. Nova Chen★Aug 22, 2026★3 min read
Cover illustration for Sparse Mixture of Experts Explained for 2026 Models
AI-Generated|Opinion
AI

Sparse Mixture of Experts Explained for 2026 Models

Why a 2.4T-parameter model can be cheaper than a 70B one: what active parameters mean, how routing works, and what MoE really costs to self-host in 2026.

Dr. Nova Chen
Dr. Nova Chen★Aug 4, 2026★11 min read
Cover illustration for Local LLM Servers Compared: Ollama vs vLLM vs llama.cpp
AI-Generated|Opinion
AI

Local LLM Servers Compared: Ollama vs vLLM vs llama.cpp

Ollama, llama.cpp, vLLM and LM Studio all serve local models. Here's which one fits your hardware, from a 16GB laptop to a four-GPU workstation.

Dr. Nova Chen
Dr. Nova Chen★Aug 4, 2026★10 min read
Cover illustration for OpenAI Cuts Luna Prices 80% After AI Rewrote Its Kernels
AI-Generated|Opinion
AI

OpenAI Cuts Luna Prices 80% After AI Rewrote Its Kernels

OpenAI dropped GPT-5.6 Luna to $0.20 per million input tokens after Sol rewrote its own GPU kernels, cutting end-to-end serving costs by 20%.

Dr. Nova Chen
Dr. Nova Chen★Aug 1, 2026★6 min read
Cover illustration for Liquid AI Encoders Hit 8K Context on CPU 3.7x Faster
AI-Generated|Opinion
AI

Liquid AI Encoders Hit 8K Context on CPU 3.7x Faster

Liquid AI's new LFM2.5-Encoders run 8,192-token inputs on a plain CPU roughly 3.7x faster than ModernBERT-base, from just 230M parameters and open weights.

Dr. Nova Chen
Dr. Nova Chen★Jul 28, 2026★5 min read
Cover illustration for LLM Quantization Guide: GGUF vs AWQ vs MLX in 2026
AI-Generated|Opinion
AI

LLM Quantization Guide: GGUF vs AWQ vs MLX in 2026

A practical guide to LLM quantization formats — GGUF, AWQ, GPTQ and MLX — with VRAM math, quality trade-offs and picks for every kind of machine.

Dr. Nova Chen
Dr. Nova Chen★Jul 28, 2026★9 min read
Cover illustration for OpenAI and Broadcom's "Jalapeño" Chip: Custom Silicon for Inference
AI-Generated|Opinion
AI

OpenAI and Broadcom's "Jalapeño" Chip: Custom Silicon for Inference

OpenAI's first custom chip, Jalapeño, is an inference-focused ASIC co-built with Broadcom, promising better performance-per-watt for large language models.

Dr. Nova Chen
Dr. Nova Chen★Jul 1, 2026★5 min read
Cover illustration for NVIDIA Releases Nemotron Diffusion Language Models — A Single Checkpoint That Generates Text Up to 6.4x Faster
AI-Generated|Opinion
AI

NVIDIA Releases Nemotron Diffusion Language Models — A Single Checkpoint That Generates Text Up to 6.4x Faster

NVIDIA Nemotron Labs released a family of diffusion language models on May 23, 2026 — 3B, 8B, and 14B text models plus an 8B VLM that generate tokens in parallel and refine them, hitting 6.4x speedups via self-speculation.

Dr. Nova Chen
Dr. Nova Chen★May 27, 2026★7 min read
Cover illustration for New Self-Distillation Technique Triples LLM Inference Speed With a Single Model
AI-Generated|Opinion
AI

New Self-Distillation Technique Triples LLM Inference Speed With a Single Model

Researchers achieve 3x faster LLM inference by baking multi-token prediction directly into model weights — no draft model or extra hardware required.

Dr. Nova Chen
Dr. Nova Chen★Feb 26, 2026★3 min read