Skip to main content
The Quantum Dispatch
Back to Home
multimodal-ai

Articles Tagged “Multimodal AI

35 articles found

Cover illustration for Qwen3.8-27B Runs a 262K-Context Vision Model Locally
AI-Generated|Opinion
AI

Qwen3.8-27B Runs a 262K-Context Vision Model Locally

Alibaba's Qwen3.8-27B lands under Apache 2.0 with vision, a 262K context, and a 17GB quantization that runs at 15-30 tokens per second on a laptop.

Dr. Nova Chen
Dr. Nova ChenAug 17, 20264 min read
Cover illustration for North Micro Vision Packs Document AI Into 2.4B Params
AI-Generated|Opinion
AI

North Micro Vision Packs Document AI Into 2.4B Params

Cohere Labs released North Micro Vision, a 2.4B Apache 2.0 vision model that reads full-resolution A4 pages and scores 0.921 on DocVQA on local hardware.

Dr. Nova Chen
Dr. Nova ChenAug 17, 20264 min read
Cover illustration for On-Device Vision AI Reads Screens in 3GB of Memory
AI-Generated|Opinion
AI

On-Device Vision AI Reads Screens in 3GB of Memory

Liquid AI's LFM2.5-VL-3B scores 69.4% average across vision benchmarks and decodes 228 tokens/second on an M5 Max, all inside roughly 3GB of memory.

Dr. Nova Chen
Dr. Nova ChenAug 13, 20265 min read
Cover illustration for Robot Model Trained on a Million Hours of Video
AI-Generated|Opinion
AI

Robot Model Trained on a Million Hours of Video

Dyna Robotics says DYNA-2 lifted manufacturing task success from 20% to 80-90% through pre-training on a million hours of egocentric human video alone.

Dr. Nova Chen
Dr. Nova ChenAug 13, 20265 min read
Cover illustration for Google Sign Language AI Ships in Gboard on Pixel 11
AI-Generated|Opinion
AI

Google Sign Language AI Ships in Gboard on Pixel 11

Google DeepMind's SL2T model brings sign-language-to-text to Gboard and Live Transcribe on Pixel 11, trained on 100,000+ hours across 50 sign languages.

Dr. Nova Chen
Dr. Nova ChenAug 12, 20266 min read
Cover illustration for AMIE Video Consultations Match Doctors in a New Study
AI-Generated|Opinion
AI

AMIE Video Consultations Match Doctors in a New Study

Google's AMIE matched or beat 30 primary care physicians across 100 clinical scenarios in the first real-time video consultation study of its kind.

Dr. Nova Chen
Dr. Nova ChenAug 12, 20266 min read
Cover illustration for Muse Glimmer Runs a 30B Agentic Model on One Consumer GPU
AI-Generated|Opinion
AI

Muse Glimmer Runs a 30B Agentic Model on One Consumer GPU

Meta's Muse Glimmer is a 30B open-weight agentic model that compresses under 20GB at 4-bit, so a single 24GB consumer GPU can run it locally.

Dr. Nova Chen
Dr. Nova ChenAug 11, 20266 min read
Cover illustration for NVIDIA Magpie TTS Hits 12 Languages With Open Weights
AI-Generated|Opinion
AI

NVIDIA Magpie TTS Hits 12 Languages With Open Weights

NVIDIA's 364M-parameter Magpie TTS adds Arabic, Korean, and Brazilian Portuguese, and reaches 32ms time-to-first-audio on a B200 GPU with open weights.

Dr. Nova Chen
Dr. Nova ChenAug 11, 20265 min read
Cover illustration for FLUX 3 Video Goes GA With 20-Second Clips and Audio
AI-Generated|Opinion
AI

FLUX 3 Video Goes GA With 20-Second Clips and Audio

Black Forest Labs opened FLUX 3 Video to every developer on August 4: 20-second clips with natively generated audio, starting at $0.06 per second.

Dr. Nova Chen
Dr. Nova ChenAug 10, 20264 min read
Cover illustration for Suno Watermarks AI Songs to Make Origins Verifiable
AI-Generated|Opinion
AI

Suno Watermarks AI Songs to Make Origins Verifiable

Suno will embed an inaudible signature in every track it generates, giving streaming platforms a way to identify AI-made music automatically.

Dr. Nova Chen
Dr. Nova ChenAug 10, 20264 min read
Cover illustration for MiniMax H3 Makes 2K Video With Native Stereo Audio
AI-Generated|Opinion
AI

MiniMax H3 Makes 2K Video With Native Stereo Audio

MiniMax H3 generates 15-second 2K clips with native stereo sound and tops the video editing leaderboard at 1,130 Elo, priced at 0.8 yuan per second.

Dr. Nova Chen
Dr. Nova ChenAug 4, 20266 min read
Cover illustration for Qwen3.7 Flash Brings 1M-Token Vision at $0.03 per Million
AI-Generated|Opinion
AI

Qwen3.7 Flash Brings 1M-Token Vision at $0.03 per Million

Alibaba's Qwen3.7 Flash is a native vision-language model with a 1M-token context window, priced at $0.03 per million input tokens with tool calling.

Dr. Nova Chen
Dr. Nova ChenAug 1, 20265 min read
Cover illustration for Newsroom AI Workflows That Give Reporters Time Back
AI-Generated|Opinion
AI

Newsroom AI Workflows That Give Reporters Time Back

OpenAI detailed on July 27 how Business Insider, WELT, and Le Monde use AI for one-tap listening, fact-checking layers, and faster translation.

Dr. Nova Chen
Dr. Nova ChenJul 27, 20265 min read
Cover illustration for Inkling Is a 975B Open-Weights Model Under Apache 2.0
AI-Generated|Opinion
AI

Inkling Is a 975B Open-Weights Model Under Apache 2.0

Thinking Machines released Inkling, a 975B-parameter Apache 2.0 model with 41B active, a 1M-token context window, and native four-modality reasoning.

Dr. Nova Chen
Dr. Nova ChenJul 21, 20265 min read
Cover illustration for Qwen3.8-Max Benchmarks: What to Watch in the Preview
AI-Generated|Opinion
AI

Qwen3.8-Max Benchmarks: What to Watch in the Preview

Alibaba's Qwen3.8-Max preview brings 2.4 trillion parameters and multimodal input, but no benchmark table yet. Here's the baseline to measure it against.

Dr. Nova Chen
Dr. Nova ChenJul 20, 20264 min read
Cover illustration for MIT's GIFT Turns 2D Designs Into CAD at 20% Compute
AI-Generated|Opinion
AI

MIT's GIFT Turns 2D Designs Into CAD at 20% Compute

MIT's GIFT system converts a 2D image and text into executable CAD code using about 20% of the compute rival methods need, with no human labeling.

Dr. Nova Chen
Dr. Nova ChenJul 19, 20264 min read
Cover illustration for Gemini Notebook Replaces NotebookLM and Runs Your Code
AI-Generated|Opinion
AI

Gemini Notebook Replaces NotebookLM and Runs Your Code

Google renamed NotebookLM to Gemini Notebook and added secure cloud code execution, serving 30 million users and over 600,000 organizations.

Dr. Nova Chen
Dr. Nova ChenJul 19, 20263 min read
Cover illustration for Kimi K3 Becomes the Largest Open-Weight AI Model Yet
AI-Generated|Opinion
AI

Kimi K3 Becomes the Largest Open-Weight AI Model Yet

Moonshot AI's Kimi K3 is a 2.8-trillion-parameter open-weight model that ranks 3rd on GDPval-AA v2, with full weights arriving July 27.

Dr. Nova Chen
Dr. Nova ChenJul 18, 20265 min read
Cover illustration for Thinking Machines Inkling: A 975B Open-Weights Model
AI-Generated|Opinion
AI

Thinking Machines Inkling: A 975B Open-Weights Model

Thinking Machines Lab released Inkling, a 975B-parameter open-weights multimodal model with 41B active per token and a 1M-token context window.

Dr. Nova Chen
Dr. Nova ChenJul 16, 20265 min read
Cover illustration for MiniMax M3 Open-Weight Model Lands With 1M Context and Native Multimodal
AI-Generated|Opinion
AI

MiniMax M3 Open-Weight Model Lands With 1M Context and Native Multimodal

MiniMax published the open weights and technical report for M3, an open-weight model pairing a 1M-token context window with native image and video understanding.

Dr. Nova Chen
Dr. Nova ChenJun 15, 20266 min read
Cover illustration for Google's Gemma 4 12B Brings Multimodal AI to a 16GB Laptop
AI-Generated|Opinion
AI

Google's Gemma 4 12B Brings Multimodal AI to a 16GB Laptop

Google DeepMind released Gemma 4 12B on June 3, 2026 — an open multimodal model that reads images and audio and runs on a 16GB laptop, free under Apache 2.0.

Dr. Nova Chen
Dr. Nova ChenJun 9, 20265 min read
Cover illustration for Gemma 4 12B Brings Full Multimodal AI to a 16GB Laptop — Free Under Apache 2.0
AI-Generated|Opinion
AI

Gemma 4 12B Brings Full Multimodal AI to a 16GB Laptop — Free Under Apache 2.0

Google DeepMind released Gemma 4 12B on June 3, 2026 — an open-weight, encoder-free multimodal model with native audio that runs locally on a 16GB consumer laptop.

Dr. Nova Chen
Dr. Nova ChenJun 4, 20265 min read
Cover illustration for Alibaba's Qwen3.7-Plus Pairs Vision With Autonomous Agent Skills at $0.40 per Million Tokens
AI-Generated|Opinion
AI

Alibaba's Qwen3.7-Plus Pairs Vision With Autonomous Agent Skills at $0.40 per Million Tokens

Alibaba's Qwen team launched Qwen3.7-Plus on June 2, 2026 — a multimodal model combining vision and video understanding with deep reasoning, tool use, and autonomous iteration.

Dr. Nova Chen
Dr. Nova ChenJun 4, 20265 min read
Cover illustration for NVIDIA's Nemotron 3 Nano Omni Lands — A 30B Open Omni-Modal Reasoning Model With 9x Higher Throughput
AI-Generated|Opinion
AI

NVIDIA's Nemotron 3 Nano Omni Lands — A 30B Open Omni-Modal Reasoning Model With 9x Higher Throughput

NVIDIA released Nemotron 3 Nano Omni on May 13, 2026 — a 30B-A3B open omni-modal model that unifies text, image, video, and audio reasoning with 9x higher throughput than other open omni models.

Dr. Nova Chen
Dr. Nova ChenMay 14, 20267 min read
Cover illustration for Google Makes Gemini API File Search Multimodal With Page-Level Citations
AI-Generated|Opinion
AI

Google Makes Gemini API File Search Multimodal With Page-Level Citations

Google's Gemini API File Search now indexes images alongside text, ties responses to original page numbers, and ships free storage and query embeddings — turning verifiable RAG into a one-call developer primitive.

Dr. Nova Chen
Dr. Nova ChenMay 7, 20265 min read
Cover illustration for Mistral Small 4 Unifies Reasoning, Multimodal, and Coding Into One Apache 2.0 Model
AI-Generated|Opinion
AI

Mistral Small 4 Unifies Reasoning, Multimodal, and Coding Into One Apache 2.0 Model

Mistral Small 4 collapses three flagship model families — reasoning, multimodal, and agentic coding — into a single 119B-parameter Apache 2.0 model with a 256k context window.

Dr. Nova Chen
Dr. Nova ChenApr 26, 20266 min read
Cover illustration for OpenAI's 'Spud' Completes Pretraining — The Next Frontier Model Is Almost Here
AI-Generated|Opinion
AI

OpenAI's 'Spud' Completes Pretraining — The Next Frontier Model Is Almost Here

OpenAI confirmed its next frontier model, codenamed Spud, finished pretraining on March 24 — a unified multimodal AI expected to arrive within weeks.

Dr. Nova Chen
Dr. Nova ChenApr 12, 20265 min read
Cover illustration for Meta Muse Spark Launches From Superintelligence Labs: Personal AI Gets Its Biggest Upgrade Yet
AI-Generated|Opinion
AI

Meta Muse Spark Launches From Superintelligence Labs: Personal AI Gets Its Biggest Upgrade Yet

Meta Superintelligence Labs launched Muse Spark on April 9 — a natively multimodal reasoning model now powering personal AI across all Meta platforms.

Dr. Nova Chen
Dr. Nova ChenApr 10, 20265 min read
Cover illustration for LG Releases EXAONE 4.5: Open-Source Vision-Language AI That Outscores GPT-5-mini
AI-Generated|Opinion
AI

LG Releases EXAONE 4.5: Open-Source Vision-Language AI That Outscores GPT-5-mini

LG AI Research's EXAONE 4.5 is a 33B multimodal VLM with Hybrid Attention architecture that outscores GPT-5-mini and Claude 4.5 Sonnet on STEM benchmarks — and it's fully open-source.

Dr. Nova Chen
Dr. Nova ChenApr 9, 20265 min read
Cover illustration for Meta Launches Llama 4 Scout and Maverick: Multimodal MoE AI Goes Open-Weight
AI-Generated|Opinion
AI

Meta Launches Llama 4 Scout and Maverick: Multimodal MoE AI Goes Open-Weight

Meta's Llama 4 Scout and Maverick bring multimodal mixture-of-experts AI to the open-source community, with an unprecedented 10 million token context window.

Dr. Nova Chen
Dr. Nova ChenApr 9, 20265 min read
Cover illustration for Google Gemma 4 Launches With Four Sizes, Apache 2.0 License, and a Top-3 Open Model Ranking
AI-Generated|Opinion
AI

Google Gemma 4 Launches With Four Sizes, Apache 2.0 License, and a Top-3 Open Model Ranking

Google's Gemma 4 arrives with model sizes from 2B to 31B, a permissive Apache 2.0 license, native multimodal support across all sizes, and the #3 spot on the global open model leaderboard.

Dr. Nova Chen
Dr. Nova ChenApr 4, 20265 min read
Cover illustration for Google Launches Gemini Embedding 2 — The First AI Model That Maps Text, Images, and Video Into a Single Search Space
AI-Generated|Opinion
AI

Google Launches Gemini Embedding 2 — The First AI Model That Maps Text, Images, and Video Into a Single Search Space

Google's new natively multimodal embedding model jointly maps text, images, and video into a unified vector space, enabling cross-modal retrieval and RAG applications.

Dr. Nova Chen
Dr. Nova ChenMar 14, 20264 min read
Cover illustration for OpenAI Is Bringing Sora Video Generation Directly Into ChatGPT — Giving Hundreds of Millions of Users Access
AI-Generated|Opinion
AI

OpenAI Is Bringing Sora Video Generation Directly Into ChatGPT — Giving Hundreds of Millions of Users Access

According to The Information, OpenAI plans to embed Sora's video-generation capabilities into the ChatGPT interface, mirroring how DALL-E image creation was integrated.

Dr. Nova Chen
Dr. Nova ChenMar 12, 20264 min read
Cover illustration for DeepSeek Unveils V4 — A Trillion-Parameter Multimodal Model That Generates Text, Images, and Video
AI-Generated|Opinion
AI

DeepSeek Unveils V4 — A Trillion-Parameter Multimodal Model That Generates Text, Images, and Video

DeepSeek's V4 model enters the frontier tier with trillion-parameter multimodal capabilities spanning text, image, and video generation plus elite coding performance.

Dr. Nova Chen
Dr. Nova ChenMar 4, 20265 min read
Cover illustration for Google Gemini 3.1 Pro Doubles Reasoning Performance With a New Three-Tier Thinking System
AI-Generated|Opinion
AI

Google Gemini 3.1 Pro Doubles Reasoning Performance With a New Three-Tier Thinking System

Google DeepMind’s Gemini 3.1 Pro scores 77.1% on ARC-AGI-2, more than doubling its predecessor’s reasoning with a three-tier thinking architecture.

Dr. Nova Chen
Dr. Nova ChenFeb 25, 20265 min read