Articles Tagged “Multimodal AI”
35 articles found

Qwen3.8-27B Runs a 262K-Context Vision Model Locally
Alibaba's Qwen3.8-27B lands under Apache 2.0 with vision, a 262K context, and a 17GB quantization that runs at 15-30 tokens per second on a laptop.

North Micro Vision Packs Document AI Into 2.4B Params
Cohere Labs released North Micro Vision, a 2.4B Apache 2.0 vision model that reads full-resolution A4 pages and scores 0.921 on DocVQA on local hardware.

On-Device Vision AI Reads Screens in 3GB of Memory
Liquid AI's LFM2.5-VL-3B scores 69.4% average across vision benchmarks and decodes 228 tokens/second on an M5 Max, all inside roughly 3GB of memory.

Robot Model Trained on a Million Hours of Video
Dyna Robotics says DYNA-2 lifted manufacturing task success from 20% to 80-90% through pre-training on a million hours of egocentric human video alone.
Google Sign Language AI Ships in Gboard on Pixel 11
Google DeepMind's SL2T model brings sign-language-to-text to Gboard and Live Transcribe on Pixel 11, trained on 100,000+ hours across 50 sign languages.

AMIE Video Consultations Match Doctors in a New Study
Google's AMIE matched or beat 30 primary care physicians across 100 clinical scenarios in the first real-time video consultation study of its kind.

Muse Glimmer Runs a 30B Agentic Model on One Consumer GPU
Meta's Muse Glimmer is a 30B open-weight agentic model that compresses under 20GB at 4-bit, so a single 24GB consumer GPU can run it locally.

NVIDIA Magpie TTS Hits 12 Languages With Open Weights
NVIDIA's 364M-parameter Magpie TTS adds Arabic, Korean, and Brazilian Portuguese, and reaches 32ms time-to-first-audio on a B200 GPU with open weights.

FLUX 3 Video Goes GA With 20-Second Clips and Audio
Black Forest Labs opened FLUX 3 Video to every developer on August 4: 20-second clips with natively generated audio, starting at $0.06 per second.

Suno Watermarks AI Songs to Make Origins Verifiable
Suno will embed an inaudible signature in every track it generates, giving streaming platforms a way to identify AI-made music automatically.

MiniMax H3 Makes 2K Video With Native Stereo Audio
MiniMax H3 generates 15-second 2K clips with native stereo sound and tops the video editing leaderboard at 1,130 Elo, priced at 0.8 yuan per second.

Qwen3.7 Flash Brings 1M-Token Vision at $0.03 per Million
Alibaba's Qwen3.7 Flash is a native vision-language model with a 1M-token context window, priced at $0.03 per million input tokens with tool calling.

Newsroom AI Workflows That Give Reporters Time Back
OpenAI detailed on July 27 how Business Insider, WELT, and Le Monde use AI for one-tap listening, fact-checking layers, and faster translation.

Inkling Is a 975B Open-Weights Model Under Apache 2.0
Thinking Machines released Inkling, a 975B-parameter Apache 2.0 model with 41B active, a 1M-token context window, and native four-modality reasoning.

Qwen3.8-Max Benchmarks: What to Watch in the Preview
Alibaba's Qwen3.8-Max preview brings 2.4 trillion parameters and multimodal input, but no benchmark table yet. Here's the baseline to measure it against.

MIT's GIFT Turns 2D Designs Into CAD at 20% Compute
MIT's GIFT system converts a 2D image and text into executable CAD code using about 20% of the compute rival methods need, with no human labeling.

Gemini Notebook Replaces NotebookLM and Runs Your Code
Google renamed NotebookLM to Gemini Notebook and added secure cloud code execution, serving 30 million users and over 600,000 organizations.

Kimi K3 Becomes the Largest Open-Weight AI Model Yet
Moonshot AI's Kimi K3 is a 2.8-trillion-parameter open-weight model that ranks 3rd on GDPval-AA v2, with full weights arriving July 27.

Thinking Machines Inkling: A 975B Open-Weights Model
Thinking Machines Lab released Inkling, a 975B-parameter open-weights multimodal model with 41B active per token and a 1M-token context window.

MiniMax M3 Open-Weight Model Lands With 1M Context and Native Multimodal
MiniMax published the open weights and technical report for M3, an open-weight model pairing a 1M-token context window with native image and video understanding.

Google's Gemma 4 12B Brings Multimodal AI to a 16GB Laptop
Google DeepMind released Gemma 4 12B on June 3, 2026 — an open multimodal model that reads images and audio and runs on a 16GB laptop, free under Apache 2.0.

Gemma 4 12B Brings Full Multimodal AI to a 16GB Laptop — Free Under Apache 2.0
Google DeepMind released Gemma 4 12B on June 3, 2026 — an open-weight, encoder-free multimodal model with native audio that runs locally on a 16GB consumer laptop.

Alibaba's Qwen3.7-Plus Pairs Vision With Autonomous Agent Skills at $0.40 per Million Tokens
Alibaba's Qwen team launched Qwen3.7-Plus on June 2, 2026 — a multimodal model combining vision and video understanding with deep reasoning, tool use, and autonomous iteration.

NVIDIA's Nemotron 3 Nano Omni Lands — A 30B Open Omni-Modal Reasoning Model With 9x Higher Throughput
NVIDIA released Nemotron 3 Nano Omni on May 13, 2026 — a 30B-A3B open omni-modal model that unifies text, image, video, and audio reasoning with 9x higher throughput than other open omni models.

Google Makes Gemini API File Search Multimodal With Page-Level Citations
Google's Gemini API File Search now indexes images alongside text, ties responses to original page numbers, and ships free storage and query embeddings — turning verifiable RAG into a one-call developer primitive.

Mistral Small 4 Unifies Reasoning, Multimodal, and Coding Into One Apache 2.0 Model
Mistral Small 4 collapses three flagship model families — reasoning, multimodal, and agentic coding — into a single 119B-parameter Apache 2.0 model with a 256k context window.

OpenAI's 'Spud' Completes Pretraining — The Next Frontier Model Is Almost Here
OpenAI confirmed its next frontier model, codenamed Spud, finished pretraining on March 24 — a unified multimodal AI expected to arrive within weeks.

Meta Muse Spark Launches From Superintelligence Labs: Personal AI Gets Its Biggest Upgrade Yet
Meta Superintelligence Labs launched Muse Spark on April 9 — a natively multimodal reasoning model now powering personal AI across all Meta platforms.

LG Releases EXAONE 4.5: Open-Source Vision-Language AI That Outscores GPT-5-mini
LG AI Research's EXAONE 4.5 is a 33B multimodal VLM with Hybrid Attention architecture that outscores GPT-5-mini and Claude 4.5 Sonnet on STEM benchmarks — and it's fully open-source.

Meta Launches Llama 4 Scout and Maverick: Multimodal MoE AI Goes Open-Weight
Meta's Llama 4 Scout and Maverick bring multimodal mixture-of-experts AI to the open-source community, with an unprecedented 10 million token context window.

Google Gemma 4 Launches With Four Sizes, Apache 2.0 License, and a Top-3 Open Model Ranking
Google's Gemma 4 arrives with model sizes from 2B to 31B, a permissive Apache 2.0 license, native multimodal support across all sizes, and the #3 spot on the global open model leaderboard.

Google Launches Gemini Embedding 2 — The First AI Model That Maps Text, Images, and Video Into a Single Search Space
Google's new natively multimodal embedding model jointly maps text, images, and video into a unified vector space, enabling cross-modal retrieval and RAG applications.

OpenAI Is Bringing Sora Video Generation Directly Into ChatGPT — Giving Hundreds of Millions of Users Access
According to The Information, OpenAI plans to embed Sora's video-generation capabilities into the ChatGPT interface, mirroring how DALL-E image creation was integrated.

DeepSeek Unveils V4 — A Trillion-Parameter Multimodal Model That Generates Text, Images, and Video
DeepSeek's V4 model enters the frontier tier with trillion-parameter multimodal capabilities spanning text, image, and video generation plus elite coding performance.

Google Gemini 3.1 Pro Doubles Reasoning Performance With a New Three-Tier Thinking System
Google DeepMind’s Gemini 3.1 Pro scores 77.1% on ARC-AGI-2, more than doubling its predecessor’s reasoning with a three-tier thinking architecture.
