Artificial Intelligence — Page 2
331 articles in this category

North Micro Vision Packs Document AI Into 2.4B Params
Cohere Labs released North Micro Vision, a 2.4B Apache 2.0 vision model that reads full-resolution A4 pages and scores 0.921 on DocVQA on local hardware.

Wispr Raises $280M and Ships Its Canto Speech Model
Wispr closed a $280M Series B at a $2B valuation and launched Canto, an in-house speech model it says cuts dictation error rates from 30% to under 10%.

Stripe Buys OpenRouter in a $7 Billion AI Gateway Deal
Stripe has agreed to acquire OpenRouter for more than $7 billion, more than five times the AI gateway startup's $1.3 billion valuation in May.

Kog Inference Engine Squeezes 30x From Existing GPUs
French startup Kog hit 3,000 tokens per second on standard AMD and NVIDIA datacenter GPUs, claiming a path to 30x faster LLM inference in software.

AI Agents Reproduced 2,226 ICML Papers in Just 19 Days
A Hugging Face community challenge used AI coding agents to audit 2,226 ICML 2026 papers, verifying claims across 6,816 public reproduction logbooks.

Gemma Downloads Top 900 Million Across Open Models
Google's Gemma open models have now passed 900 million downloads, with Gemma 4 alone contributing over 300 million since its April 2026 launch.

IBM and OpenAI Partner on Secure Enterprise AI Rollouts
IBM Consulting will embed GPT-5.6, Codex, and ChatGPT Work into its platform, backed by a new OpenAI Practice of thousands of certified staff.

Gemini 3.7 Flash Benchmarks: What Developers Get Now
Gemini 3.7 Flash arrived August 13 with DeepSWE v1.1 at 65.3%, WebDev Arena Elo of 1588, and intro pricing of $0.75 per million input tokens.

Claude Text Watermarking: What App Builders Should Know
Anthropic is embedding invisible watermarks in Claude text output and C2PA metadata in image files, with a detection API confirmed as on the roadmap.

Qwen3.8-Max Open Weights Make a 2.4T Model Downloadable
Alibaba published Qwen3.8-Max open weights on August 12: 2.4 trillion parameters, 95B active per token, and a 262K context that extends past 1M tokens.

On-Device Vision AI Reads Screens in 3GB of Memory
Liquid AI's LFM2.5-VL-3B scores 69.4% average across vision benchmarks and decodes 228 tokens/second on an M5 Max, all inside roughly 3GB of memory.

Robot Model Trained on a Million Hours of Video
Dyna Robotics says DYNA-2 lifted manufacturing task success from 20% to 80-90% through pre-training on a million hours of egocentric human video alone.

Grok 4.6 Brings a 500K Context Window to AI Agents
xAI's Grok 4.6 ships a 500K-token context window and scores 61 on the Artificial Analysis index, holding Grok 4.5's $2 per million input token price.

Nemotron 3.5 Lightning Puts 1M-Token Agents on RTX PCs
NVIDIA's Nemotron 3.5 Lightning is a 30B open model with 3B active parameters, a 1M-token context, and 4x faster output for local AI agents.

ChatGPT Restaurant Booking Arrives via Yelp and Resy
ChatGPT can now book restaurant tables through OpenTable, Resy, and Yelp, with waitlist joins across thousands of venues in the US and Canada.

Claude Code Auto Mode Turns On by Default August 14
Anthropic makes Claude Code auto mode the default for Pro, Max, and Team on August 14, after a 1,053-person study found it caught 89% of harmful actions.
Google Sign Language AI Ships in Gboard on Pixel 11
Google DeepMind's SL2T model brings sign-language-to-text to Gboard and Live Transcribe on Pixel 11, trained on 100,000+ hours across 50 sign languages.

AMIE Video Consultations Match Doctors in a New Study
Google's AMIE matched or beat 30 primary care physicians across 100 clinical scenarios in the first real-time video consultation study of its kind.

Novo Nordisk and AWS Put AI Agents on Drug Discovery
Novo Nordisk and AWS opened a London co-innovation hub to compress the path from drug target to first human dose, with AI already deployed to 25,000+ staff.

Muse Glimmer Runs a 30B Agentic Model on One Consumer GPU
Meta's Muse Glimmer is a 30B open-weight agentic model that compresses under 20GB at 4-bit, so a single 24GB consumer GPU can run it locally.

NVIDIA Magpie TTS Hits 12 Languages With Open Weights
NVIDIA's 364M-parameter Magpie TTS adds Arabic, Korean, and Brazilian Portuguese, and reaches 32ms time-to-first-audio on a B200 GPU with open weights.

FLUX 3 Video Goes GA With 20-Second Clips and Audio
Black Forest Labs opened FLUX 3 Video to every developer on August 4: 20-second clips with natively generated audio, starting at $0.06 per second.

Suno Watermarks AI Songs to Make Origins Verifiable
Suno will embed an inaudible signature in every track it generates, giving streaming platforms a way to identify AI-made music automatically.

Firmus Raises $2B for Renewable-Powered AI Factories
Firmus closed a fully subscribed $2 billion round at a $10.5 billion valuation to build renewable-powered AI data centers across Asia-Pacific.
