Artificial Intelligence
The latest breakthroughs in AI, machine learning, and neural networks.
331 articles

Intel Diamond Rapids Packs 256 Cores for Agentic AI
Intel's Hot Chips 2026 lineup pairs a 256-core Diamond Rapids Xeon with Crescent Island, a 350W inference GPU holding up to 480GB of LPDDR5X.

Nvidia Vera Rubin NVL72 Targets 30x Tokens Per Watt
Nvidia detailed the Vera Rubin NVL72 rack at Hot Chips 2026: 72 GPUs, 2 ZFLOPS of NVFP4 inference and up to 30x more tokens per megawatt.

Edge AI Model Zoo Logs 837 Reproducible NPU Tests
EdgeFirst's new model zoo publishes 837 validation sessions for YOLO models across NXP, Hailo, Jetson, Qualcomm and Apple Neural Engine hardware.

IBM Dual-Architecture Chip Runs Arm and Z on One Core
IBM's new mainframe processor packs 11 cores above 5.7 GHz on 2nm silicon, and every core runs both z/Architecture and Arm code natively - no emulation.

DeepMind Puts SIMA 2 AI Agents to Work in EVE Online
Google DeepMind and Fenris Creations will test SIMA 2 agents in offline EVE Online copies, extending 15 years of AI games research to a live universe.

Agent Memory Tuning Adds 16 Points for 5% More Tokens
IBM Research's ALTK-Evolve lifted agent task completion 16.1 points at only 5% more tokens, showing memory is a dose you calibrate per model.

GPT-5.6 Sol API Price Falls to $4 Per Million Tokens
OpenAI cut GPT-5.6 Sol API pricing on August 21: input drops 20% to $4 and output falls 33% to $20 per million tokens through November 21, 2026.

Speech Recognition Benchmarks Get a Three-Test Audit
A Hume AI study of 11 open speech models introduces three diagnostics that separate genuine transcription skill from memorized benchmark patterns.

Nvidia KV Cache Transfer Skips 7-Second Re-Prefills
Nvidia researchers moved a 32,768-token KV cache between model sizes in 278 milliseconds, replacing a 7-second re-prefill with closed-form linear math.

LFM2.5-DSpark Speeds Local AI Inference Up to 3.2x
Liquid AI released ~300M-parameter draft models for the LFM2.5 family, delivering up to 3.18x GPU throughput and 57% lower function-calling latency.

TrueForge Open Source Agent Harness Cuts Costs 75%
TrueFoundry's MIT-licensed TrueForge harness completed Enterprise-Bench tasks for $2.90 against $11.80, a 75% cost cut driven by context engineering.

ChatGPT for Teens Adds Study Mode and Quiet Hours
OpenAI's ChatGPT for Teens launched for ages 13-17 with Study Mode, scheduled Study Hours, parental Quiet Hours, and stronger content limits.

Apple Music Will Show Made With AI Labels This Year
Apple Music will make its AI Transparency Tags mandatory and surface Made With AI labels to listeners later in 2026, shifting from removal to disclosure.

ChatGPT Can Now Draft Apple Messages on Your Mac
OpenAI's macOS app gained an Apple Messages plugin that searches, summarizes, and drafts iMessage, SMS, and RCS threads locally on Apple Silicon.

Mojo 1.0 Compiler Goes Open Source Under Apache 2.0
Modular open-sourced the entire Mojo compiler and toolchain under Apache 2.0 at ModCon 2026, alongside 450,000 lines of GPU kernel code.

On-Device Piano Model Autocompletes Music on iPhone
A 125M-parameter transformer generates about 108 piano notes per second on an iPhone 15, trained on 300 million note events with no cloud call.

Cerebras CS-4 Packs Three Wafer-Scale Chips in a Rack
The Cerebras CS-4 puts three WSE-3 Turbo wafers in one rack for 750 PFLOPS, and Cerebras clocks 4,400 tokens per second per user on GPT-OSS-120B.

Ornith-1.5 Open Weights Score 86.1 on Terminal-Bench
Ornith-1.5 ships MIT-licensed weights from 9B to 397B, and the flagship posts 86.1 on Terminal-Bench 2.1 while a 35B MoE sibling runs far leaner.

Palomar Registry Puts AI Math Proofs Under Lean Review
Palomar, incubated by the Lean FRO and ICARM, is now open for submissions and gives machine-checked math proofs a public registry with real metadata.

Etched Raises $700M at a $21B AI Chip Valuation
Etched closed a $700M Series D led by Jane Street at a $21 billion valuation, doubling in just a month as inference chip orders pass $1 billion.

Anthropic Revenue Reaches a $65B Annual Run Rate
Anthropic's annualized revenue hit $65 billion at the end of July 2026, up from $47 billion in May — an $18 billion jump in roughly two months.

Embedding Models Guide: Dense vs Multi-Vector RAG
A practical 2026 guide to choosing embedding models for RAG: dense bi-encoders, multi-vector late interaction, and rerankers, with cost and memory math.

Grok Bot Beta Gives AI Agents a Real Cloud Computer
SpaceXAI's Grok Bot beta gives AI agents a persistent cloud computer and real app logins across 3 tiers, with approval gates on purchases and deletions.

Qwen3.8-27B Runs a 262K-Context Vision Model Locally
Alibaba's Qwen3.8-27B lands under Apache 2.0 with vision, a 262K context, and a 17GB quantization that runs at 15-30 tokens per second on a laptop.
