Artificial Intelligence — Page 3
331 articles in this category

ChatGPT Reasoning Slider Puts Thinking Effort in Your Hands
ChatGPT's new reasoning slider spans five effort levels, and the updated GPT-5.6 Sol makes factual errors 68% less often than GPT-5.5 Instant.

LFM2.5-2.6B Runs Tool-Calling AI Agents in 2.5GB of RAM
Liquid AI's LFM2.5-2.6B runs full tool-calling AI agents on a phone or Raspberry Pi, hitting 220 tokens per second in under 2.5GB of memory.

NVIDIA NOOA Turns an AI Agent Into One Python Class
NVIDIA open-sourced NOOA, an agent framework where a 253-line agent hits 82.2% on SWE-bench Verified using half the tokens of rival harnesses.

Meta Muse Code Pairs a Terminal Agent With Muse Spark 1.2
Meta shipped Muse Code, a terminal coding agent co-trained with Muse Spark 1.2, a model scoring 54 on the Artificial Analysis Intelligence Index.

Fable 5 Biology Safeguards Cut False Positives by 85%
Anthropic retuned Fable 5's biology classifier, cutting biology fallbacks roughly 85% and total fallbacks 67% on Claude.ai while keeping dual-use limits.

WeatherNext Cyclones Adds a Full Day of Forecast Lead
Google DeepMind published WeatherNext Cyclones in Nature and open-sourced the weights under Apache 2.0, adding over 24 hours of cyclone forecast lead time.

K-EXAONE 2.0 Ships 750B Open Weights Under Apache 2.0
LG AI Research released K-EXAONE 2.0, a 750-billion-parameter mixture-of-experts model with 37B active parameters, under a permissive Apache 2.0 license.

OpenAI Astra Proves 10 Open Math Problems in Lean 4
OpenAI published machine-checkable Lean 4 proofs for ten long-open math problems from an internal Astra model, at a total compute cost of about $2,000.

Cloudflare OS Open-Sources an Agent Workspace Platform
Cloudflare released Cloudflare OS as open source — an agent workspace with isolated code runtimes, a governance layer, and user-modifiable internal apps.

Qwen3.8-Max Packs 2.4T Parameters Into a 1M Context
Alibaba's Qwen3.8-Max is a 2.4-trillion-parameter sparse MoE model with a 1M-token window, priced at $2 per million input tokens and open weights next week.

NVIDIA Alpamayo 2 Super Ships an Open 34B AV Model
NVIDIA released Alpamayo 2 Super for commercial use — a 34B open reasoning model for robotaxis with 360-degree perception and a permissive OpenMDW license.

Sparse Mixture of Experts Explained for 2026 Models
Why a 2.4T-parameter model can be cheaper than a 70B one: what active parameters mean, how routing works, and what MoE really costs to self-host in 2026.

DeepSeek V4-Flash 0731 Tops V4-Pro at a Third the Price
DeepSeek's retrained V4-Flash 0731 beats its own V4-Pro preview on every published agentic benchmark at $0.28 per million output tokens, MIT licensed.

Local LLM Servers Compared: Ollama vs vLLM vs llama.cpp
Ollama, llama.cpp, vLLM and LM Studio all serve local models. Here's which one fits your hardware, from a 16GB laptop to a four-GPU workstation.

MiniMax H3 Makes 2K Video With Native Stereo Audio
MiniMax H3 generates 15-second 2K clips with native stereo sound and tops the video editing leaderboard at 1,130 Elo, priced at 0.8 yuan per second.

Oracle Puts Gemini Models Inside Fusion AI Agents
Oracle is bringing Google's Gemini models into Fusion Applications AI Agent Studio, letting thousands of enterprise customers build agents on Gemini.

OpenAI Cuts Luna Prices 80% After AI Rewrote Its Kernels
OpenAI dropped GPT-5.6 Luna to $0.20 per million input tokens after Sol rewrote its own GPU kernels, cutting end-to-end serving costs by 20%.

Qwen3.7 Flash Brings 1M-Token Vision at $0.03 per Million
Alibaba's Qwen3.7 Flash is a native vision-language model with a 1M-token context window, priced at $0.03 per million input tokens with tool calling.

GPT-Live Audio Gets SynthID Watermarks and a Verify API
OpenAI now embeds Google DeepMind's SynthID watermark in all GPT-Live audio and opened a verification API so any team can check provenance automatically.

ChatGPT Research Program Opens to 100,000 Scientists
OpenAI's ChatGPT for Academic Researchers gives 10,000 scientists free frontier access now and 100,000 through 2027, part of a $250M science push.

Gemini in Slides Now Builds Entire Decks From a Prompt
Google's July Workspace drop adds Gemini deck generation in Slides, Omni video editing in Vids, and Gemini in Docs support for 11 more languages.

Gemini Robotics 2 Gives Humanoids Whole-Body Control
Google DeepMind's Gemini Robotics 2 landed July 30 with whole-body humanoid control, multi-robot teamwork, and up to 89.6% gripper accuracy.

OlmoEarth Cuts Wildfire Risk Mapping to 30 Hours
Ai2's OlmoEarth Platform mapped North American wildfire risk in 30.5 hours instead of 4,737, a 155x speedup at fractions of a cent per square km.

Kimi K3 Open Weights Ship 2.8T Parameters and 1M Context
Moonshot AI released Kimi K3 open weights on July 26: 2.8 trillion parameters, 104B active per token, 1M context, and a modified MIT license.
