Articles Tagged “AI Coding”
15 articles found

Gemini 3.7 Flash Benchmarks: What Developers Get Now
Gemini 3.7 Flash arrived August 13 with DeepSWE v1.1 at 65.3%, WebDev Arena Elo of 1588, and intro pricing of $0.75 per million input tokens.

Claude Code Auto Mode Turns On by Default August 14
Anthropic makes Claude Code auto mode the default for Pro, Max, and Team on August 14, after a 1,053-person study found it caught 89% of harmful actions.

Meta Muse Code Pairs a Terminal Agent With Muse Spark 1.2
Meta shipped Muse Code, a terminal coding agent co-trained with Muse Spark 1.2, a model scoring 54 on the Artificial Analysis Intelligence Index.

Token Monitor Shows AI Coding Usage on a 4-Inch Screen
Token Monitor is an ESP32-S3 desktop display tracking Claude Code, Codex CLI, and Antigravity quota, starting at 99 euros with an Apache-2.0 local broker.

Cursor Router Auto-Picks the Cheapest Capable Model
Cursor Router, launched July 22, routes each coding request to the cheapest capable model, delivering frontier quality at up to 60% lower cost.

AI Agent Sandbox Design: 4 Lessons From New Research
Pillar Security's seven disclosures across Cursor, Codex CLI and Gemini CLI reveal four sandbox failure modes AI agent builders can design against.

M5Stack Core2 Firmware Clones a $230 Codex Macro Pad
Open-source firmware turns a roughly $50 M5Stack Core2 into a working equivalent of OpenAI's $230 Codex Micro control pad, using its touchscreen.

Meta Muse Spark 1.1 Is a Budget Agentic Coding Model
Meta's first paid model, Muse Spark 1.1, launched July 9 at \$1.25 / \$4.25 per million tokens with a 1M-token context and a computer-use mode.

Poolside's Laguna XS 2.1 Puts a Free Open-Weight Coding Model on Your Machine
Poolside released Laguna XS 2.1 on July 2, 2026 — a free, permissively licensed open-weight coding model that scores 70.9% on SWE-bench Verified and runs locally.

Cursor's First Mobile App Lets You Steer Coding Agents From Your Phone
Cursor launched its first mobile app on June 29, 2026, letting developers start and supervise autonomous coding agents from iOS or an Android PWA, anywhere they go.

GLM-5.2 Open Weights Arrive as a Top Coding Model at a Fraction of the Cost
Z.ai released GLM-5.2 open weights under an MIT license on June 16, 2026 — an open-weight coding model that rivals the best closed systems on long-horizon benchmarks at roughly one-sixth the cost.

MiniMax M3: An Open-Weight Model With Frontier Coding and a 1M-Token Context
MiniMax M3, released June 1, 2026, is an open-weight LLM pairing frontier-level coding, a 1-million-token context window, and native multimodality — and the weights are coming to Hugging Face.

Microsoft's MAI-Code-1-Flash Brings a Tiny, Fast Coding Model to Copilot's Free Tier
Microsoft launched MAI-Code-1-Flash on June 2, 2026 — a compact 5B-parameter in-house coding model now rolling out across GitHub Copilot, including the free tier.

Claude Opus 4.7 Is Here: +13% Coding, 3× Vision Gains, and a New Performance Ceiling
Anthropic releases Claude Opus 4.7 today with 87.6% on SWE-bench Verified, 70% on CursorBench, and 98.5% visual acuity — taking the top spot on agentic coding benchmarks ahead of GPT-5.4 and Gemini 3.1 Pro.

DeepSeek Unveils V4 — A Trillion-Parameter Multimodal Model That Generates Text, Images, and Video
DeepSeek's V4 model enters the frontier tier with trillion-parameter multimodal capabilities spanning text, image, and video generation plus elite coding performance.
