Artificial Intelligence — Page 4
425 articles in this category

GPT-6 Astra: What It Costs and What It Can Automate
GPT-6 Astra costs $10 per million input tokens, holds a 1.05M-token context, and finished 72.6% of OSWorld desktop tasks in roughly half the time.

WeatherNext 3 Brings 5km Hourly Forecasts to Search
Google DeepMind's WeatherNext 3 forecasts at 5km resolution every hour and improves rain accuracy up to 60% over WeatherNext 2, live in Search now.

Gemini Agentic Video Understanding Cuts Tokens 88%
Gemini agentic video understanding cuts token use up to 88% and cost up to 66% while raising accuracy 7%, with no extra fee on the Gemini API.

Gemini 3.8 Flash Cyber Fixes 2.6x More Chrome Bugs
Google's Gemini 3.8 Flash Cyber wrote 2.6x more correct Chrome patches than larger commercial models, and 3.8 Flash starts at $0.75 per million tokens.

ChatGPT Health Connects Epic Records for Clinicians
OpenAI connected ChatGPT for Healthcare to Epic EHR records with read-only access, plus a data plugin covering PubMed, RxNorm and ClinicalTrials.gov.

Claude Fable 5.1 Doubles Science Scores, Cuts Costs
Claude Fable 5.1 more than doubles agentic science benchmarks and cuts typical workload costs 25%, with cache reads dropping to $0.25 per million tokens.

Tencent Hy4 Ships 770B Open Weights Under Apache 2.0
Tencent open-sourced Hy4 preview on August 28 with 770B total parameters, 49B active per token, a 1M-token context window and Apache 2.0 weights.

WebGPU Kernels Make Local AI in the Browser 2.57x Faster
Hugging Face published 207 WebGPU kernels as a JavaScript library, reporting a 2.57x geometric-mean speedup over ORT WebGPU on an Apple M4 GPU.

South Korea Gives 52 Million Citizens Free AI Access
South Korea picked SK Telecom, KT and Kakao to give every citizen free AI agents, backed by 512 Nvidia B200 GPUs and a December 2026 launch.

Debian Adopts Responsible Generative AI Contribution Rules
Debian's 1,045 eligible voters settled an eight-option ballot on AI-assisted contributions. Option E won: use the tools, own the output, same standards.

Claude for Scientists Opens 10,000 Free Research Seats
Anthropic opened 10,000 Claude Team seats for academic labs. Standard access is free, premium seats with 5x limits run $15 a month, locked for a year.

IBM Granite 4.2 Brings Open Reasoning Models Local
IBM released Granite 4.2 under Apache 2.0 in 3B, 8B and 30B sizes, with a switchable thinking mode, a 512K context window and a 57.00 SWE-bench score.

Solid-State Cooling Uses Waste Heat, Not Electricity
A KIT and Tsukuba prototype cools using only waste heat, hitting a 13 K span in the refrigerant with zero electrical input. Published in Nature Energy.

Claude Desktop Now Runs Local Models Through Ollama
Ollama's new Claude Desktop integration is one toggle: local or cloud open models appear in Claude's picker, with the full agent toolset intact.

Qwen3.8-Flash-Next Fits in 75GB With No GPU Needed
Qwen's 125B Flash-Next MoE runs locally in 75GB of RAM with no GPU VRAM required, scoring 62.5 on SWE-bench Pro with just 6B active parameters.

Model Hardware Standard Lets AI Agents Run Lab Gear
Anthropic's Model Hardware Standard cuts lab instrument integration from weeks to minutes, with Genentech, Carnegie Mellon and QuEra among first users.

OpenAI Jalapeño Benchmarks: 1.9x Work per Kilowatt
OpenAI published the first Jalapeño inference benchmarks: a 700W part claiming up to 1.9x throughput per kilowatt and 3.6x lower end-to-end latency.

GLM-5.3-Flash 3-Bit Quant Runs on 128GB of Local RAM
Z.ai's 320B GLM-5.3-Flash now runs at 3-bit on 128GB of RAM via Unsloth GGUFs, retaining 82% of top-1 accuracy at under a fifth of its 650GB size.

Gemini 3.5 Transcribe Cuts Word Error Rate to 2.6%
Google's Gemini 3.5 Transcribe replaces Chirp 3 with a 2.6% word error rate, automatic detection across 85+ languages and 70% faster final transcripts.

Meta MTIA 400 Puts 9.4TB/s HBM3e Behind FP4 Inference
Meta detailed MTIA 400 at Hot Chips 2026: eight HBM3e stacks, 9.4TB/s of bandwidth, hardware FP4, and scale-up domains reaching 72 accelerators.

Claude Memory Now Carries Between Chat and Cowork Tasks
Anthropic unified Claude's memory across chat and Cowork, with topic-by-topic editing and sensitive categories excluded by default on Free, Pro and Max.

Gemini Enterprise for Legal Brings AI Agents to Law Firms
Google Cloud launched Gemini Enterprise for Legal with four launch firms, iManage and Thomson Reuters connectors, plus audit logging and ethical walls.

Intel Diamond Rapids Packs 256 Cores for Agentic AI
Intel's Hot Chips 2026 lineup pairs a 256-core Diamond Rapids Xeon with Crescent Island, a 350W inference GPU holding up to 480GB of LPDDR5X.

Nvidia Vera Rubin NVL72 Targets 30x Tokens Per Watt
Nvidia detailed the Vera Rubin NVL72 rack at Hot Chips 2026: 72 GPUs, 2 ZFLOPS of NVFP4 inference and up to 30x more tokens per megawatt.
