Articles Tagged “Benchmark”
5 articles found

Claude Opus 5 Brings Frontier Coding at Half the Cost
Claude Opus 5 arrives at $5 per million input tokens, delivering near-frontier coding performance at half the cost of Anthropic's Fable 5 model.

Gemini 3.6 Flash Trims Output Tokens by 17% for Devs
Google's Gemini 3.6 Flash uses 17% fewer output tokens at equal quality, costs $1.50 per million input, and jumps to 49% on the DeepSWE benchmark.

Qwen3.8-Max Benchmarks: What to Watch in the Preview
Alibaba's Qwen3.8-Max preview brings 2.4 trillion parameters and multimodal input, but no benchmark table yet. Here's the baseline to measure it against.

Real World VoiceEQ Benchmarks the Human Side of Voice AI
Hume AI and Hugging Face open a voice AI benchmark built on over one million human ratings, covering 40+ models and 60+ metrics of speech quality.

OpenAI's GeneBench-Pro Sets a Rigorous New Bar for AI in Biology Research
OpenAI open-sourced GeneBench-Pro on June 30, 2026 — a 129-problem computational biology benchmark graded against ground truth, pushing AI toward trustworthy science.
