Articles Tagged “LLM”
17 articles found

GPT-5.6 Sol API Price Falls to $4 Per Million Tokens
OpenAI cut GPT-5.6 Sol API pricing on August 21: input drops 20% to $4 and output falls 33% to $20 per million tokens through November 21, 2026.

Apple Music Will Show Made With AI Labels This Year
Apple Music will make its AI Transparency Tags mandatory and surface Made With AI labels to listeners later in 2026, shifting from removal to disclosure.

Embedding Models Guide: Dense vs Multi-Vector RAG
A practical 2026 guide to choosing embedding models for RAG: dense bi-encoders, multi-vector late interaction, and rerankers, with cost and memory math.

Claude Opus 5 Brings Frontier Coding at Half the Cost
Claude Opus 5 arrives at $5 per million input tokens, delivering near-frontier coding performance at half the cost of Anthropic's Fable 5 model.

Gemini 3.6 Flash Trims Output Tokens by 17% for Devs
Google's Gemini 3.6 Flash uses 17% fewer output tokens at equal quality, costs $1.50 per million input, and jumps to 49% on the DeepSWE benchmark.

Kimi K2.7 Code Brings Open Weights to GitHub Copilot
GitHub Copilot's first open-weight model, Kimi K2.7 Code, reached Business and Enterprise plans on July 7 — a 1T-parameter MoE with 32B active.

GPT-5.6 Tiers Explained: Which of Sol, Terra, Luna Fits
OpenAI's GPT-5.6 shipped July 9 in three tiers — Sol, Terra, Luna — priced from $1 to $30 per million tokens. Here's how to pick the right one.

Cohere Transcribe Arabic: Best Open Arabic Speech AI
Cohere Transcribe Arabic, released July 7 under Apache 2.0, hits a 25.87 word error rate — the top open-source Arabic ASR, beating Whisper Large V3.

GPT-Live Gives ChatGPT Real-Time Full-Duplex Voice
OpenAI's GPT-Live brings full-duplex voice to ChatGPT, letting the AI listen and speak simultaneously for 150M+ weekly voice users worldwide.

SWE-1.7 Brings Near-Frontier Coding Power to Devin
Cognition's SWE-1.7 scored 42.3% on FrontierCode Main at about $1.97 per task, bringing near-frontier coding into Devin at ~1000 tokens/sec.

ByteDance's Seed 2.1 Models Bring Frontier Coding to a Lower Price Point
ByteDance unveiled Seed 2.1 Pro and Turbo on June 24, 2026 — strong coding and agent models with million-token context and a dramatically lower cost of ownership.

Microsoft Launches Seven In-House MAI Models With Frontier Tuning
Microsoft unveiled a family of seven in-house MAI models spanning reasoning, coding, image, voice, and transcription — plus Frontier Tuning to customize them on your own data.

ChatGPT's New 'Dreaming' Memory Learns About You in the Background
OpenAI's 'Dreaming' lets ChatGPT synthesize and self-update memories across chats automatically, with a transparent page to review and edit what it recalls.

Meta Launches Muse Spark: Its First Closed-Weight Frontier AI Model
Meta Superintelligence Labs drops Muse Spark on April 8 — a fully closed frontier AI competing with GPT-5.4 and Gemini, marking Meta's sharpest strategic turn yet.

NVIDIA's AI-Q Blueprint Brings Enterprise Agentic AI to Adobe, Salesforce, and SAP
NVIDIA's AI-Q Blueprint gives enterprises an open framework for building AI agents that perceive, reason, and act — slashing query costs by 50% with a hybrid routing architecture.

PrismML's Bonsai Is a 1-Bit LLM That Runs on a Smartphone and Matches Full-Size Models
Caltech startup PrismML emerged from stealth with Bonsai, a 1-bit LLM family that's 14x smaller, 8x faster, and 5x more energy-efficient than standard 8B models — and runs on an iPhone.

Alibaba's Qwen3.6-Plus Delivers 1M-Token Context and Repository-Level Agentic Coding
Qwen3.6-Plus arrives with a default 1 million-token context window and breakthrough agentic coding performance, enabling AI that can navigate and rewrite entire software repositories autonomously.
