Articles Tagged “Voice AI”
16 articles found

Microsoft MAI Voice 2.1: Faster Speech for Voice Agents
Microsoft's MAI-Transcribe-2-Streaming returns first words in about 100ms across 60 languages for $0.54 an hour, alongside two new MAI-Voice-2.1 models.

MaTouch ESP32-S3 E Ink Board: A $60 Voice-Ready Dashboard
Makerfabs' MaTouch ESP32-S3 pairs a 7.5-inch four-color E Ink screen with a four-mic array for $59.80. Here are the specs and what you can build.
Gemini Live Avatar: What Real-Time AI Faces Can Do Now
Gemini 3.8 Live with Live Avatar adds a lip-synced, animated face to voice agents in 97 languages. Here is what enterprises can build with it today.

Gemini 3.8 Flash TTS: What Prompt-Built AI Voices Can Do
Gemini 3.8 Flash TTS and Flash-Lite TTS let developers design new AI voices from a text prompt in 100+ languages. Here's what both models can do.

Grok 4.7 in GitHub Copilot: What Changes for Developers
Grok 4.7 lands in GitHub Copilot on day one at $2 per million input tokens, with SpaceXAI reporting 71% on DeepSWE and a stronger agentic coding loop.

Gemini 3.8 Live Adds Reasoning to Real-Time Voice AI
Google's new Gemini 3.8 Live models reason mid-conversation across 97 languages and top the speech quality index at 82.6. Here is what changes.

iOS 27 Siri: What the New Assistant Can Actually Do
Apple shipped iOS 27 on September 14 with a rebuilt Siri, a new System Orchestrator and 250+ changes. Here is what the assistant actually does now.

GPT-Live-1 API: Build Voice Agents for $0.05 a Minute
GPT-Live-1 is now in the OpenAI API at $0.05 a minute: a full-duplex voice model that OpenAI says scores 86.2% on Tau3, up from 45.7% for its predecessor.

Gemini 3.5 Transcribe Cuts Word Error Rate to 2.6%
Google's Gemini 3.5 Transcribe replaces Chirp 3 with a 2.6% word error rate, automatic detection across 85+ languages and 70% faster final transcripts.

Wispr Raises $280M and Ships Its Canto Speech Model
Wispr closed a $280M Series B at a $2B valuation and launched Canto, an in-house speech model it says cuts dictation error rates from 30% to under 10%.

GPT-Live Audio Gets SynthID Watermarks and a Verify API
OpenAI now embeds Google DeepMind's SynthID watermark in all GPT-Live audio and opened a verification API so any team can check provenance automatically.

OpenAI Presence Puts Enterprise AI Agents on Guardrails
OpenAI launched Presence on July 22, an enterprise platform for voice and chat AI agents with guardrails, approved actions, and human escalation built in.

Real World VoiceEQ Benchmarks the Human Side of Voice AI
Hume AI and Hugging Face open a voice AI benchmark built on over one million human ratings, covering 40+ models and 60+ metrics of speech quality.

GPT-Live Gives ChatGPT Real-Time Full-Duplex Voice
OpenAI's GPT-Live brings full-duplex voice to ChatGPT, letting the AI listen and speak simultaneously for 150M+ weekly voice users worldwide.

NanoPi M6V2 Adds Dual-Mic Input, Making This RK3588S SBC a Voice-AI Pick
FriendlyELEC's NanoPi M6V2 gains dual analog microphone input on its RK3588S board — a small, smart upgrade for voice assistants and on-device audio AI projects.

OpenAI Drops Three Voice Models Into the API — GPT-Realtime-2, Translate, and Whisper
OpenAI shipped GPT-Realtime-2, Realtime-Translate, and Realtime-Whisper into the API on May 6, 2026 — bringing live voice reasoning, 70-language translation, and streaming transcription to developers.
