Gemini Live Avatar: What Real-Time AI Faces Can Do Now
Gemini 3.8 Live with Live Avatar adds a lip-synced, animated face to voice agents in 97 languages. Here is what enterprises can build with it today.
Gemini Live Avatar gives Google's real-time voice model a face. Announced on September 24, 2026, Gemini 3.8 Live with Live Avatar pairs near real-time video generation with spoken dialogue, so an AI agent can hold a conversation with lip-synced speech, natural expressions and fluid turn-taking. Google says the feature is generally available in Gemini Enterprise.
- Live Avatar generates an animated, lip-synced face in near real time alongside Gemini 3.8 Live speech.
- Google cites native speech synchronization across 97 languages.
- Tool calls can run in the background while the conversation continues.
- Every audio and video output carries an imperceptible SynthID watermark.
What is Gemini Live Avatar?
Gemini 3.8 Live already handled low-latency voice conversations; our earlier look at Gemini 3.8 Live's reasoning upgrade covered that foundation. Live Avatar adds the visual layer. According to Google's research and engineering authors, the system processes visual and audio input at the same time, which means a user can speak, show a camera view or share a screen while the avatar responds.
The practical shift is from a disembodied voice to a presence that looks at you, reacts and speaks with matching mouth movements. For training simulations, customer help desks, tutoring and guided onboarding, that visual feedback can make an exchange feel easier to follow.
How does multilingual lip-sync work across 97 languages?
Google describes native multilingual speech-to-speech synchronization, so the avatar's lip movements track whichever of the 97 supported languages is being spoken. Unite.AI's coverage adds that the avatar follows along when a speaker switches languages mid-conversation. For global support teams, that could mean one agent design serving many markets instead of a separate build per language.
Background tool calls keep the conversation moving
A real-time agent that falls silent while it looks something up breaks the illusion of a conversation. Live Avatar uses asynchronous tool calling: the agent can trigger a lookup, fetch data and keep talking while the result arrives. Unite.AI highlights Google's demonstration of an insurance-claims assistant as an example of that pattern.
Custom avatars and SynthID safeguards
Organizations can start from preset avatars. Google also supports creating a fully animated avatar from a single high-quality reference image, but custom avatar creation is currently limited to allowlisted enterprise customers. That gate is a sensible precaution for a tool that animates faces.
Every audio and video output is marked with SynthID, Google's imperceptible watermark, so generated content remains detectable. It is a constructive default: realistic synthetic video arrives with a provenance signal built in from day one.
Who can use Gemini 3.8 Live with Live Avatar?
Live Avatar is an enterprise product, available through Gemini Enterprise with API documentation in the Google Cloud console and, according to Unite.AI, US and EU endpoints. It is not part of the consumer Gemini app, and Google did not publish pricing or benchmark figures in its announcement. Developers already building voice agents on Gemini, including those experimenting with Gemini 3.8 Flash TTS voice design, now have a path to add a visual front end without stitching together a separate video pipeline.
For more model and platform releases, follow our AI coverage. The next milestones to watch are published pricing, customer deployments and whether the avatar layer reaches broader Gemini products.
Sources: Google: Gemini 3.8 Live with Live Avatar — September 24, 2026; Unite.AI: Google brings Live Avatar visual presence to Gemini 3.8 Live — September 24, 2026. Capability descriptions are attributed to Google; no independent benchmarks were published.
More AI Stories

Strands Decider 2B: What Amazon's Free Decision Model Does
Strands Decider 2B is a free 2B-parameter decision model that picks options in about 115ms on a single GPU. Here is how it works and where it fits.

Meta Muse Gadgets SDK: Build Your Own AI Agent Hardware
Meta's open-source Muse Gadgets SDKs bring its Muse agent to Raspberry Pi and ESP32 builds, and 5,000 free Home Link dongles are going to subscribers.

AstaBrief 8B: Ai2's Open Model for Cited Science Reports
AstaBrief 8B is Ai2's Apache 2.0 open model that writes cited research reports in 51 seconds, 3.5x faster than its Claude pipeline. Here is how it works.
