Trending: On-device modelsSearch
iHeartGeek
iTECH

Gemini 3.8 Live puts a talking face on enterprise AI

Google has paired its Gemini 3.8 Live speech models with low-latency video, giving enterprise agents a lip-synced, multilingual face with custom avatars built from a single reference image.

Google's announcement card for Gemini 3.8 Live with Live Avatar: the product name set in black across the centre of a soft blue gradient, the multicoloured Gemini star mark beneath it and a blurred white speech bubble to the upper right.

Google has given its Gemini 3.8 Live speech models a face. Live Avatar, announced by the company on 24 September, brings near real-time visual presence to Gemini's conversational AI by coupling the live dialogue models with low-latency streaming video, and it is available immediately inside Gemini Enterprise. That last detail is the one worth noticing: Google is not shipping a companion for consumers, it is selling a front-of-house interface to companies that currently answer questions with a text box.

A face for the voice models

The avatar speaks with precise lip-syncing, natural expressions and fluid turn-taking, which is the part most voice assistants still fail. Google frames conversation as inherently multimodal - we listen, look, speak and read each other's faces - and Live Avatar processes visual and audio input at the same time, so it responds to what it sees and hears rather than only to a transcript of the question asked.

Tool calls that do not interrupt the conversation

The more practical advance is asynchronous tool calling. Live Avatar can trigger a tool call and fetch data in the background while the dialogue carries on, so a booking lookup or a stock check no longer produces the polite pause that makes most voice assistants sound scripted. Google's own demonstration is a hotel front desk, where the avatar pulls up a guest's details mid-sentence without dropping the thread.

97 languages and a custom face

Live Avatar carries native speech-to-speech synchronisation across 97 languages, adapting its lip-sync and expressions as it switches mid-conversation instead of degrading into a dubbed recording. Alongside a library of preset avatars, organisations can generate their own from a high-quality reference image and keep the result recognisably close to that picture, a capability that will appeal to brands and trouble anyone who has watched a likeness be reused without permission.

Google says the feature was built with safeguards designed to respect identity, and that all audio and video output carries a SynthID watermark, the imperceptible marker the company embeds in its media products. API documentation is published for developers who want to build on the feature.

Our opinion

Enterprise first is the tell, and it is a sensible order of play: a brand can decide what its avatar may say in a way a consumer cannot. The risk that matters here is not the uncanny valley, it is fraud. A face that lip-syncs convincingly in 97 languages is a better social engineering tool than any amount of generated text, and Google's own prominence for SynthID suggests the company knows it. Custom avatars push the consent question straight onto whoever supplies the reference photograph, and no watermark answers that. The tool-calling behaviour is the genuinely new engineering, because a conversation that keeps running while software works is the difference between a demo and a colleague, and it is also where a mistake gets expensive.