Google introduces Gemini 3.8 Live with Live Avatar for multimodal enterprise agents

Google’s Gemini 3.8 Live with Live Avatar combines near-real-time video and speech, background tool execution, and synchronized multilingual interaction for enterprise agents.

Google announced Gemini 3.8 Live with Live Avatar on September 24, 2026, combining live dialogue, streaming video, and a digital avatar into an enterprise-agent experience. Google says the feature is available in Gemini Enterprise and is intended for customer service, interactive walkthroughs, and other services that need more than text or voice responses.

Live Avatar processes visual and audio input together and responds with video, expressive behavior, synchronized lips, and speech. Google presents it as a more natural multimodal interaction, but organizations still need to decide which scenarios benefit from visual presence. A richer interface also raises content, identity, accessibility, and disclosure requirements.

A notable technical feature is asynchronous tool execution. An agent can call tools and fetch data in the background while keeping the conversation active. Google uses hotel check-in as an example of a complex task in which the agent maintains dialogue while interacting with external systems. Deployments still need to distinguish readable data, executable actions, and steps that require human approval.

Google says Live Avatar can switch across 97 languages while adapting speech, lip-sync, and expressions. That is the product’s stated support range, not a guarantee of equal performance across accents, domain terminology, or cultural contexts. Real deployments should test accuracy and naturalness with the conversations of the intended market.

Organizations can choose from preset avatars or use a high-quality reference image to generate a custom avatar that preserves a brand style or character identity. Google says custom avatar creation is currently available only through enterprise allowlisting. Rights to likeness, consent from employees or customers, impersonation risk, and disclosure should therefore be part of the launch process rather than left to a product setting.

For transparency, Google says output generated by its AI products is watermarked with SynthID so audio and video can be detected more easily, helping reduce misinformation and misattribution. Watermarking can support provenance, but it does not replace visible AI disclosure, human oversight, complaint paths, or broader deepfake governance.

The direction of Gemini 3.8 Live with Live Avatar is a move from agents that answer questions to agents that remain present and coordinate tools. Its value will depend on whether the agent can retrieve the right data without breaking the conversation, pause for approval at the right moment, and remain understandable, traceable, and controllable across video and multilingual settings.

MODULE.002 //

More insights

Ideas on websites, AI automation, digital marketing, AI news, and VMTS updates.