
Tavus previews Griffin for full-duplex video conversations
The AMW Read
Griffin meaningfully updates Tavus's conversational-video architecture, but limited access and a nonstandard human-resemblance trial keep demonstrated impact within the digital-human subsegment.
Tavus previews Griffin for full-duplex video conversations
U.S. startup Tavus announced Griffin on October 1, introducing a full-duplex video-to-video architecture that processes speech and visual cues while generating responses. In a company-run test, 26 of 54 participants, or 48%, judged its Griffin-Lite AI to be human after a one-minute video call. The previous system recorded one such judgment among 41 participants, or 2.4%. Griffin-Lite remains a research preview available to selected testers, with additional safety, alignment and AI identity disclosure work planned before general release.
For the generative-media market, the meaningful shift is from rendering a digital human to coordinating an ongoing conversation. Griffin separates continuous dialogue modeling from audiovisual generation, allowing nods, interjections and expression changes while a user is still speaking. Tavus reports average audio-to-video latency of 0.43 seconds on an NVIDIA H100. These capabilities make conversational timing and responsiveness potential differentiators alongside visual quality. The human-misidentification result needs qualification: participants were told they would chat with another participant and were not warned that the counterpart might be AI. That limits what the result establishes about passing a conventional Turing test.
Builders evaluating customer-support or training applications should test sustained interaction quality and clear identity disclosure before adopting the system. Tavus reports leading results on multiple NVIDIA VideoFDB metrics, but benchmark performance and a short human-resemblance trial do not establish production reliability or customer ROI. For investors, the concrete next evidence is whether the limited preview can translate audiovisual responsiveness into useful, repeatable workflows as access broadens.

