AI Scout Daily
News

Tavus Says Its Griffin Video Model Passed a Face-to-Face Turing Test. The Test Was a One-Minute Call, and You Can't Use It Yet

48% of participants thought a Griffin-powered video agent was a real person after a one-minute call. Tavus says it is holding the model back from customers over deception risks.

Priya Nair·October 3, 2026·5 min read

TL;DR

Tavus announced Griffin, which it calls the first Human Interaction Model: a video-to-video model that listens, watches, speaks, and moves at the same time. After one-minute live calls, 48% of participants thought they were talking to a real person, against 2.4% for earlier systems on a similar test, per Tavus. It also reports 0.43-second video generation latency and a first-place result on NVIDIA's VideoFDB benchmark. Griffin-Lite is a research preview for select testers only, with no pricing or release date, and Tavus says it is building disclosure features first.

Tavus, which sells AI video agents, says it has built a model that can hold a live video call convincingly enough to fool nearly half the people who try it. All the numbers below are Tavus's own, from its Griffin page; none has been independently reproduced.

What Griffin is

Most video agents are a chain: speech is transcribed, a language model writes a reply, a voice model reads it, and an avatar lip-syncs it. Tavus describes Griffin as a single video-to-video system that perceives and responds at once. It can interrupt, give backchannel cues, track how long a silence has lasted, and generate every pixel, including body movement, shadows, and background changes, according to Tavus. It uses a custom audio codec, Tavec, that compresses audio into 40-value frames at 100 frames per second so it can stream in 10ms packets.

The numbers Tavus reports

MeasureTavus's resultComparison given
Thought it was a real person (one-minute live call)48%2.4% for previous systems on a similar test
NVIDIA VideoFDB, generation track3.83 out of 52.80 for the next-best system; human reference 3.92
NVIDIA VideoFDB, perception track3.73 out of 53.44 for the strongest baseline; human reference 4.20
Video generation latency0.43 seconds averageAbout 50% faster than comparable systems
All figures from Tavus's Griffin page; not independently verified.

How to read the 48%

Tavus presents it as the first model to pass a face-to-face Turing test. The details are narrower than the headline: calls lasted one minute, and AI Weekly's roundup reports the figure as 48% of testers. The Tavus page does not give the number of participants, how they were recruited, or whether there was a control group, and the only version it says testers can access is Griffin-Lite, a research preview. A one-minute chat shows the model can hold a short exchange, not that people couldn't tell over a long call.

Availability and risk

Griffin-Lite is a research preview for select testers. Tavus says it is not available to customers at this time, and has not announced pricing or a timeline. It also states plainly that the properties that make the model useful "allow them to deceive a human into believing it is not AI", and says it is developing disclosure features and working with AI safety organisations before wider release.

What this means if you buy video agents

Nothing here is something you can buy or test today, so don't change plans yet. If you're building customer-facing video agents, the useful takeaways are the benchmarks to ask vendors about (full-duplex interaction, not just lip-sync) and the disclosure question: decide now how your product will tell users they are talking to an AI, because the quality that makes the experience good is also what makes that disclosure matter.

PN
Priya Nair

AI Industry Reporter

Priya covers model releases, industry announcements, and the gap between what labs claim and what independent evaluators actually find. She reads the primary source - the paper, the system card, the benchmark org's own statement - before writing a word.