# New AI Model Achieves Near-Human Conviction in Live Video Conversations
## A Breakthrough in Real-Time Human-AI Interaction
A recent development in artificial intelligence has raised both excitement and concern across the tech world. A new AI model, designed specifically for live face-to-face video interactions, has demonstrated an unprecedented ability to convince human participants that they are conversing with a real person. In controlled testing, the model managed to fool nearly half of all participants during one-minute video calls.
The technology represents a significant leap from earlier conversational AI systems. While previous generations struggled to maintain the nuance of a natural conversation — especially on video — this new approach integrates real-time expression tracking, voice modulation, and conversational awareness into a single unified model. It responds to visual cues, interprets pauses, and adapts its behavior mid-conversation, much like a human would.
## Testing Methodology and Results
The evaluation involved 54 participants who were each paired with the AI model for a brief one-minute video call. The topic of conversation was a simple, open-ended prompt: what each person was looking forward to in the coming year. Participants were not informed that they might be speaking with an AI system.
After the call, each participant was asked whether it had ever crossed their mind that their conversation partner might not be human. Of the 54 individuals, 26 — roughly 48% — admitted that they had believed the person on the other end was real.
To put this in perspective, an earlier version of the same company’s technology had been tested under similar conditions and only convinced a single participant out of 41. The dramatic improvement suggests that the underlying architecture has reached a new level of sophistication.
## Benchmark Performance
Independent benchmarking through NVIDIA’s VideoFDB evaluation suite further underscores the model’s capabilities. In the generation track — which assesses how natural and expressive a model’s responses appear — the system scored 3.83 out of 5. The next-best competing system managed only 2.80, while the human reference score was 3.92.
The perception track, which measures a model’s ability to understand and react to what it sees and hears, showed a more modest performance. Here the model scored 3.73 compared to a human reference of 4.20. Still, this represents a meaningful step forward in multimodal AI understanding.
NVIDIA conducted the evaluation independently, lending additional credibility to the benchmark results.
## Technical Architecture
Unlike earlier systems that stitched together multiple separate models — one handling visuals, another managing dialogue, and a third overseeing perception — the new model operates as a unified system. This integration allows for smoother, more cohesive interactions where all channels of communication flow simultaneously.
The model supports full-duplex communication, meaning it listens, watches, speaks, and reacts all at once. During demonstrations, it has been shown coaching someone through solving a Rubik’s cube by observing their hands in real time and pausing when the user takes time to think. The audio-to-video delay on NVIDIA H100 chips averages just 0.43 seconds, which the developers say is roughly half that of the next fastest competing method.
## Security and Ethical Concerns
The announcement has inevitably drawn attention to the security implications of increasingly realistic AI-generated video. Deepfake technology has already been weaponized in scams. In early 2025, North Korean-linked hackers were reported to have used AI-generated video on platforms like Zoom and Microsoft Teams, impersonating trusted contacts to trick victims into installing malware.
Security experts have warned that as these models become more convincing, traditional visual and verbal verification methods — such as asking someone to wave at the camera or answer personal questions — may no longer be reliable. Some companies have already begun adapting their hiring and verification processes in response, incorporating spontaneous, context-specific questions that are difficult for AI to handle convincingly.
The developers acknowledge these risks and have stated that the model is currently restricted to select trusted testers as part of a research preview. They are actively working with AI safety organizations to develop disclosure features and safeguards before any public release.
## Funding and Context
The company behind the technology recently secured $40 million in Series B funding led by CRV, signaling strong investor confidence in the direction of human-like conversational AI.
Industry analysts note that this development follows a broader trend. A separate study by researchers at the University of California, San Diego found that OpenAI’s text-based GPT-4.5 model convinced human judges it was real in 73% of text-only conversations when prompted to adopt a specific persona. The new video-capable model extends this concept into real-time audiovisual interaction, which many consider a far more challenging frontier.
—
## Frequently Asked Questions (FAQ)
**What is a Human Interaction Model?**
A Human Interaction Model is a type of artificial intelligence specifically designed to engage in face-to-face video conversations in a way that mimics human social behavior. Unlike text-based chatbots, these models process visual and audio information simultaneously, responding with appropriate facial expressions, vocal tone, and conversational timing.
**How was the model tested?**
Participants were paired with the AI for a one-minute video call about what they were looking forward to that year. They were told they would be speaking with another person and were only asked at the end whether they had ever suspected their partner might be AI.
**Why do only 48% of participants believe the AI is human?**
This is actually a significant jump from the previous version, which only fooled about 2.4% of participants. The remaining 52% likely noticed subtle inconsistencies in expression, timing, or conversational flow that gave the AI away, particularly within the first 20 seconds of the call.
**What is NVIDIA’s VideoFDB benchmark?**
VideoFDB is an evaluation framework developed by NVIDIA that tests AI models on live audio and video conversations. It measures both the naturalness of generated responses and the model’s ability to perceive and understand what it sees and hears in real time.
**Is this model available to the public?**
No. The model is currently restricted to select trusted testers as part of a research preview. The company has stated that safety measures and disclosure features are still being developed before a public launch.
**What are the security risks associated with this technology?**
As AI-generated video becomes more convincing, it becomes easier for bad actors to impersonate trusted individuals in video calls. This has already been exploited in scams, including cases where deepfakes were used to impersonate colleagues or contacts in order to deliver malware.
**How fast does the model respond in real time?**
On NVIDIA H100 hardware, the audio-to-video delay averages 0.43 seconds, which is approximately half the latency of the next fastest competing method. This near-instantaneous response is critical for maintaining the illusion of a natural conversation.
**What is full-duplex communication?**
Full-duplex communication means the AI can listen, watch, speak, and react simultaneously — much like a traditional phone call — rather than processing input and output in separate turns like a walkie-talkie.
—
## Conclusion
The rapid advancement of AI models capable of realistic, real-time video interaction marks a turning point in how we think about digital communication. While the technical achievements are impressive — from near-human benchmark scores to full-duplex conversational fluency — they also underscore the urgent need for robust safety frameworks and public awareness. As these systems move closer to widespread availability, the line between authentic human interaction and AI-generated conversation will continue to blur, demanding new tools and safeguards to protect individuals and organizations alike.
Thank you for reading



