Skip to content

The Data Scientist

Why Deepfake Detection Needs a New Foundation – And Why It Can’t Wait

By Zohaib Ahmed, CEO and co-Founder, Resemble AI

A finance employee gets a video call from their CFO, asking for an urgent wire transfer. A recruiter interviews a candidate over Zoom. A claims adjuster opens a photo of storm damage submitted through an insurance app. Each interaction looks and sounds exactly as it should, yet increasingly, there is no guarantee that any of these scenarios are actually real.

This isn’t a hypothetical scenario for security teams anymore. It’s a Tuesday.

The scale problem is now measurable

For years, deepfakes were treated as an edge case: rare enough that most organizations could reasonably bet against ever encountering one. That bet no longer pays off. Incident data my team tracks at Resemble AI shows 821 verified deepfake attacks in the first six months of 2026 alone, drawn from more than 1,700 news reports. Combined potential media reach from those incidents topped 290 billion, matching the full-year total for 2025 in half the time. You can view the entire 2026 Deepfake report here.

Those are just the incidents that made the news. The real number is almost certainly higher, since organizations that catch and stop a deepfake attack quietly have no incentive to publicize it. Our data shows a third of reported fraud attempts were caught before any money moved, and completed frauds tend to stay out of headlines entirely. Corporate fraud in particular tends to surface through a single report and then go dark, which suggests the incidents we can count understate the problem by a wide margin.

The reason for this growth isn’t mysterious. Generative AI models are multiplying at a pace no security team could have planned around two years ago. Hugging Face’s public model count has grown from roughly 2 million to nearly 3 million in the past nine months, with a growing share built specifically to generate images, audio, and video. Every one of those tools is a new way to fake a voice, a face, or a scene, and many are free, fast, and require no technical skill to use.

Gartner made this explicit in June, naming identity impersonation via deepfakes one of four critical, unpredictable threats where attackers currently hold the advantage over defenders. Its guidance to CISOs is blunt: deepfake detection alone won’t solve the problem, but it’s an essential layer among several needed to protect biometric verification, real-time meetings, and call center authentication. That framing matters, because it puts detection back on the roadmap for security leaders who might otherwise treat it as someone else’s problem, whether that’s fraud, trust and safety, or brand protection.

Why detection alone has struggled to keep up

Here’s the uncomfortable part: most deepfake detection technology built over the last several years was designed to recognize what it had already seen. Feed it a training set of known generators, and it gets very good at flagging their fingerprints – the sub-visual noise patterns and mathematical artifacts particular AI tools leave behind. But show it output from a brand-new model, one that shipped last week, and accuracy can drop off a cliff until the detector is retrained.

That gap is exactly what attackers are learning to exploit. When X opened its Grok image tool to every user for a two-week window earlier this year, our data shows it produced millions of images at a rate few dedicated detection systems, or trust and safety teams, were built to absorb. Whatever the eventual legal reckoning for that episode looks like, the operational lesson for every enterprise is the same: the next tool capable of that kind of output might not come with two weeks’ warning.

Retraining a detector every time a new generator appears isn’t sustainable against an industry releasing thousands of new models a day. It also means the detection gap for a genuinely novel attack is often widest at the moment it matters most.

A different question to ask

At Resemble AI, we built our newest detection architecture, DETECT-World, around a different premise. Instead of asking only “have I seen this pattern before,” it also asks whether what’s on screen or in a call could actually happen in the physical world. Does the lighting on a face track the light source in the room? Does a shadow move the way gravity and geometry say it should? Is the audio actually synchronized with the video, frame for frame?

This is the same class of architecture, a world model, that leading AI labs are racing to build for generation itself: systems that understand physical reality well enough to simulate it convincingly. We think the same understanding is just as valuable when pointed the other direction, at verification. A fake that mimics every known artifact pattern perfectly can still break the laws of physics, and that’s a signal no amount of pixel-level pattern matching will catch on its own. Just as important, a model trained to reason about physical plausibility doesn’t need to see a specific new generator to flag its output; it can generalize to zero-day tools from day one, which shortens the dangerous window between a new model’s release and the point where defenders catch up.

This isn’t only an enterprise problem

It’s tempting to file deepfake defense under “problems only Fortune 500 companies need to worry about.” The data doesn’t support that. Private individuals make up the majority of documented victims once you exclude the single largest mass-incident case in our dataset. Community banks, regional insurers, mid-market recruiters, and local government offices are all exposed to the same executive-impersonation calls, synthetic ID submissions, and fabricated claims photos as any global enterprise, often with fewer resources to catch them. A generative tool doesn’t check a target’s revenue before it’s used against them.

Where this leaves security leaders

No detection system, including ours, will ever catch everything, and no vendor should claim otherwise. But the direction is clear. As the AI generation gets better at mimicking reality, detection has to understand reality too, not just recognize the fingerprints of yesterday’s tools. Pair that with layered controls around biometric verification, call authentication, and human review, and organizations of any size can close a gap that’s currently working against them.

The call, the photo, the video call: increasingly, none of them can be taken at face value. The organizations that adapt their verification posture now, rather than after their own incident makes the news, will be the ones still trusted when it matters.

Zohaib Ahmed is the CEO and co-founder of Resemble AI, a deepfake detection model company.