# Video Forensics Gets a Fingerprint: Researchers Can Now Trace Deepfakes to Their Source Model
The interview went fine. The candidate answered every question, their webcam was steady, their voice was clear. Someone had already run the background check. So the company hired them — and handed over VPN credentials, a laptop, and access to internal systems.
What they hired was a face that didn't exist.
The North Korean IT worker scheme has been running long enough that it has its own FBI advisory, its own CISA bulletin, and its own body of corporate casualties. Companies across the United States have unknowingly paid salaries to individuals running laptop farms, using deepfaked video identities to pass remote interviews, then exfiltrating data or transferring paychecks back to Pyongyang. The face on the video call was AI. The voice was AI. The person was somewhere else entirely.
Detection tools have been trying to catch up. But catching up — flagging a video as synthetic — turns out to be only half the problem.
## Knowing It's Fake Isn't Enough Anymore
Researchers at UC Riverside recognized this gap and built something that goes further. Their framework, called SAGA (Source Attribution of Generative AI), doesn't just answer *is this video real?* It answers *what made it, what version, and who built the model.*
That's a meaningful leap. Think of the difference between a ballistics analyst saying "this is a bullet" versus "this bullet came from a Glock 19, manufactured in 2022, and we've seen rounds from this same weapon at three other crime scenes." One observation tells you something happened. The other one builds a case.
Rohit Kundu, a doctoral candidate at UC Riverside and research intern at YouTube, describes the progression his team went through: first building a single detection model that could handle video from any generative tool (rather than requiring separate models for each generator), then adding explainability — not just *fake* but *why fake* and *which artifacts gave it away* — and finally pushing into attribution, tying a synthetic video back to its originating model and development team.
Kundu, who studies AI-generated video daily, has said he himself can no longer reliably distinguish real from synthetic with the naked eye. That's the environment SAGA was built to operate in.
## The Attribution Problem Is Where the Real Accountability Lives
The question of which model generated a piece of synthetic media matters more than it might seem.
Right now, AI-generated disinformation and deepfakes exist in an accountability vacuum. A video surfaces. Someone flags it as synthetic. The platform takes it down — maybe. But there's no trail connecting the content back to the tool that made it, the actor who deployed it, or the model provider whose API was abused. Accountability stops at "this is fake."
SAGA creates the missing forensic layer. If investigators can identify that a cluster of fraudulent deepfake interview videos were generated using the same model version, or that a disinformation campaign targeting an election used a specific commercial generator, that's the kind of evidence that can support platform enforcement, regulatory action, or criminal referral.
The research team's goal isn't just building a detection gadget — it's promoting industry collaboration on shared standards for AI-generated content tracking. That framing matters. The long-term play here is closer to how the security community handles malware attribution: shared indicators, shared databases, provenance chains that let analysts compare findings across incidents.
## What Defenders Should Actually Do Right Now
SAGA isn't a commercial product yet. It's a research framework built from publicly available data. But the operational implications are immediate enough to act on.
For enterprises doing remote hiring, the deepfake interview attack is no longer theoretical. A layered verification protocol is the minimum: live liveness checks that go beyond simple webcam video, identity proofing against government documents with a human in the loop, and — critically — flagging any candidate who is evasive about in-person verification as a hard disqualifier. The North Korean campaign has specifically targeted tech companies seeking remote engineers, which means if you're hiring remotely for roles with privileged access, you're in the threat model.
For legal and compliance teams, attribution capability changes the landscape for AI-generated fraud cases. When forensic tools can identify the model used, it opens questions about provider liability, terms-of-service enforcement, and the evidentiary chain in fraud investigations. Legal teams advising clients on deepfake incidents should start building relationships with researchers who can provide this kind of forensic testimony.
For platform trust and safety teams, the framework points toward what should become standard infrastructure: provenance metadata embedded at generation, verified by platforms at upload. Several AI model providers have started implementing C2PA (Coalition for Content Provenance and Authenticity) standards voluntarily. SAGA's approach — model-level fingerprinting from the video artifact itself, without relying on embedded metadata — provides a backstop for when that provenance data is stripped or spoofed.
## HackWire Analysis
There's a pattern worth naming here, and it runs deeper than this single research project.
The security industry spent years learning that detecting malware wasn't sufficient — you needed attribution. Knowing you had a trojan told you to clean the machine; knowing it was from Lazarus Group told you to assume full network compromise and call the FBI. The same intellectual evolution is now happening with synthetic media, just five years behind. We've had deepfake detection for a while. Attribution is the missing piece, and SAGA is one of the first serious efforts to operationalize it.
What's underreported in coverage of this tool is how it closes a specific accountability loophole that AI companies have quietly benefited from. When a bad actor uses a commercial image or video generator to create fraudulent content, the provider faces no forensic trail connecting their product to the harm. That's starting to change. If SAGA-style attribution becomes standard — embedded in platform review pipelines, cited in court filings — the question of which model was used in a deepfake fraud case becomes answerable. That changes the risk calculus for model providers who've been lax about abuse prevention. It also creates pressure for regulators who've been waiting for attribution tools to mature before writing enforcement rules.
The North Korean IT worker scheme is the most visible current abuse case, but it's not the only one. Synthetic media is already being used in business email compromise schemes where video "proof" is sent to verify wire transfer requests. It's being used in romance fraud at scale. It's being used to fabricate crisis footage. Detection flagged these as problems. Attribution is what makes prosecution and platform enforcement possible.
The researchers at UC Riverside have built the forensic equivalent of ballistics for video. The industry needs to treat it that way.
— HackWire Editorial
---
## Related Coverage