That viral clip of a woman water-skiing inside a tornado? Probably fake. The one where a skyscraper folds like origami? Almost certainly AI. The problem isn’t that AI-generated video exists — it’s that it’s gotten good enough to fool people who should know better, including you, including me.
So what do you actually do when something looks a little off but you can’t quite prove it?
Why AI Video Is So Hard to Catch
The tells that used to give away AI video — melting fingers, garbled text, faces that shimmer at the edges — are getting rarer. Modern video generators handle lighting, motion blur, and even crowd scenes with unsettling competence. Your brain is wired to trust moving images in a way it doesn’t trust still photos, which makes the problem worse.
There’s also a cat-and-mouse dynamic at play. As detection tools improve, generation models adapt. It’s not a solved problem, and anyone claiming their detector is 100% accurate is selling something.
What Current Detection Tools Can Actually Do
A few options exist today, each with real limitations:
Content credentials (C2PA) — Some AI platforms embed invisible metadata tags into the videos they produce. Tools that check for these credentials can confirm, definitively, that a clip came from a specific AI system. The catch: only a handful of platforms participate, and nothing stops someone from stripping the metadata or using a platform that doesn’t tag its output at all.
Computer vision models — Services like SightEngine use frame-by-frame analysis to flag regions that match known patterns of AI generation: unnatural grain structure, inconsistent motion vectors, artifacts around hair and edges. These work surprisingly well on current-generation AI video, but they’re tuned to what AI video looks like now. They’ll need constant retraining.
Multimodal LLMs — Models like Gemini can actually watch a video and reason about it. In theory, a model that understands physics, object permanence, and human behavior should be able to notice when something violates those rules. In practice, today’s models are inconsistent at best. They’ll confidently declare a clearly synthetic clip to be real, then hedge on something ambiguous. Don’t rely on them as your primary signal.
The most reliable approach right now is to combine signals: check for content credentials, run a computer vision scan, and use a reasoning model as a tiebreaker — not as an oracle.
How to Manually Spot AI Video
While you’re waiting for tools to catch up, your own eyes are still useful. Train them on these:
- Scale inconsistencies. Objects change size relative to each other between cuts, or even within a single shot. A car that’s truck-sized in one frame shrinks two seconds later.
- Physics that almost works. Water, smoke, and fabric are hard to generate correctly. Look for liquid that moves too uniformly or cloth that doesn’t bunch and release naturally.
- Background loop artifacts. Crowd scenes and backgrounds sometimes repeat or pulse in ways that look wrong if you focus on the edges of the frame instead of the subject.
- Audio-visual mismatch. AI video generators and audio generators are often separate systems stitched together. Listen for ambient sound that doesn’t match the environment, or speech that doesn’t quite sync with jaw movement.
- Impossible camera behavior. Real cameras have inertia, focus lag, and lens distortion. AI video sometimes has a frictionless, weightless camera move that no physical rig could produce.
None of these alone is proof. Together, they build a case.
The AGI Footnote Worth Thinking About
Here’s a genuinely interesting wrinkle: if the best AI models available to the public can’t reliably tell AI video from real video, that says something about where we actually are in AI development. Humans with no special training can often spot a fake in seconds. The gap between human visual intuition and current AI detection capability is a useful reality check on some of the bolder claims you’ll hear about AI reaching human-level intelligence.
Computer vision tools trained specifically on AI detection patterns outperform general-purpose language models at this task — which makes sense. A specialized tool beats a generalist when the task is narrow and technical.
What to Do Before You Share
Building a habit is more useful than memorizing a checklist. Before you forward a wild video to anyone:
- Pause for five seconds and watch it again, this time looking at the background and edges, not the main subject.
- Search the key visual claim in the video on a news site. If it’s real and dramatic, someone covered it.
- Run the URL through a computer vision detection API or service if you have access to one.
- If it still feels off, don’t share it. The cost of not sharing a real thing is zero. The cost of spreading convincing slop is real.
The tools will get better. Until they do, skepticism is a feature, not a flaw.