How to Spot AI-Generated Video: The Second-Look Checklist
How to spot AI generated video: the QA checklist our lab runs in reverse — hands, teeth, text melt, physics, audio drift — and where detection stops working.
We spend our working hours making AI video that survives inspection, which puts us in a reasonable position to write the manual for the other side. This post is our internal QA checklist turned inside out: everything we verify before a clip ships is, read in reverse, a guide to how to spot AI generated video. The tells below are specific and checkable in under a minute — and, to be honest up front, every one of them has a shorter life expectancy than you’d probably like.
Detection is the second look, run in reverse
Our whole standard — the reason for the company name — is that realistic AI output survives a second look. Modern video models clear the first glance easily now; a person scrolling a feed at speed will not catch a well-made generation. The second look is different territory. Humans are ruthless second-look detectors, and the checklist below is that instinct made systematic.
It’s also, nearly word for word, the QA pass we run before our own AI videos ship. That symmetry matters, and we’ll return to what it implies near the end: any tell that’s publishable is also fixable, by exactly the people most motivated to fix it.
Two ground rules before the list.
Watch it twice, then scrub. Once at natural speed on a phone, once full-screen on the biggest display you have. Then pause and step frame by frame through any moment of physical contact — a hand grabbing a cup, a person sitting down. Video artifacts live between frames; scrubbing is the single highest-yield move in detection.
Tells are asymmetric. A genuine snag is strong evidence of generation. A clean pass is weak evidence of nothing. Keep that asymmetry in mind for everything below.
How to spot AI generated video: the second-look checklist
Here is the checklist in scannable form; the sections after it explain what each row actually looks like in practice.
| # | Where to look | Real footage | The generated tell |
|---|---|---|---|
| 1 | Hands in motion | Fingers keep their count; grips deform objects | Fingers merge or spawn mid-gesture; hands hover on contact |
| 2 | Teeth and eyes | Individual teeth; irregular but complete blinks | A single white band that shimmers; metronomic or missing blinks |
| 3 | Hair and fabric edges | Strands stay strands under motion | Edges smear or “boil” against the background |
| 4 | Text and logos | Legible from any angle, stable between frames | Background signage melts; letters rearrange |
| 5 | Physics | Weight, impact, splash; shadows stay attached | Objects settle without landing; liquids move like gel |
| 6 | Audio | Room tone; sync holds to the last frame | Studio-clean voice in a kitchen; lip-sync drift compounding |
| 7 | Continuity | Jewelry, patterns, background people persist | Details change across cuts; extras loop |
Hands, teeth, and hair: the biology checks
The six-finger hand is mostly a museum piece. As of mid-2026, frontier models render a static hand about as well as a camera does; the tell has migrated to hands doing things. Watch the moment of contact: a real grip compresses the cup, whitens the knuckles, casts a coherent shadow. Generated grips tend to hover — contact without pressure — and fingers still occasionally merge or spawn mid-gesture, which is exactly why you scrub those frames.
Teeth remain harder than they look, because they’re small, high-contrast, and repetitive. Generated teeth often read as one continuous white band, and the band shimmers between frames — pause twice, half a second apart, and compare. For eyes, check the blink cadence (real blinking is irregular but always completes) and the reflections, which should agree with the room the person is supposedly in.
Hair and fabric edges are where models quietly spend the least effort. Under motion, individual strands smear into the background or “boil” — a fine, crawling instability you’ll recognize immediately once you’ve seen it once.
Text: the cheapest check, with one big caveat
Freeze any frame that contains background text — signage, a product label, a screen. Generated text still melts under scrutiny more often than any other element: letterforms that almost spell the word, then rearrange two frames later.
The caveat: skilled makers know this, so they either keep text out of frame or start from a real photograph and animate it, which hands the model a label it cannot misspell. Melted text proves generation; pristine text proves nothing. There’s that asymmetry again.
Physics: weight is expensive to fake
Watch anything land, pour, or collide. Real objects arrive with impact — a thud, a bounce, a settle. Generated objects tend to simply arrive, decelerating into place as if the last inch of gravity were optional. Liquids are worse: pours that move like gel, splashes that resolve too neatly. Shadows and reflections deserve a scrub too — they should move in lockstep with their owners, and generated ones sometimes lag or detach.
One flag that isn’t about physics at all: convenient degradation. Heavy compression, sudden grain, the filmed-off-a-screen look — these can hide every tell in this section, and in 2026 that is frequently the point. Potato quality on a consequential video is itself a reason for suspicion.
Audio: let the last two seconds testify
Lip-sync drift compounds over a clip, so the end of the video tells the truth the beginning hides. Watch the final two seconds mouth-only. A sync that started perfect and slid by tens of milliseconds is visible there, even if you can’t name what’s wrong.
Then close your eyes and just listen. Real rooms have tone — HVAC, echo, the faint hum of a refrigerator — and real speakers breathe in ways that interact with the room. A studio-clean voice over a kitchen scene is a mismatch a surprising number of generated videos still ship with. We keep a full checklist for synthetic voice, and most of it doubles as detection.
The tells that come from the workflow, not the model
Some of the strongest signals aren’t rendering defects at all — they’re fingerprints of how AI video gets assembled.
Suspicious editing rhythm. Generation is most coherent in the 5–10 second range, so AI video is typically built from short beats joined by jump cuts — that is literally the pipeline. Jump cuts are normal in real creator video too, but watch where they land. A cut that arrives exactly when a hand reaches for an object, every single time, is someone editing around their artifacts.
Nothing incidental. Real footage is full of accidents: a stranger crossing the background, a phone buzzing, a car horn. Generated scenes are curated by construction — everything in frame is on-topic. An environment with zero unplanned events is a soft tell that gets stronger the longer the video runs.
Identity drift across a channel. For a recurring “creator,” compare today’s video with one from a few months back. Keeping a synthetic person consistent takes deliberate, ongoing effort, and channels that skip it show drift — a jawline that changed, a voice half a tone off, wardrobe details that don’t persist. Real people age; they don’t oscillate.
Where detection loses
We’d be selling you something dishonest if we ended at the checklist, so here is the uncomfortable part.
Every tell above is a current-model defect, and defects get patched. Published tells become training targets. Hands went from a running joke to mostly solved in roughly two years; text is visibly improving with each model generation. We revise our internal version of this checklist quarterly, and rows fall off it every time. Treat this post as dated the day it was published — that isn’t modesty, it’s the mechanism.
The good stuff has already passed this checklist. Anyone producing AI video seriously runs second-look QA before shipping, which means the clips that reach you have survived the exact checks you’d apply. That’s not a paradox; it’s an arms race in which the attacker moves last. It’s also the other direction of the same craft — the making side is the entire curriculum of Realistic AI Club, and none of it is secret.
Automated detectors are weakest exactly where it matters. Classifier-style detection tools exist, and they’re fine as a tiebreaker. But they misfire in both directions: compressed, low-light real footage gets flagged, and careful generations pass. Their misses cluster on precisely the high-effort fakes you’d most want caught. Use them as one input, never as a verdict.
What ages better: provenance and context. Pixel forensics has a shelf life; source questions don’t. Who posted this, and what else have they posted? Does an independent second source show the same event? Is there provenance metadata — content credentials attached at creation — and does the platform’s synthetic-media label apply? (The labeling side is a genuine compliance regime now; we map it in the disclosure rules post.) For anything consequential — a public figure saying something explosive, “evidence” arriving at a convenient moment — provenance beats pixels, and it will keep beating them long after this checklist is obsolete.
The takeaway
If you remember four things about how to spot AI generated video, make them these:
- Watch twice, then scrub. Feed speed first, full-screen second, frame by frame through any moment of physical contact.
- Check the big five in order: hands in motion, teeth and eye shimmer, hair and fabric edges, background text, the physics of weight and liquid.
- Let the last two seconds of audio testify. Sync drift compounds; the end of the clip is where it confesses.
- Weight provenance over pixels for anything that matters. Tells prove generation; their absence proves nothing. Source, corroboration, and content credentials age far better than any artifact list — including this one.
And if the flip side interests you — making video that passes these checks instead of failing them — that skill is process, not luck. It’s what we mean by realistic AI, it’s documented step by step in our production pipeline, and it’s what Realistic AI Club teaches for ten dollars a month.
FAQ / Common questions
How can you tell if a video is AI-generated?
Watch it twice — once at normal speed, once full-screen — then scrub frame by frame anywhere a hand touches an object. Check finger count and grip, teeth and eye shimmer between frames, hair and fabric edges under motion, any text in the background, the physics of liquids and landings, and lip-sync in the final two seconds, where drift accumulates. A snag on any of these strongly suggests generation; a clean pass proves nothing either way.
What are the most common signs of AI-generated video?
As of mid-2026, the most reliable tells are hands interacting with objects (contact without pressure, fingers merging in motion), teeth that read as one shimmering band, hair edges smearing into the background, background text that melts or rearranges between frames, liquids that move like gel, and lip-sync drift that compounds toward the end of the clip. Static faces and scenery are largely solved; motion and physics are where current models still fail.
Are there apps or tools that detect AI-generated videos?
Automated detectors exist, but treat their verdicts as weak evidence. They misfire in both directions: heavily compressed real footage gets flagged, while careful, high-effort generations pass. Provenance signals are sturdier when present — content credentials embedded at creation, platform AI labels, invisible watermarks some providers add — though re-encoding and screen capture can strip them. For anything consequential, checking who posted the video and whether an independent source corroborates it beats any pixel-level tool.
Why is AI-generated video so hard to detect now?
Because every published tell becomes a training target. The visual defects that made detection easy — extra fingers, warped faces, melted logos — get fixed within a model generation or two of becoming well known, and skilled creators run detection checklists as quality control before publishing, shipping only clips that pass. Detection by artifact is a moving standard with a shelf life; detection by provenance and context ages much better.