Synthetic video fails at boundaries and across time. Generators are good at producing a plausible surface and poor at maintaining consistency where one object meets another, and poorer still at keeping arbitrary details identical from one second to the next.
That gives you a practical method: stop staring at the face, and start checking edges and sequences. A single frame can be flawless. Ten seconds of the same shot usually is not.
What are the highest-yield things to check?
Ordered by how often they produce a clear answer, not by how obvious they are.
| Check | What to look for | Why it works |
|---|---|---|
| Hands and fingers | Count fingers; watch fingers crossing each other or gripping | Hands have high articulation and frequent self-occlusion — the hardest geometry to keep coherent |
| Earrings and jewellery | Asymmetry between ears, shape drift, items that disappear behind hair and return different | Small rigid objects have no strong prior; generators improvise them per frame |
| Hair strands against background | Individual strands at the silhouette edge; wisps that merge into the background | Fine alpha detail at a boundary is where matting artefacts concentrate |
| Teeth and tongue in speech | Tooth count and shape changing mid-sentence; tongue geometry | Interior mouth structures are frequently regenerated frame to frame |
| Temporal consistency | Freeze at 0s, 2s, 4s of the same shot; compare moles, scars, tattoos, nail length, fabric pattern | Details with no structural constraint drift because nothing forces them to persist |
| Skin at contact points | Where a hand presses skin, or skin meets fabric | Contact deformation requires physics the model only approximates |
| Text in frame | Signage, packaging, on-screen captions | Text is the classic failure — plausible letterforms, incoherent words |
| Eye reflections | Catchlights that differ between eyes, or reflect nothing consistent | Reflections require a scene model that does not exist |
The temporal consistency row is the strongest single technique and the most under-used. Pause the same shot at several points and compare a specific, arbitrary detail — the exact shape of a mole, the pattern on a bedsheet, the length of a fingernail. A camera records the same object; a generator re-invents it.
Why do single frames fool people now?
Because per-frame quality is the thing generation research optimises hardest, and it has largely been solved for still imagery.
The older tells — waxy skin, uncanny symmetry, mangled ears — are artefacts of earlier models, and checking for them gives false confidence when they are absent. Judging a 2026 generation by 2022 artefacts produces confident wrong answers in both directions.
What has not been solved as thoroughly is persistence: keeping every unconstrained detail identical across hundreds of frames while the camera and subject move. That is where the remaining evidence lives.
How do you check consistency in practice?
- Pick one continuous shot — no cuts. Cuts reset everything and give the generator a free pass.
- Pause at three or four points across that shot.
- Choose an arbitrary, non-structural detail: a mole, a freckle cluster, a tattoo edge, a fabric print, a nail shape, a strand of jewellery.
- Compare it across your pauses. Real detail is stable. Synthetic detail drifts, changes shape, or vanishes and returns different.
- Repeat with a second detail in a different region of the frame. One drifting detail can be compression; two is a pattern.
- Watch a hand cross in front of something and step through it. Occlusion boundaries are where coherence breaks most visibly.
Can face search tell you whether a video is synthetic?
Not directly — and it is important to be precise about what it can and cannot contribute.
Face search compares your frame against indexed faces and returns similarity scores. What that gives you is a provenance signal, not a synthesis verdict:
- A strong match to a real, indexed performer tells you the face resembles a real person. That is compatible with an authentic release and with a face-swap onto other footage. It narrows the question rather than answering it.
- No match tells you almost nothing. Our index holds 241,792 faces with 2,333 linked to a named performer across 106 sites (our index, 2026-08 snapshot), and 2,246 of those 2,333 have their representative image from a single source. Absence is far more often a coverage or frame-quality problem than evidence of synthesis.
The useful combination is a strong facial match plus failures on the temporal checks above. That pattern is consistent with a real performer's likeness applied to generated or substituted footage — which is a serious matter, because it is done without consent.
Boundary worth stating: none of these checks is proof. They shift probability. Anyone offering certainty about synthetic media from visual inspection alone is overstating what the evidence supports.
What about automated detectors?
Treat their output as one input, not a verdict.
Detectors are trained on the outputs of specific generator families and lose accuracy as those families change. They also degrade on the exact material you are likely to be examining: re-encoded, resized, watermarked video, where compression has destroyed the fine statistical traces detectors rely on. Both false positives and false negatives are common.
How much they degrade has been measured, in the research literature at least. An independent blind evaluation run alongside a 2025 computer-vision conference had its strongest entrant score 0.93 AUC on untouched generated video, 0.74 once ordinary post-processing was applied, and roughly 0.62 on footage that had been screen-recorded — the last of those not far from a coin toss. Separate work found detectors scoring above 96% against the generator they were trained on dropping to the mid-fifties against one they were not.
Those are research systems evaluated under controlled conditions. For the consumer-facing "is this AI?" tools you can actually reach, there are no independently reproduced accuracy figures at all — the one evaluation that did include commercial video detectors was contractually obliged to anonymise the vendors, which is the opposite of reproducible. So this page quotes no accuracy figure for any named tool, and a tool that quotes one about itself has told you nothing you can check.
Sources: SAFE Challenge video track (ICCV 2025); GenVidBench and Deepfake-Eval-2024 evaluations; checked 2026-08-03.
Why does this matter beyond curiosity?
Because non-consensual synthetic depictions of real people are a genuine harm, and misidentification compounds it.
If a video appears to depict a real, identifiable person and fails the consistency checks, the responsible response is to treat it as unverified and to avoid redistributing it — not to publish a conclusion. A wrong accusation that footage is synthetic damages a real performer; a wrong assumption that it is authentic spreads a fabricated depiction. Both errors have a person on the other end.
Related questions
- How face search works, step by step
- How to tell an official source from a pirated one
- How to find a video from a single screenshot
- Why can't I find the title I'm looking for?
- Common mistakes people make when searching for a source
If you want the provenance half of the question answered, upload a frame here. It will tell you whether the face matches an indexed performer, with a score — detection runs in your browser and the original never leaves your device.