Back to Blog
AI Guides

How to Tell If a Video Is AI-Generated (2026 Guide)

Ravi Prajapati

Author

Ravi Prajapati

August 13, 2026
/api/uploads/1786649799191-how-to-tell-if-a-video-is-ai-generated.webp

There is no single trick that spots every AI video. Here is the layered check that works: provenance, Content Credentials, watermarks, reverse search and more.

A clip lands in your WhatsApp group. A politician says something they would never say. A doctor you have never heard of endorses a supplement. A disaster unfolds in a city you recognise. It looks real. It also feels wrong.

The question you want answered is not academic. You want to know, in the next few minutes, whether to believe it or forward it.

Here is the uncomfortable part. There is no single visual trick that identifies every AI-generated video, and anyone selling you one is selling you a 2023 answer to a 2026 problem. What does work is a layered check, where each step narrows the possibilities and no single step is treated as proof.

How can you tell if a video is AI-generated?

Start with provenance, not pixels. Check who posted it first, look for a platform AI label, then upload the file to a Content Credentials checker to see its signed creation history. Replay slowly and inspect faces, hands, text and shadows. Compare lip movement with audio. Finally, screenshot a frame and reverse-search it to find the original.

The 12-step check, in order

Work down this list. Most videos are resolved in the first four steps.

  1. Check who originally posted the video, not who sent it to you

  2. Look for platform AI labels or a creator disclosure

  3. Check Content Credentials or other provenance information

  4. Watch it again slowly, at reduced speed if the player allows

  5. Examine faces, hands, objects, on-screen text and backgrounds

  6. Check lighting, shadows and reflections for physical consistency

  7. Compare lip movement against the audio

  8. Listen for unnatural speech, missing breaths and mismatched room tone

  9. Extract a distinctive keyframe and reverse-search it

  10. Inspect metadata if you have the original file

  11. Run an AI detector, treating the result as one input

  12. Cross-check the underlying event with sources that do their own reporting

The order matters. Steps 1 through 3 answer more questions in less time than steps 4 through 8, which is the opposite of how most guides are written.

What counts as an AI-generated video?

The phrase covers several different things, and conflating them causes bad calls in both directions.

Fully AI-generated video is created from a text prompt or reference image with no camera involved. Nothing in the frame was ever photographed.

AI-edited video is real footage altered with generative tools: an extended background, a removed object, an upscaled resolution, a repaired audio track. The underlying event happened.

Face swaps replace one person's face with another's across frames. Voice cloning synthesises a specific person's voice from samples of their speech. Lip-sync manipulation reshapes a real person's mouth to match audio they never spoke.

Deepfake is the narrower term for synthetic media that depicts a real, identifiable person saying or doing something they did not. Every deepfake is AI-generated media. Most AI-generated video is not a deepfake.

Traditional CGI, compositing and video editing predate generative AI by decades and remain in constant use. A green screen is not a deepfake. Neither is a colour grade, a stabilisation pass or a speed ramp.

The distinction that trips up even careful viewers is dubbing versus lip-sync. A video dubbed into another language is authentic footage with replaced audio, and the lips will visibly not match. A lip-synced manipulation is inauthentic footage where the lips have been reshaped to match, and the sync will look almost right. Researchers building the Deepfake-Eval-2024 benchmark listed exactly this confusion among their most common labelling errors, and they were trained analysts working without a deadline.

Why AI video became harder to spot

Text-to-video models improved sharply through 2024 and 2025 on precisely the failure modes older detection advice depended on. Human gait looks plausible. Faces hold their identity across cuts. Objects fall at roughly the right speed. Camera motion includes handheld shake and rack focus. Lip synchronisation ships as a default rather than a bolt-on. Clip lengths stretched from seconds to a minute or more, removing the tell of suspiciously short footage.

Count fingers and watch for blinking, the two pieces of advice that dominated coverage in 2023, now generate more false accusations than correct identifications. Hands are still imperfect, but imperfect in ways that also describe motion blur on a real phone camera. Visual inspection has moved from being the primary method to being a supporting one.

Start with the source, because context breaks more fakes than pixels do

The most common deception in a viral video is not synthesis at all. It is authentic footage carrying a false caption: real protest video relabelled as a different country, real storm damage relabelled as this week, real speech clipped to remove the sentence that changes its meaning. No amount of frame-by-frame inspection will catch this, because there is nothing wrong with the frames.

So ask source questions first.

Who posted it first? The account that sent it to you is rarely the origin. Search the caption text, a distinctive quoted phrase or a visible landmark and look for earlier posts of the same footage.

Is the account credible? Age, posting history and topical consistency matter more than follower count. An account created recently that posts nothing but high-emotion clips from unrelated events is an aggregator at best.

Does the date match the claim? Check whether weather, foliage, clothing, vehicle models and signage fit the claimed date and place. Snow in a clip captioned as a heatwave is a faster debunk than any detector.

Is anyone who does original reporting covering it? A dramatic event from three hours ago that appears nowhere except one account is a strong signal about the claim rather than the file.

Are there other angles? Real public events generate multiple recordings from different positions. Synthetic events generate exactly one.

One warning worth internalising: do not ask a chatbot to verify a video for you. When the Atlantic Council's Digital Forensic Research Lab analysed roughly 130,000 posts involving Grok during the Israel-Iran conflict, it found the assistant struggled to authenticate AI-generated media, gave contradictory answers to near-identical prompts within minutes, and told users that AI-generated footage of building collapses appeared to be real. Chatbots reason fluently over text about an event. They are not running forensic analysis on your pixels.

Platform AI labels: useful, incomplete, and easy to misread

Every major platform now labels some AI content, and knowing what each system does and does not cover tells you how much weight to give a missing label.

YouTube has required creators to disclose realistic AI-generated or altered content since 2024. In an update published in May 2026, it moved that label somewhere you will actually see it: directly below the player on long-form video, and as an overlay on Shorts. The same update introduced automatic labelling, so a creator who skips disclosure may still get labelled when YouTube's systems detect significant photorealistic AI use. Two categories of disclosure are permanent and cannot be removed by the creator: content made with YouTube's own AI tools, and content carrying C2PA metadata indicating it was fully generative.

Read YouTube's disclosure requirements closely, though, because the exemptions are broad. Creators need not disclose beauty filters, colour and lighting adjustment, background blur, caption generation, upscaling, audio repair, or cloning their own voice for voiceovers and dubs. A video can be substantially touched by AI tools, carry no label, and be entirely within policy.

TikTok became the first video platform to implement Content Credentials, letting it auto-label AI content made elsewhere as soon as the file arrives carrying the right metadata. Meta applies an "AI info" label across Facebook and Instagram using C2PA and IPTC indicators alongside creator disclosure.

All three share a blind spot: they label reliably when a file carries machine-readable provenance and unreliably when it does not. Someone who screen-records a generated clip and re-uploads it has stripped that signal entirely.

The regulatory picture is also shifting. Article 50 of the EU AI Act, which requires providers to mark AI-generated output in a machine-readable format and requires deployers of deepfakes to disclose them, became applicable on 2 August 2026. Content generated before that date does not need retroactive labelling, and a limited grace period runs to 2 December 2026 for the marking obligation on systems already on the market. Expect labelling coverage to improve unevenly over the next year rather than all at once.

Content Credentials and C2PA, explained without the jargon

This is the step most people skip, and it is often the most decisive.

Content Credentials are tamper-evident records attached to a media file describing how it was made. The technical standard behind them comes from the Coalition for Content Provenance and Authenticity, an industry body whose steering members include Adobe, Google, Microsoft, OpenAI, Sony, BBC and Truepic. Think of it as a nutrition label with a cryptographic signature: it can record that a file was captured by a specific camera model, opened in a specific editor, extended with a specific generative tool, and exported at a specific time.

The signature is what makes it more than a claim. If the credential says a clip came out of a particular AI video generator, and the signature validates, that is verifiable evidence about origin rather than a guess about texture and lighting. A visual artifact tells you something looks odd. A valid credential tells you what the file's own creation chain asserts.

How to check one. Download the video if you can, then open the Content Credentials verification tool and upload it. The check runs in your browser and takes under a minute. If a credential is present and valid, you will see the creation and edit history. If it has been broken or stripped, the tool will say so.

Google has extended provenance checking into its own products, bringing Content Credentials into Search and the "About this image" feature, and its Gemini apps can now report on Content Credentials in a file you upload.

Now the limits, which matter as much as the capability.

A missing credential proves nothing. Most cameras do not sign their output, most platforms re-encode uploads in ways that destroy the manifest, and screenshots, screen recordings and forwards strip provenance routinely. The absence of Content Credentials is the normal state of internet video, not a red flag.

A present credential proves origin, not truth. C2PA records what a signer asserts about a file's history. It cannot tell you whether the footage is captioned honestly, filmed where claimed, or cut to mislead.

There is also a false-positive path worth knowing. Editing software embeds credentials whenever an AI-assisted feature is used, including generative fill, AI denoise or AI masking. A genuine photograph that received one AI-assisted retouch can inherit provenance markers describing AI involvement and get labelled on upload. That label is technically accurate and practically misleading.

AI watermarks and what SynthID can prove

Watermarking works differently from provenance. Rather than attaching metadata that can be stripped, it embeds an imperceptible signal into the pixels or audio itself, designed to survive compression, cropping and re-encoding.

Google's SynthID is the most accessible example for ordinary users. You can upload a video to the Gemini app and ask whether it was generated using Google AI. Gemini scans both the visual and audio tracks and reports which segments carry the watermark, with output along the lines of a watermark detected in the audio between ten and twenty seconds but not in the visuals. That segment-level detail is genuinely useful, because it distinguishes a fully synthetic clip from a real one with an AI-generated audio track spliced in.

The constraints are specific. Per Google's own documentation, uploads are limited to 100 MB and videos must run under 90 seconds. More importantly, Gemini currently recognises SynthID from Google's AI tools only. A clip made with a competing generator will return nothing.

OpenAI takes a layered approach with Sora, describing in its launch documentation a visible moving watermark on videos downloaded from its apps, C2PA metadata embedded in output files, and internal reverse-search tooling that only OpenAI operates.

Hold this line firmly: a detected watermark is strong evidence a video is AI-generated. A watermark that fails to appear is evidence of nothing. Different generators use different systems, some use none, and processing pipelines can degrade verification. Visible watermarks specifically are trivially removable by cropping.

Video metadata: sometimes useful, frequently misleading

Metadata embedded in a video file can record creation and modification timestamps, device make and model, camera settings, GPS coordinates and the software used to encode or edit it. When you have an original file, downloaded directly from the source rather than re-shared through three apps, that information is worth checking. Right-click and view properties on Windows, or Get Info on macOS, gets you the basics.

Two cautions, and both are important enough that a lot of published advice on this topic is wrong.

Social platforms strip metadata as a matter of routine, partly for privacy and partly as a side effect of re-encoding. Anything you save from a feed will almost certainly have none. Missing metadata is therefore the default condition of shared video and says nothing about whether AI was involved.

Metadata is also trivially editable. Free tools rewrite timestamps and device fields in seconds. Treat present metadata as a lead worth following, never as proof.

Visual signs that still deserve attention

None of these prove anything alone. Look for clusters, and weigh them against compression and lighting as alternative explanations.

Faces remain the richest source of tells, though the tells have moved. Rather than obvious distortion, watch for identity drift: a jawline, hairline, mole or ear that shifts slightly between frames as the head turns. Watch the eyes for a gaze that does not track what the person is looking at, and the mouth for teeth that change count or shape mid-sentence.

Hands are still imperfect but deserve less weight than they once got. Fingers passing through objects or gripping nothing are more telling than an extra finger in a blurred frame.

Object permanence is a strong signal. Something in the background appearing, vanishing or merging into another object across a continuous shot has no natural explanation, and neither do clothing details that change: a stripe becoming a check, a bag switching shoulders, jewellery that arrives halfway through.

Physics deserves specific attention, because water, smoke, fabric and crowds are computationally expensive to get right. Watch how debris falls, how a flag moves, how people at the edge of a crowd behave.

Reflections and shadows encode the whole scene's geometry, which makes them hard to fake consistently. Look for a shadow pointing against the visible light source, a window reflecting something that is not there, glasses reflecting a different room, or lighting that changes intensity mid-shot with no visible cause.

Text is the cheapest test available. Pause on any frame containing a sign, subtitle, phone screen, jersey number or logo. Generated text tends to be almost right and unstable between frames, and distorted logos are common. The same instinct covers spatial coherence: walls that do not meet plausibly, a road whose perspective bends, floors that do not line up.

Audio clues, and why voice alone is not enough

Lip-sync mismatch is worth checking first, with the dubbing caveat above firmly in mind. Watch the ends of sentences, where drift tends to accumulate.

In the speech itself, listen for absent breath sounds, missing filler words and hesitations, an unusually even cadence, and emotional flatness that does not match the speaker's face. Voice consistency across a clip matters too, since a timbre that shifts subtly at cut points can indicate splicing.

Background audio is often the weakest link. Real recordings capture room tone, traffic, wind, and the acoustic signature of the space. Audio that is too clean for a busy street, or whose room tone changes abruptly without a visible cut, deserves scrutiny.

Do not judge a clip by voice alone. Modern voice cloning replicates timbre, accent and cadence closely enough that experienced listeners are fooled, which is why the FBI's Internet Crime Complaint Center warned in 2025 about AI-generated voice messages impersonating senior US officials and advised verifying through independently obtained contact details rather than replying to the message.

Reverse search: the step that finds the original

Reverse image search on a video frame is the highest-yield technical step available to a non-specialist, and almost nobody does it.

Pause the video on a frame that is visually distinctive and not motion-blurred: a wide shot with a landmark, a recognisable storefront, a clear face. Screenshot two or three such frames from different points in the clip. Run each through reverse image search, and search separately for the claimed event in plain language.

Compare upload dates. If the earliest copy you can find predates the claimed event, you have a recontextualised video rather than a generated one, and you have your answer.

The InVID-WeVerify verification plugin automates the tedious parts. It is a free browser extension maintained by AFP's media lab, it extracts keyframes from a video automatically, and it fans a single query out across multiple reverse image search engines at once. It is the tool professional fact-checkers reach for, and it costs nothing.

Reverse search has real gaps. It misses very recent uploads, private and closed-platform content, and heavily cropped or filtered versions. A miss is not a result.

AI video detection tools: supporting evidence, not verdicts

Detectors are worth running. They are not worth trusting on their own, and the research on this is unambiguous.

The Deepfake-Eval-2024 benchmark, built from 45 hours of video and 56.5 hours of audio that real users flagged as suspicious across 88 websites and 52 languages, tested detectors against material actually circulating rather than curated lab datasets. Open-source models that scored near-perfectly on academic benchmarks saw AUC fall by around 50 percent for video, 48 percent for audio and 45 percent for images. Commercial detectors performed better, with the best video model reaching 78 percent accuracy, but no commercial model tested reached 90 percent, which the authors estimated as the floor for trained human forensic analysts.

The failure patterns matter because they describe most social media video. Accuracy dropped 31 percent on videos where only some faces were manipulated and the rest were real, and around 17 percent on manipulations outside the face. Audio detectors lost roughly 18 percent accuracy when background music was present, with false negatives rising sharply, and music is standard in short-form video. Compression, screen recording, cropping and platform transcoding degrade performance further, and every one of those happens to a clip on its way to your feed.

Four categories are worth knowing: AI media detectors score the likelihood that content is synthetic and work as a tiebreaker; provenance verification checks Content Credentials and is stronger evidence when present; watermark verification checks for signals like SynthID and is conclusive when positive, uninformative when negative; and keyframe extraction with reverse image search finds the original, which is often the decisive step.

One practical note on tool churn. TrueMedia.org, a free nonprofit detector still recommended in articles published this year, shut down in January 2025 and open-sourced its models. Detection tools appear and disappear constantly. Check that a tool is live before relying on a recommendation, including this one.

The 60-second AI video check

When you need an answer before you decide whether to forward something:

0 to 10 seconds. Read the uploader and the caption. Is this the origin account or a re-share? Does the caption make a specific, checkable claim?

10 to 20 seconds. Look for a platform AI label near the player, in the Shorts overlay, or in the expanded description.

20 to 35 seconds. Replay at reduced speed. Scan faces, hands, on-screen text, shadows and background objects. Look for two or more independent oddities, not one.

35 to 45 seconds. Watch the mouth against the audio. Listen for breath, room tone and whether the sound matches the visible environment.

45 to 60 seconds. Screenshot a distinctive frame, reverse-search it, and search the claimed event separately.

If sixty seconds does not settle it, the honest position is "unverified," and unverified content does not get forwarded.

The ReadInBrief verification confidence framework

Different signals carry different evidential weight. This is how to rank them.

Signal

What It Can Tell You

Reliability

Limitation

Valid Content Credentials

Signed creation and edit history, including whether AI tools were used

High when present and validating

Rarely present; stripped by re-encoding; describes origin, not truthfulness of the claim

Detected watermark (e.g. SynthID)

That a supported generator produced or edited part of the file

High when positive

Only covers participating generators; a negative result is uninformative

Original-source verification

Whether the event occurred and the footage depicts it

High when a primary source is found

Slow; not always possible for recent or local events

Reverse image search on keyframes

Whether the footage existed earlier, and in what context

Moderate to high when a match is found

Misses recent, private, cropped or heavily filtered uploads

Platform AI label

That the creator disclosed AI use, or the platform detected it

Moderate when present

Depends on self-disclosure or intact metadata; many AI edits are exempt from disclosure

Commercial AI detector result

A probability estimate, useful as a tiebreaker

Moderate; best video models tested below 80% on in-the-wild content

Degrades on compression, music, partial manipulation and unfamiliar generators

Visual artifacts

That something in the frame is inconsistent

Low to moderate; depends on clustering

Compression, low light and stabilisation mimic AI artifacts

Audio anomalies

That speech or ambience is inconsistent

Low to moderate

Dubbing, noise reduction and phone codecs produce similar effects

Metadata

Device, software and timestamps, when present

Low on shared video

Routinely stripped by platforms; trivially edited

Read this top to bottom. If you resolve a question at the top three rows, you do not need the bottom three.

Four scenarios, four different workflows

A celebrity endorses a product

Skip visual analysis and go to the source. Check the person's verified accounts and the brand's official channels, because real endorsements are cross-posted; that is the point of them. A promotional clip that exists only on an ad account, attached to a payment link, is a scam regardless of how the pixels look.

A political figure says something inflammatory

Find the full speech or press conference first. A large share of political video deception is real footage clipped to remove context rather than synthesis. Check whether venue, date, clothing and backdrop match a documented event, then check the platform label, since this is the category platforms label most aggressively.

Disaster footage from a breaking event

Recontextualised real video dominates here, and speed pressure causes the most sharing errors. Reverse-search a keyframe first, because the likeliest answer is authentic footage of a different event. Check whether weather, foliage, signage and vehicles fit the claimed location, and wait for corroborating angles, which real events produce and synthetic ones do not.

An executive on video asks you to move money

Do not verify the video. Verify the person, through a channel you established previously: a number from your internal directory rather than one supplied in the request, or a colleague sitting near them. The FBI's guidance on impersonation campaigns is consistent on this, and it applies to live video calls as much as recordings. A pre-agreed passphrase for payment authorisation defeats a convincing fake without requiring anyone to become a forensic analyst. Urgency plus an unusual payment channel is the signal here, not pixel quality.

Two rules that prevent most wrong conclusions

Suspicious does not mean AI-generated

Heavy compression produces smearing and blockiness that resemble generative artifacts. Low light produces noise and detail loss. Stabilisation warps edges and background geometry. Beauty filters smooth skin into a plastic finish. Dubbing desynchronises lips. Repeated re-uploads compound all of it. A shaky, blurry, oddly-lit clip from a compressed WhatsApp forward will look somewhat synthetic no matter what created it.

An absence of artifacts does not mean the video is real

Current generators produce clean output, and a short, well-lit, simply-composed clip gives them very little opportunity to fail. Finding nothing wrong means you found nothing wrong. It does not mean nothing is wrong.

Holding both of these at once is what separates useful verification from confident guessing. The goal is not to reach a verdict on every video. It is to know which videos you have actually verified, which ones you have merely failed to disprove, and to treat those two states differently when you decide what to share.

Frequently asked questions

How can I check if a video is AI-generated?

Work from provenance to pixels. Identify the original uploader, check for a platform AI label, upload the file to contentcredentials.org/verify to inspect its signed history, watch it slowly for clustered inconsistencies in faces, hands, text and shadows, compare lip movement to audio, then reverse-search a keyframe to find the earliest version.

Can AI-generated videos be detected?

Sometimes, and less reliably than tool marketing suggests. Watermark and provenance checks give strong evidence when a signal is present. Statistical detectors are weaker: on the Deepfake-Eval-2024 benchmark of real circulating content, the best commercial video model reached 78 percent accuracy and none tested reached 90 percent.

What are the most common signs of an AI-generated video?

Identity drift in a face across frames, hands interacting incorrectly with objects, objects appearing or merging, unstable on-screen text, shadows or reflections that contradict the scene, physics that looks slightly wrong in water, smoke or crowds, and audio with no breath sounds or mismatched room tone. One sign alone proves nothing.

Are AI video detectors accurate?

Accuracy varies widely and drops sharply on real-world content. Detectors lose significant accuracy on compressed video, clips with background music, videos where only part of the frame is manipulated, and output from generators they were not trained on. Use a detector as one input alongside source verification, never as the verdict.

Can metadata tell if a video was made with AI?

Occasionally, if you have an original file that names generation software. Usually not. Social platforms strip metadata during upload and re-encoding, and metadata is easy to edit. Missing metadata is the normal state of shared video and is not evidence of AI generation.

Can Google detect AI-generated videos?

Partially. Gemini can scan an uploaded video for Google's SynthID watermark and report which segments carry it, subject to a 100 MB and 90-second limit, but it only recognises content from Google's own AI tools. Google also surfaces Content Credentials in Search and its "About this image" feature. YouTube separately applies AI labels through creator disclosure, C2PA metadata and automatic detection.

What is the difference between an AI-generated video and a deepfake?

A deepfake depicts a real, identifiable person saying or doing something they did not, usually built from footage of that specific person. AI-generated video is the broader category, covering fully synthetic scenes, invented characters and AI-edited real footage. All deepfakes are AI-generated media; most AI-generated video is not a deepfake.

Can an AI video look completely real?

Yes. Short, well-lit clips with simple composition can be indistinguishable from camera footage on visual inspection alone. This is why provenance, source verification and cross-checking the underlying claim now matter more than looking harder at the frames.

Comments (0)

No comments yet. Be the first to share your thoughts!

Leave a Reply