Quick answer
No single sign proves a voice is AI-generated in 2026. Listen for breath and pause timing, flat or misplaced intonation, inconsistent room tone, identical renderings of repeated words, and mispronounced names; then check the file itself. Run the saved clip through the free EyeSift AI voice detector for a private first pass, use the vendor classifier if you suspect a specific generator, look for provenance, and, if the clip is asking you to do something, verify the person through a number you already trust.
Why your ears are no longer enough
A few years ago the tell was obvious: robotic timing, a metallic edge, words glued together without air. Modern voice cloning removed most of that. Current systems model breathing, hesitation, and emotional colour well enough that a short clip of a familiar voice, delivered through a phone speaker or a messaging app, can pass a relative, a colleague, or a journalist on first listen.
The good news is that the same conditions that make a clone convincing also make it fragile. Voice generators are optimised for the words, not for everything around the words. Breath placement, the acoustics of a real room, the way a speaker stumbles on an unfamiliar name, and the container the audio travels in still leak information. The checks below target exactly those edges. None of them is proof on its own; together they give you a defensible reason to escalate or to relax.
One framing to keep in mind throughout: you are not trying to prove the voice is fake. You are trying to decide whether it is safe to act on it. For a voicemail asking for money, the answer to the second question is almost always "not until I have verified the person another way," regardless of what any detector says. Our companion guide on verifying a caller during a voice cloning scam covers that side in detail.
8 manual checks for an AI-generated voice
Listen on headphones, not a phone speaker, and play the clip at least twice. The first pass is for meaning; the second is for everything that is not meaning. For each check, the table also records honestly whether the EyeSift browser screen can flag it, because most of these are things a human hears and a waveform statistic does not.
| # | Check | What to listen or look for | Can EyeSift flag it? |
|---|---|---|---|
| 1 | Breath and pauses | Real speakers inhale before long phrases, run short of air at the end of sentences, and pause unevenly while thinking. Cloned speech often has breaths that are perfectly regular, oddly placed, missing entirely, or pasted in at identical volume every time. | Partly. Silence ratio and micro-variation reflect pause structure, but EyeSift cannot hear a breath as a breath. |
| 2 | Prosody and intonation | Human pitch rises and falls with meaning: questions lift, lists step, emphasis lands on the word that matters. Synthetic voices can drift into a pleasant but flat melody, or place a rising inflection where nobody would. | No. Pitch contour is a frequency-domain property; EyeSift only measures time-domain waveform statistics. |
| 3 | Room tone and reverb | Every real room has a noise floor and a reflection pattern that stays constant across a recording. Generated audio is often unnaturally dry, or the reverb changes between sentences because segments were rendered separately. | Partly. Very low silence-floor energy and flattened dynamics can raise the score, but EyeSift does not model reverb. |
| 4 | Plosives and sibilance | Hard consonants (p, b, t) push air into a microphone and cause small pops; s and sh sounds hiss with a texture that varies by distance. Clones can render every plosive identically or smooth them into nothing. | No. Consonant texture lives in the spectrum, which EyeSift does not analyze. |
| 5 | Background consistency | Traffic, a fan, a TV, or wind should continue smoothly under the voice. If background noise cuts in and out exactly at word boundaries, or disappears during pauses, the voice was likely generated or spliced over a bed. | Partly. Abrupt energy changes affect the silence and micro-variation metrics, but EyeSift cannot separate voice from background. |
| 6 | Unnatural word stress | Names, acronyms, numbers, and homographs (read, lead, live) are where text-to-speech still stumbles. Listen for a familiar name pronounced slightly wrong, stress on the wrong syllable, or a number read digit by digit. | No. This is a linguistic judgment only a listener who knows the speaker can make. |
| 7 | Repeated micro-patterns | When the same word or phrase occurs twice, a human never says it exactly the same way. If two renderings sound identical down to the breath and timing, they may share one generated source. | Partly. Unusually smooth waveform transitions and low zero-crossing variation across a long clip are among the signals that push the score up. |
| 8 | Metadata, provenance, and file history | Where did the file first appear? Was it exported by a known app, forwarded through a messenger, or screen-recorded? Does it carry C2PA credentials or a vendor watermark? Has the sender ever produced the original? | Partly. EyeSift reports container facts (duration, sample rate, channels, estimated bitrate) and flags low-bitrate or re-encoded audio, but it does not cryptographically verify C2PA manifests or decode watermarks. |
Checks 1–3: the air around the words
Breath, pauses, and room tone are the most useful first-listen signals because they are rarely what a generator was asked to produce. A real person recording a voice note in a kitchen breathes in before "listen, I need you to...", and the refrigerator hum sits under the whole message at one level. A clone rendered sentence by sentence may have a breath that is the same length every time, a silence that drops to digital zero between phrases, or reverb that subtly changes from one line to the next because each line was generated on its own.
Intonation is harder to judge but worth the effort when you know the speaker. Everyone has a melodic habit: a friend who ends statements on a rising note, a manager who flattens the last word of every sentence. Clones borrow timbre well and melody less well. If the voice is right but the tune is wrong, note it.
Checks 4–6: consonants, background, and pronunciation
Plosives and sibilance are microphone physics. A real "p" close to a phone mic pops; a real "s" hisses differently depending on how far the mouth is from the mic and where in the room the speaker turned. Generated speech tends to render these consonants with suspicious uniformity, or to sand them off entirely so nothing ever pops. Background consistency is the same idea applied to the environment: if a passing car fades out precisely when the speaker pauses and returns when they resume, the voice and the environment were not recorded together.
Word stress is the check that most often catches a clone of someone you actually know. Text-to-speech still guesses at names, acronyms, street names, and words with two readings. Your grandmother knows how to say your cousin's name; a model reading a script may not. A single mispronunciation is not proof, since real people misspeak, but a mispronunciation of something the real speaker says every day is a strong reason to verify.
Checks 7–8: repetition and the file itself
Humans never say the same phrase twice identically. If a clip contains a repeated word, greeting, or number, compare the two renderings. Identical timing, identical breath, identical background at both points suggests copy-paste generation. Then leave the audio and look at the file. Which app exported it? Has it been through a messenger that re-encodes to a low-bitrate Opus stream? Is there a C2PA manifest or a vendor watermark to check? Can the sender produce the original? Provenance is covered in depth in our explainer on ElevenLabs watermarks, SynthID, and C2PA for audio; the short version is that a present, valid signal is strong evidence and an absent one is no evidence at all.
What EyeSift can and cannot flag
It is worth being precise about the free tool on this site, because "AI voice detector" covers everything from a browser heuristic to an enterprise fraud platform. The EyeSift AI voice detector decodes the audio locally in your browser, never uploads it, and measures time-domain waveform statistics: duration, sample rate, channels, estimated bitrate, RMS energy, silence ratio, clipping, crest factor, zero-crossing rate, and average sample-to-sample micro-variation. It then combines those into an indicative score with a reliability label and a list of next verification steps.
That is useful as a first pass and as a way to structure your own review. It is not spectral analysis, it is not a trained classifier, and it does not decode SynthID or ElevenLabs watermarks or cryptographically verify C2PA manifests. It can miss a high-quality clone and it can raise the score on a heavily compressed real recording. Every result on that page is indicative, not proof.
| Signal | What EyeSift does | Limit |
|---|---|---|
| Duration, sample rate, channels, estimated bitrate | Reported for every decodable file | Says nothing about who or what produced the voice |
| RMS energy and silence ratio | Flags silence-heavy or unusually dry clips | Cannot tell a quiet room from a generated one |
| Clipping and crest factor | Flags hard limiting and flattened dynamics that reduce trust | Heavy processing is common in real podcasts too |
| Zero-crossing rate and micro-variation | Flags overly smooth transitions that can appear in synthetic speech | High-quality clones can look natural on these metrics |
| Watermarks (SynthID, ElevenLabs) and C2PA manifests | Not decoded or verified | Use the vendor tool or a C2PA verifier |
| Spectral or machine-learning classification | Not performed | Use a provider classifier or an enterprise detector |
A decision workflow that does not depend on one score
- Triage privately. Get the earliest, least compressed copy of the clip you can (WAV, FLAC, or a high-bitrate M4A/MP3 beats a forwarded voice note) and run it through EyeSift. Note the reliability label; a clip under about 15 seconds is weak evidence whatever the score says.
- Run the manual checks. Headphones, two listens, the eight checks above. Write down what you noticed. Specific observations ("the name is stressed on the wrong syllable, and the fridge hum drops out during pauses") are worth more than a percentage.
- Try the provider classifier. If you suspect a specific generator, its own classifier is the strongest scoped signal. The ElevenLabs AI Speech Classifier analyzes the first minute of a sample with no login, is built to detect speech generated with ElevenLabs only, does not reliably classify audio from its Eleven v3 model, and cannot detect other providers' output. A "not ElevenLabs" result therefore clears nothing else.
- Check provenance. Look for C2PA credentials and vendor watermarks where the generator supports them. Google's SynthID audio watermark applies to Lyria and NotebookLM output and can be checked in the Gemini app; the broader SynthID Detector portal was still limited to early testers as of September 2026. Absence means nothing; presence is strong.
- Verify the person. If the clip asks for money, credentials, or action, contact the person on a number or channel you already trusted before the clip arrived. This step beats every detector and is the one official guidance agrees on.
- Escalate when stakes are high. Keep the original file, the source context, and your notes, and hand them to someone qualified.
When to involve a forensic audio expert
Consumer tools, including EyeSift, produce indicative scores rather than calibrated probabilities, and enterprise detectors report accuracy under their own test conditions. Vendors such as Resemble (Detect) and Pindrop (Pulse) publish high self-reported figures for their platforms, but those products are built for contact centers, meeting platforms, and fraud teams under contract, not for a single contested clip on a phone. If a recording could affect a court case, an employment decision, a news story, an insurance claim, or someone's safety, the right move is an audio forensics professional who can examine the spectrum, the codec history, and the edit points and testify to the method.
What you can do in the meantime is preserve evidence well: do not re-record the clip from a speaker, do not trim or normalize it, keep the message thread or call log it arrived in, and record when and from whom you received it. That chain of custody is often more valuable to an expert than any score you generated.
For a comparison of the free tools mentioned here and what each one actually covers, see free AI voice detectors in 2026 and the broader best AI voice detectors guide. If you are building a policy for a classroom, a podcast, or a hiring pipeline rather than checking one clip, the workflow guide for teachers, podcasters, and HR covers limits and policy language.
Frequently Asked Questions
Can you tell if a voice is AI-generated just by listening?
Sometimes, but not reliably. Older text-to-speech had obvious robotic timing; current voice clones can pass casual listening, especially in short, compressed clips such as voice notes or phone recordings. Listening for breath, pauses, room tone, word stress, and repeated micro-patterns still helps, but it should be combined with file checks, provider classifiers, provenance signals, and verifying the person through a channel you already trust.
What is the fastest free way to check if a voice is AI?
Run the saved clip through a browser-side screen such as the EyeSift AI voice detector, which checks waveform and file-quality signals without uploading the audio, then run the ElevenLabs AI Speech Classifier if you suspect an ElevenLabs voice. Both are free and need no login. Treat every result as a screening signal, not proof.
Can EyeSift detect ElevenLabs or OpenAI voices?
Not specifically. EyeSift is generator-agnostic: it measures duration, bitrate, energy, silence, clipping, crest factor, zero-crossing, and micro-variation, and it can miss high-quality synthetic speech from any vendor. It does not decode watermarks or attribute a clip to one provider. For ElevenLabs attribution, use the ElevenLabs AI Speech Classifier, which is built only for ElevenLabs audio.
Does a missing watermark or missing C2PA credential mean the voice is real?
No. Watermarks such as SynthID apply only to supported generators (for audio, Google names Lyria and NotebookLM), and C2PA credentials survive only when the file was created by a supporting tool and never stripped by re-encoding, messaging apps, or screen recording. Absence of a signal tells you nothing either way.
How long should a clip be to check for AI voice signals?
Longer is better. EyeSift labels clips under about 15 seconds as low reliability and treats very short clips as weak evidence. Provider classifiers also work on limited windows, for example the ElevenLabs classifier analyzes the first minute of a sample. Use the earliest, least compressed version of the recording you can obtain.
When should I involve an audio forensics expert?
When the answer affects money, employment, legal proceedings, safety, or publication. Consumer detectors and heuristic screens produce indicative scores, not calibrated probabilities, and they degrade on compressed or re-recorded audio. Keep the original file, the source context, and your notes, and escalate rather than acting on a single score.
Sources checked September 17, 2026
- ElevenLabs AI Speech Classifier, scope, first-minute analysis, Eleven v3 limitation.
- Google DeepMind SynthID, audio coverage (Lyria, NotebookLM), Gemini check, Detector portal status.
- C2PA Specification 2.2, audio container support (RIFF/WAV, ID3/MP3, AIFF).
Screen a Voice Clip Now: Free & Private
EyeSift's AI voice detector runs entirely in your browser. No signup, no email, no server upload. Indicative results, not proof.
Open the AI Voice DetectorRelated Articles
AI Voice Cloning Scams: How to Verify a Caller
FTC and FBI guidance, family code words, and a 60-second verification checklist.
ProvenanceDoes ElevenLabs Watermark Its Audio?
SynthID, C2PA, and what a missing signal does and does not mean.
ComparisonBest AI Voice Detectors 2026
Free browser tools, provider classifiers, and enterprise platforms compared.