EyeSift
Workflow GuideSeptember 17, 2026· 12 min read

AI Voice Detection for Teachers, Podcasters and HR: Workflow, Limits, Policies

Reviewed by Sasha Brenner·Last updated September 17, 2026

Three groups now meet synthetic voices in ordinary work: teachers grading recorded speech, podcasters booking guests and selling ad reads, and HR teams interviewing by phone and video. This guide gives each of them a repeatable workflow, an honest account of what a heuristic detector can and cannot show, and policy wording you can adapt.

Quick answer

Use an AI voice detector to decide whether to follow up, never to decide the outcome. Get the original file, screen it privately with the free EyeSift AI voice detector, add a provider classifier or provenance check where one applies, then verify the person through a channel you already trust. Write the policy down before you need it.

What EyeSift actually measures

EyeSift decodes the audio in your browser and reports time-domain waveform measurements: duration, estimated bitrate, RMS energy, silence ratio, clipping ratio, crest factor (peak-to-average dynamics), zero-crossing rate, and sample-to-sample micro-variation. It compares those against ranges typical of natural spoken recordings and returns an indicative score with a reliability label.

  • Runs entirely in the browser; the file is never uploaded.
  • No spectral machine-learning classifier and no vendor fingerprinting.
  • Does not decode SynthID or any other watermark.
  • Does not cryptographically verify C2PA manifests.
  • Output is a triage signal, not proof. Short, compressed, or heavily processed audio lowers reliability, and the tool says so.

The shared six-step workflow

The scenarios differ, but the order of operations does not. Every group below follows the same six steps; what changes is who you call and what you document.

1. Collect the original file

Ask for the earliest export: the WAV or M4A from the recorder, the LMS upload, the raw interview recording. Messenger re-exports, screen recordings, and clips played back through a speaker lose the detail every detector needs.

2. Run a private browser-side triage

Open the EyeSift AI voice detector, load the file, and read the reliability label before the score. The file stays on your device, which matters for student, candidate, and unreleased-episode audio.

3. Use a provider classifier if you suspect a specific generator

If the clip may have come from ElevenLabs, its AI Speech Classifier checks the first minute without a login. It only detects ElevenLabs output and does not reliably classify audio from its Eleven v3 model, so a negative result clears nothing else.

4. Check provenance where a supported signal could exist

Google SynthID currently marks audio from Lyria and NotebookLM, and Gemini can check for that watermark. C2PA Content Credentials can travel with some files. A missing signal is not evidence either way.

5. Verify the human through a known channel

Call back on a number you already have, run a short live oral check, or ask a question only the real person could answer. This step decides the outcome; the detector only decides whether you take it.

6. Document what you did

Record the file source, the reliability label, the follow-up conversation, and the decision. If the matter escalates, this trail is what a dean, an editor, or an employment lawyer will ask for.

For the listening checks that go with step two, see how to tell if a voice is AI-generated. For a comparison of the free tools mentioned in steps two and three, see free AI voice detectors in 2026.

Teachers: recorded speech is now an assessment format

Oral assessments, recorded presentations, language-lab pronunciation submissions, and voice-note homework moved online during the last few years and stayed there. The same period made consumer text-to-speech good enough that a student can paste a script into a voice generator and submit narration they never spoke. The temptation is strongest where the recording is graded on delivery rather than on ideas.

The teacher's advantage is that the student is available. That changes the workflow: the detector exists only to decide which submissions get a two-minute live follow-up. Ask the student to read a paragraph from their own submission, or to answer one question about it, on a call or in class. Real authorship shows up immediately; a generated narration does not survive the question "say that last sentence again, slower."

Limits matter more in education than anywhere else. Non-native accents, careful read-aloud delivery, school-issued laptops with aggressive noise suppression, and low-bitrate LMS transcodes can all push a heuristic score toward "possibly synthetic" with no AI involved. Treat a score as a reason to talk, never as a reason to penalise. The same principle already applies to text: university guidance covered in our AI detector policy roundup says detector output should not be sole evidence, and voice deserves the same restraint. Our classroom triage page and AI detector for teachers guide cover the text side.

Syllabus template (adapt with your institution's policy office)

Recorded speaking assignments in this course must be spoken by you, in your own voice, at the time of recording. Text-to-speech, voice cloning, and AI narration are not permitted unless the assignment says otherwise. If a submission raises questions, you may be asked to discuss it or repeat part of it live; screening tools are used only to decide whether that conversation happens, and no grade or integrity decision will be based on a tool result alone. Keep your original recording file until the course ends.

Podcasters: guests, ad reads, and your own voice

Podcasters face voice questions from three directions. A guest booked over email may not be the person they claim to be. A sponsor may ask whether an ad read was voiced by the host or by a clone of the host. And the host's own voice, published for hundreds of hours, is the easiest voice on the internet to clone and attach to a scam, an endorsement, or a fake episode.

For guest verification, the workflow starts before the recording: confirm the booking through a channel the guest already uses publicly, such as an address listed on their organisation's site, and do a short live video pre-call. If a pre-recorded contribution arrives as a file, screen it with the browser-side detector so the unreleased audio never leaves your machine, then send the guest one specific question about the content by the known channel.

For ad reads and any synthetic voice you use deliberately, disclosure is the norm to adopt now. If you clone your own voice for corrections or translations, say so in the show notes and tell sponsors in writing. That transparency is also your defence: when a fake clip surfaces, you can point to a published policy that says every real episode ships through your feed and your voice is never licensed for third-party reads.

Protecting your own voice is mostly about making impersonation easy to catch rather than impossible. Keep the project files and mastered exports of every episode so you can prove what you did release. Tell your audience and business contacts which channels are real. When a suspicious clip attributed to you appears, screen it, compare it to your archive, and respond through your own feed. If the clip is being used in a scam, the reporting steps in our voice cloning scam guide apply.

Guest agreement template (adapt with counsel)

Both parties confirm that any audio they supply for this episode was spoken by them and not generated or altered with voice-synthesis tools, except where disclosed in writing before publication. The show may screen supplied files and will confirm any concern directly with the guest before publishing. Neither party will use the other's voice, or a synthetic imitation of it, outside this episode without separate written permission.

HR: interviews, references, and the urgent payment call

HR meets synthetic voices in two very different situations. The first is candidate screening: recorded one-way video interviews, phone screens, and voice-note references can be scripted, voiced by a generator, or delivered by a stand-in. The second is fraud against the company: a call that sounds like the CFO, the founder, or a payroll vendor asking for an urgent transfer or a change of bank details.

One more guardrail for HR teams: a voice or text score should never feed a pay decision. If a candidate's interview is flagged, verify the person, not the offer. For the offer itself, anchor on published market data instead of gut feel, the BLS-based metro-by-metro median wage rankings on Salario (OEWS May 2025) are a solid, sourced starting point for benchmarking compensation across locations.

For candidates, the rule is the same as for teachers: the detector only picks who gets a live conversation, and the live conversation is where the decision is made. A phone-quality recording, a non-native accent, or a candidate who rehearsed a scripted answer can all look unusual to a heuristic tool. Rejecting someone on a score is unfair and, in several jurisdictions, invites scrutiny of automated decision-making. Schedule the follow-up, ask an unscripted question, and document that the decision rested on that exchange. Our hiring page and guide to checking AI-written resumes cover the written side of screening.

For the executive-impersonation call, no file-based detector helps in the moment, and the official guidance is behavioural. The FTC's April 8, 2024 alert tells consumers to call the person back on a number they know is theirs. The FBI's May 15, 2025 public service announcement on AI voice impersonation recommends independently finding a phone number for the person and calling to verify, and agreeing a secret word or phrase in advance. Build those two habits into finance procedure: any voice request to move money or change payment details is confirmed by callback on a directory number, and the requester must give the pre-agreed phrase. The FCC's February 8, 2024 ruling that AI-generated voices in robocalls are artificial under the TCPA adds a legal route after the fact, but it does not stop the call from landing.

Interview notice template (adapt with employment counsel)

Recorded and live interviews must be completed by the applicant in their own voice without voice-synthesis or real-time voice-alteration tools. We may screen recordings for signs of synthetic audio; a screening result is used only to decide whether to schedule an additional live conversation and is never the basis of a hiring decision. If you need an accommodation that affects how you record or speak, tell us and we will arrange an alternative.

Detector output versus acceptable evidence

SignalWhat it isWhat it can support
Detector percentage from a heuristic toolA reason to look closerNot evidence of authorship, fraud, or misconduct
Provider classifier positive result (e.g. ElevenLabs)Strong signal inside that vendor’s scopeStill needs context: who uploaded it, when, and why
Detected SynthID or valid C2PA credentialSigned or watermarked origin signalAbsence proves nothing; presence does not show intent
Live oral follow-up or callback on a known numberDirect verification of the personThe evidence that should carry a decision
Original file plus recording metadata and draftsProcess evidenceBest combined with the conversation above

Where every voice detector fails

The failure modes are shared across free and enterprise tools, so they belong in every policy. Clips under about fifteen seconds give too little signal. Audio re-encoded by WhatsApp, Teams, or a learning platform loses the fine detail that separates natural and generated speech. New generators narrow the statistical gap each release, and provider classifiers only cover their own output: the ElevenLabs classifier analyses the first minute, requires no login, but by ElevenLabs' own statement does not reliably classify its Eleven v3 model and cannot detect other providers. Watermarks are narrower still: SynthID for audio covers Google's Lyria and NotebookLM output, is checkable through Gemini, and its Detector portal is still on a waitlist for journalists and researchers. Whether the biggest voice vendor marks its files at all is a separate question, answered in does ElevenLabs watermark its audio?

The practical conclusion for all three audiences is identical. Heuristic scores are indicative, not proof. Nothing punitive should follow from a score alone. The strongest evidence you can collect is still a live human interaction through a channel you already trust, and the detector's only job is to tell you when that interaction is worth scheduling. For the broader tool landscape, our best AI voice detectors comparison sorts free, provider, and enterprise options by use case, and the voice cloning scam statistics page shows why the fraud side of this is not hypothetical.

Frequently Asked Questions

Can a teacher use an AI voice detector as evidence of cheating?

Not on its own. EyeSift and similar tools produce heuristic, indicative scores, not proof. A flagged recording is a reason to ask the student follow-up questions, request the original file, or run a short live oral check. Any academic-integrity action should rest on the conversation, drafts, and process evidence, not on a detector percentage.

Can HR reject a candidate based on an AI voice score?

That is a poor practice and, depending on jurisdiction, may create legal exposure. Heuristic detectors can misread compressed phone audio, non-native accents, and processed recordings as synthetic. Use a score only to decide whether to schedule a live follow-up conversation, and document that the hiring decision was based on that live interaction. Review any automated-screening policy with employment counsel.

How can podcasters protect their voice from cloning?

You cannot stop a determined clone, but you can make impersonation easier to catch. Publish through channels listeners already trust, keep original project files and mastered exports, tell sponsors and guests that you will confirm ad reads and bookings through a known email or phone number, and screen suspicious clips attributed to you with a browser-side check before responding publicly.

Does the audio leave my device when I use EyeSift?

No. EyeSift decodes and measures the file in your browser using the Web Audio API. The audio is not uploaded, stored, or shared, which matters when the clip is a student submission, a candidate interview, or an unreleased episode.

What should I do if a caller sounds like my manager and asks for an urgent payment?

Hang up and call the person back on a number you already have, or confirm through a second channel such as a known internal chat or in person. The FTC and the FBI both recommend independent verification for urgent money or identity requests, and the FBI suggests agreeing a secret word or phrase in advance. Do not rely on a detector during a live call.

Which audio formats give the most reliable screening result?

WAV or FLAC originals of at least 15 seconds of clear speech give the browser the most waveform detail. MP3, M4A, and OGG work when the bitrate is reasonable. Voice notes re-exported through messaging apps, screen recordings, and clips re-recorded from a speaker lose detail and should be treated as weak evidence in either direction.

Screen a Recording Now: Free & Private

EyeSift's AI voice detector runs entirely in your browser. No signup, no email, no server upload. Indicative results, not proof.

Open the AI Voice Detector