ONE RECORDING. TWO MODES.

Would you trust
this voice?

Play the real-or-fake listening game. Learn why convincing voices can mislead us. Then put a recording through HEARSAY and inspect the evidence.

TRUST YOUR EARS? TRY IT.

Same voice. Which recording is fake?

Two recordings attributed to the same person: one genuine, one synthetic. Listen to both, then choose. No access code needed to play.

0 correct / 0 answered

Loading the listening game…

A

Recording A

B

Recording B

Listen for yourself before choosing.

What does this game measure?

Your score is for this session and these demonstration clips only. Each pair contains exactly one synthetic recording, so random guessing averages 50%. That setup is different from screening an unknown recording. Speakers and clips may overlap training data; this is not a scientific test of you or HEARSAY. These are different utterances, not necessarily matching sentences.

The reveal uses the dataset labels, not a detector prediction. After answering, analyze either clip to see whether HEARSAY agrees. Audio and labels are public; the game is for learning, not a secure exam.

Recording sources and credits

A FAMILIAR VOICE IS NO LONGER ENOUGH

When a voice sounds right, a false story can travel.

A fabricated recording can impersonate a relative asking for money, a manager requesting a transfer, or a public figure saying something they never said. Voice synthesis also supports accessibility and creative work. The risk is deception: passing generated speech off as a real person's words.

$3.5B

Reported imposter-scam losses

Reported to the FTC for 2025. This includes impersonation across many channels; it is not a voice-cloning loss estimate.

FTC · June 2026
73%

Fake speech spotted in one study

A 2023 study with 529 listeners tested English and Mandarin synthetic speech. Listeners missed roughly a quarter of the fakes. This is a study result, not a universal human detection rate.

UCL / PLOS ONE · August 2023
48 pairs

Experience the ambiguity yourself

14 speakers and 96 paired clips in this game. The full library has 122 recordings, including modern synthetic speech and genuine readers.

Try a round →

Listen critically. Verify independently.

Odd pauses, changing background noise or flattened delivery may prompt a closer look, but compression, editing and genuine speech can produce similar effects. A smooth voice can still be synthetic. No single listening cue establishes authenticity.

If a caller makes an urgent request, pause and contact the person using a number you already trust. For published audio, seek the original recording and corroborating sources. FTC guidance on fake emergencies.

FROM A GUT FEELING TO A CHECKABLE RESULT

What HEARSAY adds

Learn by listening

Compare two recordings of one speaker, commit to an answer, then inspect the supplied labels. Build awareness of uncertainty without needing a technical background.

Choose your analysis budget

Speed runs the original detector. Accuracy uses a fusion head over two pretrained feature paths. Run both to inspect agreement and measured processing time on the same recording.

Keep the evidence visible

See the raw decision score, runtime and failures. Model workers run offline; uploaded recordings are deleted after each job. The access code is shared, so results are visible to other code holders.

  1. Listen or upload. Choose a library clip or your own recording.
  2. Run a detector. HEARSAY analyzes learned audio features in an isolated worker.
  3. Read the result in context. Compare modes, check coverage and verify the recording's provenance before acting.

Measured here, with limits

ModeWarm runFirst loadAudio coverage
Speed0.89 s8.64 sUp to the first 12 seconds
Accuracy2.28 s14.12 sUp to three spaced 4.0375-second windows

Deployment smoke checks on September 27, 2026: warm timings use synthetic-2.flac; first-load timings use synthetic-1.flac. CPU server: 6 vCPUs, 16 GiB RAM; workers limited to 2 CPUs. These are individual observations on short samples, not latency guarantees or an accuracy benchmark. Accuracy mode's name is not a promise that it wins on every recording.

What a score can—and cannot—tell you

Higher scores lean synthetic. The displayed cutoff is 0.5. A score of 0.9 does not mean a 90% chance of being fake: these are decision scores, not calibrated probabilities. A false positive flags genuine audio; a false negative misses synthetic audio. Both matter.

The detectors sample bounded portions of a recording and can miss edits elsewhere. New generators, languages, background noise and encoding changes can shift performance. We do not claim a universal detection rate, official ranking, speaker identity verification or forensic proof.

Model sources, coverage and usage restrictions

Why even a good detector needs context

Illustrative example: screen 1,000 recordings with a hypothetical detector that catches 90% of fakes and falsely flags 5% of genuine audio. Change how common fake recordings are.

Teaching assumptions, not HEARSAY performance measurements. When fakes are rare, false alarms can account for much of the flagged queue.

Offline workers · One detector at a time · No writable host mounts · Bounded analysis time

Model credits and usage limits

Start a comparison

Enter your access code, choose Speed or Accuracy, and upload a recording. Uploaded audio is deleted when the job finishes. Results survive dashboard reloads.

Audio analysis runs in offline workers. Use HTTPS at hearsay-ai.tech to protect uploads in transit. Benchmark timings were measured on the original reference server.

Need audio to try? Download four test clips (ZIP)

Two real and two synthetic recordings, with labels and source credits. These are demonstration samples with possible training overlap, not a fair accuracy benchmark.

Real 1 · Real 2 · Synthetic 1 · Synthetic 2 · Audio credits & labels

Explore the recording library

Listen to 122 recordings, including real and synthetic JFK clips. Pick a recording, then test it with Speed, Accuracy, or both.

Download recording · Source credits

Choose how to analyze

First loads are slower. Timing depends on recording length and server load.

For a quick comparison, use a short speech clip. Scores are synthetic-high model outputs, not interchangeable probabilities or proof of authenticity.

Comparisons

Jobs run sequentially. Results refresh every two seconds. Our models stay ready briefly between uploads for faster repeat tests. Interpretations use a 0.5 decision cutoff, not calibrated confidence. A match means agreement with your supplied label.

No results loaded.