Earshot

Sonic investigations for communities affected by corporate, state, and environmental injustice

What we do

Expand

Featured investigations

Expand

Hondurasgate: Audio Authentication of Twelve Recordings

The recordings behind a plot to return Juan Orlando Hernández to power.

Since May 2026, Earshot has been analysing several alleged recordings of prominent political figures involved in the Hondurasgate scandal, at the request of Drop Site News, in order to determine their authenticity or whether they were AI-generated. If authentic, the recordings would expose a transnational effort to reinstate Juan Orlando Hernández, the former president of Honduras and a convicted criminal, to the Honduran presidency following his pardon by Donald Trump. His return would "reposition the country as a strategic ally for the US and Israel in the region" and create "a media outlet with the express mission to undermine leftist governments in Mexico and Colombia."

Following a first round of analysis of three recordings, Earshot was asked by an anonymous journalist from Honduras to analyse an additional twelve recordings containing the alleged voices of Honduran President Nasry Asfura, former President Juan Orlando Hernández, Vice President María Antonieta Mejía, and Cosette López1 .

By identifying several audio artefacts consistent with recordings made in real-world conditions and performing machine-learning voice comparisons, Earshot concluded that all twelve recordings were highly likely to be authentic and not AI-generated. This strongly indicates that plots to undermine democratic processes in Honduras by the United States are being actively pursued.

Authentication Methodology

This conclusion rests on two lines of analysis: critical listening to identify audio artefacts consistent with authentic speech, and machine-learning voice comparison - using a programme called Resemblyzer - with known recordings of the alleged speakers. The critical-listening findings were further tested against AI-generated speech produced with Coqui, a voice-cloning programme2 . A separate line of analysis, examining each recording's sample rate, supports the findings of the first two.

I. Critical listening for audio artefacts, tested against Coqui

Multiple audio artefacts could be identified across each of the twelve recordings that are consistent with natural speech recorded in real-world conditions. These fall under two categories: speech artefacts and acoustic artefacts. Speech artefacts, such as breaths and vocal hesitations, and acoustic artefacts, such as background noises and microphone distortions, could be heard between sentences.

The presence of multiple artefacts in a single recording points to a coherent acoustic web that is consistent with speech recorded in real-world conditions. However, advances in AI-generated audio are progressively incorporating these sounds into the generative process. This poses challenges to authentication methodologies. To test how reliable each artefact is as a marker of authenticity, Earshot checked whether they could be reproduced using Coqui.

Several artefacts proved resistant to reproduction by current generative-AI programmes - among them breaths preceded by the distinct sound of the tapping of teeth or lips, yawns, and background noises with the same room resonance as the voice. Together with the coherent acoustic web identified across all twelve recordings, this makes it unlikely that the artefacts were produced by AI-generated speech.

II. Machine-learning voice comparison

For all twelve recordings, Earshot ran a machine-learning voice comparison using Resemblyzer, a programme that analyses short samples of a voice and translates their patterns of rhythm, prosody, frequency composition, and harmonics into a numerical value. This value is compared to the value from a known recording of the speaker in question; the resulting score indicates how similar the two voices are, and by extension how likely the alleged voice is to be genuine3 . Across all twelve recordings, these results supported the conclusions reached through critical listening.

Breaths with taps

0:00

Clip from audio "art4-FILTRACION_146" of Juan Orlando Hernández containing two breaths characterised by a brief tapping of the lip or teeth preceding the intake of air.

Microphone distortions

0:00

Clip from audio "art5-FILTRACION_190" of María Antonieta Mejía containing two microphone distortions caused by direct exhalations onto the microphone.

III. Low sample rate

The relatively low sample rate found across all twelve recordings provides further evidence of their authenticity. Sample rate refers to how many times per second a microphone measures sound pressure. Digital audio is typically recorded at 44,100 or 48,000 samples per second, capturing a wide range of frequencies and therefore more detail. Telephone audio, by contrast, is usually recorded at a lower sample rate, since it only needs to capture speech frequencies (0 to 8,000 Hz) rather than the full audible range (0 to 20,000 Hz). All twelve recordings contain frequencies only up to 4,500 Hz, indicating they were recorded at a relatively low sample rate.

Earshot's ongoing experiments have not encountered AI-generated audio with such a low sample rate. While it is possible to manually reduce any recording's sample rate, Earshot considers this more likely to be a product of authentic telephonic conversations.

The strength of these conclusions reflects both the number of audio artefacts identified in each recording and the score obtained through machine-learning voice comparison. That said, the findings on speaker comparison are limited by the low telephonic quality of all twelve recordings, and ongoing advances in machine-learning algorithms are continually improving the quality of AI-generated speech, making it harder to detect. Further analysis of accent and dialect should be conducted by a Spanish-speaking sociolinguist.

Press

Expand

Partners

Expand