How Shazam's 20-Year-Old Audio Fingerprinting Algorithm Actually Works
Audio fingerprinting technology, most famously used by Shazam, can identify a recording from just ten seconds of noisy, degraded audio captured in a loud environment. The algorithm, described by Avery Wang in a 2003 paper titled 'An Industrial-Strength Audio Search Algorithm,' works by detecting local peak points in a audio spectrogram and discarding all amplitude data. These sparse 'constellation' points are then paired and hashed using only their frequencies and the time gap between them, making the lookup key independent of where in a track the audio clip begins. This design means the system can match a query clip against millions of recordings without needing to know its position in the song, while remaining robust to re-encoding, background noise, and audio filtering. The approach outperforms alternatives like cryptographic hashing, waveform correlation, and neural embeddings, each of which breaks down under real-world audio degradation or fails to scale efficiently.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in