AUDIOMIND← analyzer

about

A music-genre classifier that runs entirely in your browser.

Drop in a clip and AudioMind turns it into the same kind of picture its neural network trained on — a mel spectrogram — and predicts one of 10 genres. The model is downloaded once and runs on-device with TensorFlow.js, so there’s no server, nothing to wait on, and your audio never leaves your machine.

10genres
85%song-level accuracy
0.42 MBmodel download
100%runs in your browser
0bytes of audio uploaded

How it works

The trick is that a spectrogram is a 2-D image, so classifying a genre becomes classifying an image — a problem CNNs are very good at. The new work is at the front: turning sound into that image, exactly the way the training pipeline did.

  1. 01

    Decode & resample

    Your file is decoded and resampled to mono at 22,050 Hz using the browser's Web Audio API — the same sample rate the model was trained on. wav, mp3 and flac all work; nothing is uploaded.

  2. 02

    Segment

    The clip is cut into non-overlapping 3-second windows (66,150 samples each). Each window is classified on its own, and the results are averaged — a whole song is easier to place than any single moment of it.

  3. 03

    Log-mel spectrogram

    Each window becomes a picture the model can read: a short-time Fourier transform (n_fft 2048, hop 512), mapped onto 128 perceptual mel bands, then converted to decibels. The result is a 128 × 130 image of how energy is spread across pitch and time.

  4. 04

    CNN → genre

    A small convolutional neural network reads that spectrogram like an image and outputs a probability for each of the 10 genres. Averaging those probabilities across the clip's windows gives the final call.

The model

A compact convolutional network — three convolution blocks (32 → 64 → 128 filters, each with batch-norm and max-pooling), global average pooling, then a dense layer into a 10-way softmax. Its input is a single 128 × 130 × 1 spectrogram; its output is 10 probabilities that sum to 1.

Accuracy

85.3% at the song level (averaging a clip’s windows), 78.3% on a single 3-second window. Chance is 10%.

Where it slips

Mistakes cluster between genres that genuinely sound alike:

  • rock ↔ metal — shared distorted guitars and driving drums
  • disco ↔ pop — similar four-on-the-floor production and tempo
  • reggae ↔ hiphop — prominent bass and off-beat, sparse arrangements
  • blues ↔ country — overlapping acoustic instrumentation and phrasing

Everything runs on-device

The Keras model was converted to a TensorFlow.js graph model — about 0.42 MB — and runs in the page. Two things make that trustworthy:

  • Privacy. Audio is decoded and analyzed locally. Nothing is uploaded; there’s no backend at all.
  • Fidelity. The browser must build the exact spectrogram the model trained on, or predictions drift — so librosa’s mel filterbank is exported verbatim rather than re-derived in JavaScript. The result is verified two ways: the in-browser spectrogram matches librosa to ~1.7×10⁻⁴ dB, and the converted model reproduces the original Keras predictions.

The data it learned from

AudioMind was trained on GTZAN, the standard benchmark for this task: 1,000 clips, 30 seconds each, 100 per genre, across the 10 genres below.

bluesclassicalcountrydiscohiphopjazzmetalpopreggaerock

GTZAN is a product of its time (assembled in the early 2000s) and is well known in the research literature to contain repeated tracks, some mislabelled clips, and a few corrupted files. It skews toward Western, mostly 20th-century recordings. It’s excellent for learning and benchmarking — and a fair reminder that a model is only ever as broad as the data behind it.

What it can’t do

  • It knows these 10 genres and nothing else — hand it lo-fi, afrobeats, or a podcast and it will still confidently pick one of the ten.
  • It judges timbre and texture over ~3 seconds, not lyrics, structure, or artist. A genre-blending track can land anywhere.
  • It reflects GTZAN’s era and taste, so modern or non-Western music may be classified less reliably.
  • It’s a focused demonstration of the audio → spectrogram → CNN pipeline, not a production music-tagging service.

Glossary

Spectrogram
A 2-D picture of sound: time runs left-to-right, frequency bottom-to-top, and brightness is how much energy is at that pitch and moment.
Mel scale
A warping of frequency that matches how humans hear — we tell low pitches apart far better than high ones. Mel bands give the model perceptually meaningful features.
STFT
Short-Time Fourier Transform. Slide a small window along the audio and, for each position, measure how much of each frequency is present. Stacking those columns builds the spectrogram.
CNN
Convolutional Neural Network. The kind of model that excels at images; since a spectrogram is an image, genre classification becomes image classification.

Built with

Keras / TensorFlow and librosa for training and feature extraction; TensorFlow.js for in-browser inference; Next.js, TypeScript and the Web Audio API for the app. The spectrogram pipeline is checked against librosa and the original model in the test suite, so accuracy claims stay honest as the code changes.

Try it →