Almost every machine-listening model that has touched music works on a window — thirty seconds, often less. That is fine for genre tags and instrument recognition, and it is structurally blind to the thing that makes a piece of music a piece of music: what happens over its whole length.
MERT-v2-FullSong, published to Hugging Face by Multimodal Art Projection (m-a-p) on September 9, 2026, takes the same architecture out to 30–360 seconds of context. Six minutes. A whole song.
What it is
MERT is a bidirectional music encoder — it does not generate audio, it produces representations of it. You feed it a recording and it hands back a numerical description that downstream models can classify, search, or cluster.
The specifics:
- 632M parameters across 24 layers
- 24 kHz mono input
- 1,024-dimensional outputs at 25 Hz — so twenty-five vectors per second of audio
- Loads through standard Hugging Face Transformers via
AutoModel.from_pretrained()withtrust_remote_code=True - CC-BY-NC-4.0 — non-commercial only
Crucially, it continues pretraining from MERT-v2-30s and preserves the same feature interface. If you already have a pipeline built on the 30-second model, swapping in the full-song version does not mean rewriting the downstream half.
Why the context length is the whole story
At 25 Hz, a 360-second track is 9,000 frames the model attends across. That buys you things a clip model cannot represent no matter how good it is:
- Form. Verse, chorus, verse, bridge. Where sections begin and end, and which ones are the same section returning.
- Repetition and variation. A chorus that comes back louder, or with a countermelody added, is legible only if you saw the first one.
- Long-range development. Ambient, drone, classical, DJ mixes, and most electronic music do their work across minutes. A thirty-second window samples them roughly the way a single frame samples a film.
The model card points at MARBLE, the standard benchmark suite for music understanding, with strong results across its ten tasks — genre classification, emotion detection, instrument recognition, chord estimation and the rest.
Where it’s actually useful
For creative work the interesting applications are not classification, they are retrieval and organisation:
- Search your own library by sound, not metadata. Pool the frame-level output into a single recording-level embedding and you have a vector you can do nearest-neighbour search against — “find me things that sound like this” across a sample collection nobody ever tagged.
- Structure-aware editing. Section boundaries recovered from a full-song encoder are a better basis for automatic edits than beat-grid guessing.
- Audio-reactive visuals with memory. Most audio-reactive work responds to the last 50 milliseconds. Features that encode where you are in the song’s arc open a different kind of piece.
The licence, plainly
CC-BY-NC-4.0 means non-commercial. You can research with it, build tools you don’t sell, and use it in personal or non-commercial artistic work. You cannot ship it inside a product. Given how much of the MIR ecosystem is academic this is a normal licence for the field, but it is worth knowing before you build a workflow on it.
For anyone who wants to try it: start by pooling across time for recording-level embeddings, which is the simplest useful thing, and only reach for frame-level output when you specifically need the time axis.