AI & Creative Tools

MERT-v2-FullSong Listens to a Whole Six-Minute Track at Once

Multimodal Art Projection's new 632M-parameter music encoder extends MERT's context from 30 seconds to 360, keeping the same feature interface — which means existing MIR pipelines get song-level structure for free.

Almost every machine-listening model that has touched music works on a window — thirty seconds, often less. That is fine for genre tags and instrument recognition, and it is structurally blind to the thing that makes a piece of music a piece of music: what happens over its whole length.

MERT-v2-FullSong, published to Hugging Face by Multimodal Art Projection (m-a-p) on September 9, 2026, takes the same architecture out to 30–360 seconds of context. Six minutes. A whole song.

What it is

MERT is a bidirectional music encoder — it does not generate audio, it produces representations of it. You feed it a recording and it hands back a numerical description that downstream models can classify, search, or cluster.

The specifics:

  • 632M parameters across 24 layers
  • 24 kHz mono input
  • 1,024-dimensional outputs at 25 Hz — so twenty-five vectors per second of audio
  • Loads through standard Hugging Face Transformers via AutoModel.from_pretrained() with trust_remote_code=True
  • CC-BY-NC-4.0 — non-commercial only

Crucially, it continues pretraining from MERT-v2-30s and preserves the same feature interface. If you already have a pipeline built on the 30-second model, swapping in the full-song version does not mean rewriting the downstream half.

Why the context length is the whole story

At 25 Hz, a 360-second track is 9,000 frames the model attends across. That buys you things a clip model cannot represent no matter how good it is:

  • Form. Verse, chorus, verse, bridge. Where sections begin and end, and which ones are the same section returning.
  • Repetition and variation. A chorus that comes back louder, or with a countermelody added, is legible only if you saw the first one.
  • Long-range development. Ambient, drone, classical, DJ mixes, and most electronic music do their work across minutes. A thirty-second window samples them roughly the way a single frame samples a film.

The model card points at MARBLE, the standard benchmark suite for music understanding, with strong results across its ten tasks — genre classification, emotion detection, instrument recognition, chord estimation and the rest.

Where it’s actually useful

For creative work the interesting applications are not classification, they are retrieval and organisation:

  • Search your own library by sound, not metadata. Pool the frame-level output into a single recording-level embedding and you have a vector you can do nearest-neighbour search against — “find me things that sound like this” across a sample collection nobody ever tagged.
  • Structure-aware editing. Section boundaries recovered from a full-song encoder are a better basis for automatic edits than beat-grid guessing.
  • Audio-reactive visuals with memory. Most audio-reactive work responds to the last 50 milliseconds. Features that encode where you are in the song’s arc open a different kind of piece.

The licence, plainly

CC-BY-NC-4.0 means non-commercial. You can research with it, build tools you don’t sell, and use it in personal or non-commercial artistic work. You cannot ship it inside a product. Given how much of the MIR ecosystem is academic this is a normal licence for the field, but it is worth knowing before you build a workflow on it.

For anyone who wants to try it: start by pooling across time for recording-level embeddings, which is the simplest useful thing, and only reach for frame-level output when you specifically need the time axis.