Research & Innovation

A Reference-Listening Tool That Makes You Choose What to Compare

Diptych lets you scope the comparison — whole tracks or specific segments — then shows audio features alongside an AI reading of them. Around 90% of what users found held up to expert review.

Every mix engineer does reference listening. You load a commercial track you admire next to your own, switch between them, and try to work out what the difference is. It is the most reliable technique in the craft and the hardest to do well, because the honest answer to “what’s different?” is usually everything, slightly.

The tools that claim to help mostly make it worse. Diptych: Scoped, AI-Interpreted Comparison for Reference Listening in Music Production, posted 30 September 2026 by Chongjun Zhong, Abhinaba Roy, Archishman Ghosh, Kejun Zhang and Dorien Herremans, identifies why.

The diagnosis

Existing comparison tools fail to let users decide precisely what deserves comparison.

That is the whole problem in one sentence, and anyone who has used a reference-matching plugin will recognise it. You load two tracks and the tool reports a difference in spectral balance, loudness, stereo width, dynamic range — across the entire file, averaged.

Which is almost never the question you had. The question was why does their chorus hit harder than mine, or why does my low end disappear on the second verse. Those are scoped questions about specific moments, and a whole-file average actively obscures them: the reference’s quiet intro drags its spectral average down, your dense bridge drags yours up, and the resulting comparison describes two songs that neither of you made.

What Diptych does

Users define what to compare — whole tracks or specific segments — and then get structured audio feature analysis alongside AI-generated interpretations tailored to the chosen scope.

Three things in that sentence, in the right order:

  1. The user sets the boundaries. Not the tool.
  2. Measured features come first. Actual numbers from the audio.
  3. The AI interprets those features, for the scope you chose.

The AI is reading the measurements, not replacing them. Which means the interpretation can be checked against the evidence it was derived from — a property almost no AI audio assistant currently offers.

The results

A study with 12 musicians plus expert validation:

  • Participants discovered additional differences between tracks that they had not previously noticed
  • About 90% of those findings received at least partial expert support
  • Positive usability, and “greater clarity about possible next steps”

The 90% figure is the one to dwell on, and it is reported with appropriate care — at least partial expert support, not vindication. What it says is that the tool surfaced things that were really there, rather than generating plausible-sounding observations about audio.

That is the failure mode this category is prone to. A language model asked to describe the difference between two tracks will produce fluent, specific, confident mix notes whether or not it has any grounding. “The reference has more presence in the upper mids and a tighter low end” is a sentence that is true of most pairs of tracks and tells you nothing. Grounding the interpretation in measured features and then checking the output against experts is how you find out whether you built an analysis tool or a sentence generator.

“Greater clarity about possible next steps” is the other result worth noting, because it is the actual user need. Nobody wants to know that their track differs from the reference. They want to know what to do on Tuesday morning.

The design conclusion, which generalises well past audio

The authors’ recommendation: AI comparison tools should emphasise user-defined scope, visible evidence, and practical guidance — rather than delivering definitive judgements unsupported by the underlying data.

Each of those three is a corrective to a specific, common, current mistake:

User-defined scope against tools that analyse everything and therefore nothing. The user knows which eight bars are the problem; let them say so.

Visible evidence against the black-box verdict. If the system says the reference has more low-mid energy, show the measurement. This is what makes the tool teachable — you learn to hear the thing because you can see it alongside what you are hearing.

Practical guidance against description masquerading as help.

And no unsupported definitive judgements — which is the hardest one commercially, because confidence sells. A tool that says “your mix needs 2dB less at 400Hz” feels more valuable than one that says “there is more energy here than in your reference, in this section, and here is the plot.” The second is honest and more useful, and it will lose a demo to the first.

For anyone building creative tools with a model in them, that trio is a better specification than most product documents: let the user scope it, show your work, suggest an action, and don’t pretend to certainty you can’t evidence.