Research & Innovation

Benchmark Tutorial: How to Measure a Model on Your Own Work Instead of Trusting a Table

Build a 50-case eval from your own material, score it the same way twice, and find out whether the model that wins the leaderboard wins your job.

Every model release comes with a table, and the tables are mostly useless to you. A number on GPQA-Diamond or AIME tells you how a model handles graduate physics and competition mathematics; it tells you almost nothing about whether it will write usable TouchDesigner Python, caption your archive consistently, or decide correctly whether a command is safe to run.

The fix is unglamorous and takes an afternoon: build a small evaluation out of your own work and run it yourself.

Adam Lucek on what public benchmarks do and do not measure, and how to run your own.

Step 1 — Decide what question you are answering

Write it down in one sentence before you collect anything. Not “is model A better than model B,” which is unanswerable, but something like:

Does this model produce Hydra shader code that runs without edits, from a one-line description?

A good eval question has a specific task, a specific input shape, and an outcome you can check. If you cannot say what a right answer looks like, you are not ready to build the eval — and that is useful information, because it usually means the task is one where you are the judge and no benchmark will help.

Step 2 — Collect 30 to 50 real cases

Not 10, not 1,000.

Why not 10: the arithmetic is brutal. With 10 cases, one case is 10 percentage points. Two models scoring 70% and 80% differ by a single item, which is noise. You cannot distinguish anything.

Why not 1,000: you will not do it, and you do not need to. The gap between models that matters in practice is usually large, and 50 cases resolves a large gap.

Where cases come from: your actual history. Old prompts, real project briefs, files from finished work, the questions you asked last month. Pull them from things you did before you started building this eval, which keeps you honest — cases invented for the eval drift toward what you expect the model to handle.

Include the hard ones. An eval made only of representative cases will have every model scoring 90% and tell you nothing. Deliberately include the cases you remember going badly. Roughly: half typical, a third difficult, a sixth genuinely adversarial.

Step 3 — Pick a scoring method, and pick the strictest one that fits

Three options, in descending order of how much you should want them.

Exact or programmatic match — best, use it wherever the task allows. The output is code: does it run? The output is JSON: does it validate against the schema? The output is a classification: does it equal the label? No judgement, no drift, re-runnable forever.

This is why typed outputs are worth engineering for. If you can restructure the task so the answer is a choice from a list, a number, or code that either executes or does not, you get a benchmark you can trust. If the task is “write me something nice,” you do not.

A written rubric scored by you — for anything subjective. The rules:

  • Three to five levels, not ten. You cannot reliably distinguish 7 from 8 on your own scale; you can distinguish fails / works / good.
  • Write the rubric before you see any output. Afterwards you will write a rubric that matches what you got.
  • Score blind. Strip the model names, shuffle, then score. This is not optional — knowing which model produced an output changes how you read it, reliably and unconsciously.

Pairwise comparison — show two outputs for the same input, pick the better one. Easier and more consistent than absolute scoring, because relative judgement is something humans are good at and absolute calibration is not. The cost is that you get a ranking, not a level: you learn A beats B, not whether either is good enough.

A fourth option, model-as-judge, is tempting and belongs behind a caveat — see below.

Step 4 — Control the variance before you compare anything

This is where most home-made evals quietly fail.

Set temperature to 0 for everything, or if the task needs sampling, run each case 3–5 times and record the spread. A single sampled generation per case produces differences between models that are entirely sampling noise.

Pin everything you can pin: the exact model version, the quantisation, the context length, the system prompt, the seed if available. “Qwen3.8-27B” is not a specification — a Q4_K_M GGUF and the fp16 original are measurably different models, and six months later you will not remember which you tested.

Measure your own repeatability first. Score ten outputs, wait a day, score them again blind. If you disagree with yourself on three of ten, your resolution is about 30 percentage points and any difference smaller than that is not real. This step is annoying and it is the one that tells you whether your eval means anything.

Record a baseline. Run the eval against whatever you use today, and against something deliberately weak. If a model you know is poor scores 65%, then 70% is not a result.

Step 5 — Use lm-evaluation-harness when a standard benchmark is what you want

Sometimes the question genuinely is “how does my fine-tune compare on an established task,” and then you want the standard tooling rather than your own.

pip install lm-eval

lm_eval --model hf \
  --model_args pretrained=Qwen/Qwen3-4B,dtype=float16 \
  --tasks hellaswag,arc_challenge \
  --device cuda:0 \
  --batch_size 8 \
  --output_path ./results

It supports Hugging Face models, local GGUF through llama.cpp, and OpenAI-compatible endpoints, which means you can point it at your own server.

Two things to understand about what it gives you:

The numbers are comparable only to numbers produced the same way. Prompt formatting, few-shot count and answer extraction all move scores by several points, which is why the same model appears with different numbers in different papers. Record your exact invocation alongside the result.

Contamination is real and unfixable from your side. The standard benchmarks are old and public, which means they are in training data. A good score on HellaSwag may mean the model is good or may mean it has seen HellaSwag. Your own eval, made of your own material, has this problem to a far smaller degree — which is the main reason to prefer it.

On using a model as the judge

It works better than you would expect and it has specific, documented failure modes you need to design around.

Position bias. Judges prefer whichever answer came first. Mitigate by running every comparison in both orders and discarding cases where the verdict flips — those are ties, and the flip rate is itself a measure of how reliable the judge is.

Length bias. Judges prefer longer answers almost regardless of quality. If one model is verbose, this alone can produce a win rate.

Self-preference. A judge tends to favour outputs from its own family. Never use a model to judge itself or its siblings.

Validate the judge against yourself. Score 20 cases by hand, have the judge score the same 20, and measure agreement. If it agrees with you 85% of the time, you have a usable instrument with a known error rate. If it agrees 60% of the time, you have a random number generator with good manners. Do not skip this, and report the agreement figure whenever you report judged results.

Step 6 — Report it in a way your future self can use

Keep a single file per eval run with:

  • the question in one sentence
  • the number of cases and where they came from
  • the exact model identifiers, quantisation and sampling settings
  • the scoring method and the rubric verbatim
  • your own self-agreement figure
  • the baseline scores
  • the result, with the number of cases the difference represents

That last item is the one everyone omits and the one that prevents the most mistakes. “78% vs 72%” sounds decisive. “39/50 vs 36/50 — a difference of three cases” is the same fact and invites the right amount of scepticism.

Where to go next

  • Keep the eval and re-run it. Its real value is longitudinal: the same 50 cases against every model you consider for the next two years becomes the most useful instrument you own.
  • Add cases when something fails in production. Every real failure is a case you were missing.
  • Publish it if you can. An eval built from a specific practice is more informative to people in that practice than any general benchmark, and almost nobody publishes them.