AI & Creative Tools

Julia-1 Is a 144M-Parameter Model That Only Makes Decisions

Give it a context, a question and between 2 and 20 options, and it picks one — or scores them, or answers yes/no. It cannot do anything else, which is the point.

Most interactive work that reaches for a language model doesn’t need one. The actual requirement is usually far smaller: given what’s happening, pick one of these things. Which response, which sequence, which branch, which of these five states best describes the visitor’s input.

Doing that with a general-purpose LLM works and is absurd — you’re running billions of parameters, probably over an API, to get back one of five words.

Julia-1, from Supersonic Labs, is 144.3 million parameters and does only this.

What it is

A specialised decision-making model that transforms context and questions into clear choices. It’s built for classification, routing and scoring rather than general language understanding, and it handles three decision types through one interface:

  • Choice selection — pick the best option
  • Ordered scoring — rank or rate
  • Boolean decisions — yes or no

The input shape is: a state (context), a question, and 2 to 20 possible answers. It returns the appropriate option.

It’s Apache 2.0, and built on JHU CLSP’s mmBERT-small multilingual encoder as its foundation.

The reported numbers

  • 73.15% on typed-decision benchmarks
  • 86% on emotion classification
  • 94% on news categorisation

And the stated limitation, which the model card is admirably direct about: it isn’t designed for knowledge retrieval or multi-step reasoning.

Read those numbers carefully. 73% on the general typed-decision benchmark is not high — roughly one decision in four is wrong. The much better task-specific figures suggest it does well when the decision is well-posed and the options are genuinely distinct, and less well on harder or vaguer choices. That’s the right way to think about deploying it: it’s good at decisions you could have written rules for, and it saves you writing the rules.

Why a 144M decision model is genuinely useful for creative work

Three reasons, and they’re all about it being small.

It runs anywhere. 144M parameters is embeddable — on a laptop, in a browser via ONNX or WebGPU, on a single-board computer inside the artwork. No API, no key, no per-call cost, no network dependency on the one evening the venue’s Wi-Fi fails. For an installation that has to run unattended for months, removing a network dependency is worth more than accuracy.

It’s fast enough to be interactive. A model this size responds in milliseconds rather than the second or two an LLM round-trip costs. In interaction design that’s the difference between a piece that feels responsive and one that feels like it’s buffering.

The structured interface removes prompt fragility. Asking an LLM to “reply with only one of these five words” works most of the time and fails in ways that break your parser — extra punctuation, an explanation you didn’t ask for, a sixth option it invented. A model whose interface is “here are the options, return one” cannot do that. Anyone who has written defensive parsing around an LLM’s output will recognise the value.

Where it fits, and where it doesn’t

Good fits: routing visitor input to one of N behaviours; classifying sentiment or emotion to drive visuals; choosing the next section of a generative composition; scoring which of several responses suits the current state; boolean gates in an interactive narrative.

Bad fits: anything needing knowledge it wasn’t trained on, anything conversational, anything requiring a chain of inference. The model card says so and it should be believed.

The broader pattern worth noticing: after two years of everything being solved by making models bigger, small task-specific models are becoming genuinely interesting again — precisely because creative work runs on constrained hardware, in rooms with bad networks, and needs answers in milliseconds.