Visualization grammars are among the better-designed things in software. Vega-Lite, and the Grammar of Graphics tradition behind it, let you describe a chart declaratively — this field on x, that one on y, colour by category — and the constraints exist because a human had to write and read it. Sensible defaults, terse syntax, implicit behaviour where the obvious thing is almost always right.
“Where LLMs Fail with Visualization DSLs”, posted 1 October 2026 by Chang Han, Andrew McNutt and Katherine Isaacs, starts from the observation that the author has changed.
As LLMs take up the role of authoring charts using visualization domain-specific languages, the human constraints that shaped those languages may no longer apply, as what is easy for a person is not necessarily easy for a model.
The study
- 10 JSON-style visualization DSLs
- 41 tasks
- 3 LLMs
- Generated specifications assessed with JSON validity checks and rendering checks
- Qualitative coding of the failures
Result: four recurring failure patterns, each linked to specific DSL features, with design considerations for future DSLs.
Why the premise is right, and sharper than it sounds
The claim that human ergonomics and machine ergonomics diverge is worth unpacking, because the divergences run in both directions.
Terseness helps a person and can hurt a model. A grammar that lets you omit anything inferable is pleasant to type. But omission means the correct output depends on knowing what the system will infer — and a model that has seen a thousand variants of a spec across different versions and dialects has to guess which defaults are in play. Verbosity a human would resent is free for a model, which does not get bored.
Implicit behaviour is the same trap. “If you don’t specify a scale, we’ll pick a sensible one” is excellent design for an author iterating in a notebook and a source of silent wrongness for a generator that cannot see the rendered result.
One syntactic position doing several jobs is a human convenience. Overloaded fields — a property that accepts a string, or an object, or an array, meaning different things — save typing and are a reliable source of model error, because the model must commit to a shape before it has finished reasoning about the content.
And the training-data distribution matters enormously. Vega-Lite has a large public corpus; a newer or more obscure DSL does not. A model’s competence in a language is roughly its exposure to it, which means DSL design and DSL popularity are now entangled in a way they never were for human authors — a person can learn a new language from a spec in an afternoon.
Meanwhile models are better than people at some things: they do not mind writing four hundred lines of fully-explicit JSON, they do not make typos in long enum names, and they do not get tired three charts in.
Why JSON specifically, and why that is the hard case
All ten DSLs studied are JSON-style, which is not an arbitrary choice — it is where this whole family lives, from Vega and Vega-Lite to ECharts, Plotly and the rest.
JSON is a particularly interesting substrate for generation. Structural validity is cheap to check, which is why the study’s first gate is a JSON check. But structural validity is nearly uncorrelated with semantic correctness: a model can emit perfectly well-formed JSON that specifies a chart which fails to render, or renders and shows the wrong thing. Hence the second gate — a rendering check — and hence the qualitative coding, because “it rendered but it’s wrong” is not machine-detectable.
That three-stage evaluation is the methodological contribution and it is the right way to measure this. Counting parse failures alone would have made every DSL look fine.
What to do with it
If you are generating charts with a model: pick the DSL with the largest public corpus, be explicit rather than relying on defaults, and always render-check the output before showing it to anyone. A spec that parses is not a chart that is correct.
If you design a DSL, an API or a library — and a great many people reading this do — the broader lesson is that a model is now one of your users, and the usual ergonomic instincts can work against it. Concretely, the things that help:
- Explicit over inferred. Let callers state everything; don’t require them to know defaults.
- One shape per field. Overloaded parameter types are a generation hazard.
- Fail loudly and specifically. A model given a precise error message can repair its output; a silent fallback to a default produces a wrong result nobody notices.
- Publish examples. The corpus is the documentation now, in a very literal sense.
- Keep names unambiguous and stable. Renaming a property across versions means the model has two conflicting memories of your API.
None of which means designing for models over people. It means the two audiences have different failure modes, and the paper’s useful move is identifying which language features produce which failures rather than concluding that models are bad at DSLs.