Research & Innovation

ONNX From Zero: Get a Model Out of PyTorch and Into Anything

Export once, run on a browser, a phone, a Raspberry Pi or a C++ installation binary — without shipping Python or a two-gigabyte framework.

You have a model that works. It works in a notebook, on your machine, in a conda environment with 4GB of dependencies and a specific CUDA build. Now it needs to run inside a TouchDesigner installation, or on a Raspberry Pi in a gallery, or in a browser, or in a C++ binary that starts in under a second.

ONNX — Open Neural Network Exchange — is how you get it out.

The ONNX project’s own overview of the PyTorch converter.

What it is

ONNX is a file format for a computation graph. An .onnx file contains the network’s operations, their connections, and the trained weights — expressed in a standard operator set rather than in any framework’s internal representation.

Export once, then run it with ONNX Runtime, which has bindings for C++, C#, Python, JavaScript (including WebAssembly and WebGPU), Java, Rust, Objective-C and Swift, on Windows, macOS, Linux, Android, iOS and the web.

The thing you are buying is the removal of Python from your deployment. That is the whole value proposition and it is larger than it sounds.

The export

import torch

model.eval()
dummy = torch.randn(1, 3, 224, 224)      # shape of a real input

torch.onnx.export(
    model,
    dummy,
    "model.onnx",
    input_names=["input"],
    output_names=["output"],
    dynamic_axes={
        "input":  {0: "batch"},
        "output": {0: "batch"},
    },
    opset_version=17,
    do_constant_folding=True,
)

Four things in there that matter:

model.eval() — before anything. Dropout and batch-norm behave differently in training mode, and an export taken in training mode bakes in the wrong behaviour. This is the single most common silent error.

A real dummy input. The exporter traces the model by running it, so the dummy’s shape and dtype must match real input. Garbage shape, garbage graph.

dynamic_axes — without this, your exported model accepts exactly the dummy’s shape forever. Mark batch as dynamic always; mark height and width dynamic if your model handles variable resolution. People discover this when their production model refuses a batch of 4.

opset_version — the operator-set version. Newer supports more operators; older is compatible with more runtimes. 17 is a safe modern default. If a runtime rejects your model, this is the first thing to lower.

Then verify, immediately:

import onnx, onnxruntime as ort, numpy as np

onnx.checker.check_model(onnx.load("model.onnx"))

sess = ort.InferenceSession("model.onnx")
onnx_out = sess.run(None, {"input": dummy.numpy()})[0]
torch_out = model(dummy).detach().numpy()

print(np.abs(onnx_out - torch_out).max())   # want < 1e-4

Always do this numerical check. An export can succeed and be wrong, and finding out in the gallery is expensive.

The three things that reliably break

1. Python control flow that depends on tensor values.

if x.sum() > 0:          # traced once, baked in forever
    return self.a(x)
else:
    return self.b(x)

Tracing runs the model once with your dummy input and records what happened. A branch taken on dummy data is the only branch in the exported graph. The fix is torch.jit.script on that module, or restructuring to avoid data-dependent branching, or torch.onnx.dynamo_export which handles more cases.

2. Unsupported operators. You get Unsupported: ONNX export of operator .... Options, in order: raise the opset, replace the op with an equivalent composition of supported ops, or register a custom symbolic function. Exotic attention implementations and custom CUDA kernels are the usual offenders.

3. Shapes baked in from .shape arithmetic.

b, c, h, w = x.shape
x = x.view(b, c, h * w)        # h*w may become a constant

Use x.flatten(2) or -1 instead of arithmetic on .shape values. This is why a model exports fine and then rejects a different resolution.

Running it where you want it

C++ — link onnxruntime, load the file, feed a float buffer. This is the path for a TouchDesigner CPlusPlus TOP, an openFrameworks addon, a JUCE audio plugin, or a standalone installation binary. No interpreter, fast startup, one dependency.

Browser — onnxruntime-web runs via WebAssembly, with WebGPU and WebGL backends. A model in a web page with no server is a genuinely different distribution model, and it is the same trend we keep noting in WebGPU vector graphics, Web MIDI and Web Audio tools.

Raspberry Pi / ARM Linux — onnxruntime has ARM builds. Combined with quantisation this is how a model ends up in an unattended installation.

Mobile — ONNX Runtime Mobile, a reduced build with only the operators your model uses.

Unity — Unity’s Sentis (formerly Barracuda) consumes ONNX directly, which is the standard route for ML in a game engine.

Quantisation, in one paragraph

ONNX Runtime includes quantisation tooling, and it is the main reason to bother on small hardware:

from onnxruntime.quantization import quantize_dynamic, QuantType
quantize_dynamic("model.onnx", "model.int8.onnx", weight_type=QuantType.QInt8)

Dynamic int8 quantisation is roughly a 4x size reduction and a meaningful speedup on CPU, for a usually-small accuracy cost. Measure the cost on your own data — usually small is not always small. Static quantisation is better and needs a calibration dataset.

When not to use it

If you are still iterating, stay in PyTorch. ONNX is a deployment step, not a development environment, and exporting on every change is friction with no payoff.

If you need to train. ONNX is for inference. Training export exists and is not what you want.

If the model is enormous. A 78B-parameter MoE is not going through this pipeline into a browser. ONNX is for the small-to-medium models that fit the places you want to put them — which, for installation and interactive work, is most of them.

And if a better-supported path exists for your target. Core ML on Apple platforms, TensorRT on NVIDIA, llama.cpp/GGUF for language models. ONNX’s advantage is breadth, not peak performance on any one platform.