Research & Innovation

Docker for Creative Work: Make the Thing That Worked Keep Working

Four commands, one Dockerfile, and the GPU flag nobody tells you about — so a research repo you got running in October still runs in March.

We published a guide three days ago on getting a research paper’s code actually running, which ended with a habit: write down what worked. This is the version of that habit that a machine can read.

The scenario Docker solves is specific and everyone here has lived it. A thing worked. You changed something unrelated. The thing is now broken and you cannot remember what you did.

OpenInfra Foundation on Docker for machine learning specifically — the GPU material is the part worth staying for.

What it actually is, in one paragraph

A container is a process running on your kernel with its own filesystem. Not a virtual machine — there is no second operating system and almost no overhead. It sees the Python, the libraries, the CUDA runtime and the files you put in its image, and nothing else from your machine.

An image is the frozen filesystem. A Dockerfile is the recipe for building one.

The consequence: your environment becomes a text file in version control. pip install disasters stay inside the container, and “it works on my machine” becomes a file you can hand over.

The minimum you need

Four commands:

docker build -t myproject .          # build an image from the Dockerfile here
docker run -it myproject             # run it, interactively
docker run -it -v $(pwd):/work myproject    # ...with the current dir mounted at /work
docker ps -a                         # what have I got running or left behind

And a Dockerfile for a typical GPU project:

FROM pytorch/pytorch:2.4.0-cuda12.1-cudnn9-runtime

WORKDIR /work
RUN apt-get update && apt-get install -y --no-install-recommends \
      git ffmpeg libgl1 libglib2.0-0 \
    && rm -rf /var/lib/apt/lists/*

COPY requirements.txt .
RUN pip install --no-cache-dir -r requirements.txt

COPY . .
CMD ["python", "demo.py"]

Five things in there are doing real work:

Start from an official PyTorch image with CUDA baked in. This removes the single most painful part of ML setup. The tag encodes the whole stack — 2.4.0-cuda12.1-cudnn9 — so the version question is answered in the first line.

ffmpeg libgl1 libglib2.0-0 are the three system packages that virtually every vision or video project needs and that nobody lists in their README. libGL.so.1: cannot open shared object file is the most common failure when containerising this kind of code, and libgl1 is the fix.

Copy requirements.txt before the rest of the code. Docker caches each layer. If requirements are copied and installed first, editing your Python does not re-run the install — which turns a rebuild from four minutes into four seconds. This one line is the difference between Docker being pleasant and being unbearable.

--no-cache-dir keeps the image from carrying pip’s download cache, which is often a gigabyte of nothing.

runtime not devel in the base image, unless you need to compile CUDA kernels. devel includes the full toolchain and is several gigabytes larger. If something needs nvcc at install time — custom ops, flash-attn, many splatting repos — you need devel.

The GPU flag nobody mentions

docker run --gpus all -it myproject

Without --gpus all the container cannot see your GPU, torch.cuda.is_available() returns False, and you will spend twenty minutes debugging your PyTorch install. It needs the NVIDIA Container Toolkit on the host, which is a one-time install.

Verify immediately:

docker run --gpus all -it myproject python -c \
  "import torch; print(torch.cuda.is_available(), torch.cuda.get_device_name(0))"

Mounting, which is where people get confused

A container’s filesystem disappears when it stops. Anything you want to keep must be on a mount:

docker run --gpus all -it \
  -v $(pwd)/data:/work/data \
  -v $(pwd)/output:/work/output \
  -v ~/.cache/huggingface:/root/.cache/huggingface \
  myproject

That third one is the one to copy. Mounting your Hugging Face cache means the container reuses models you have already downloaded instead of pulling forty gigabytes every time you rebuild.

Three creative-work-specific notes

GUI applications are awkward. Running something with a window from a container requires forwarding X11 or using a VNC image, and on macOS and Windows it is worse. For anything with an interface, Docker is the wrong tool. Containerise the headless part — the model, the renderer, the batch process — and keep the GUI native.

Serving is the natural pattern. The clean architecture for creative work is a container exposing an HTTP endpoint, with TouchDesigner, Max, Unity or a browser talking to it:

docker run --gpus all -p 8000:8000 myproject

Your creative tool stays native, the fragile Python dependency hell stays contained, and the interface between them is a URL. This is the same argument as our ONNX guide from a different direction — there you remove Python from deployment, here you quarantine it.

And it does not run on a Raspberry Pi the way you expect. ARM images exist but most ML images are x86-64 only. For an installation on small hardware, ONNX or a native build is the realistic path.

When not to bother

While you are iterating on a project you own. A venv or uv is lighter and faster, and rebuilding images to change one line is friction with no payoff.

For anything with a GUI, per above.

And if you only need it once. Docker pays off on the second or third time you come back to something. For a repo you will clone, run and delete, it is overhead.

Where it genuinely pays: a research repo with ugly dependencies you want to keep working; an installation you need to redeploy on a different machine in two years; anything you want to hand to a collaborator; and the model-serving pattern above.