Research & Innovation

Hugging Face Datasets Tutorial: Publishing Your Own Archive So It's Actually Usable

A loader that works in one line, a card that tells people what they're holding, a licence you're entitled to grant, and the git-lfs and parquet details that decide whether any of it loads.

If you have been making work for a decade you probably have a dataset and do not call it one: a folder of scans, a run of generated variations, field recordings, motion capture takes, screenshots of every version of a project. The difference between that and a dataset is a layout something can load and a document explaining what it is.

This is how to publish one properly.

Orbit OCP on the upload mechanics — files, git and the README.

Step 0 — Decide whether you are entitled to publish it

Do this first, because it is the step that cannot be fixed later. Publishing is irreversible in practice: people clone datasets, mirror them, and train on them within hours.

Three questions, honestly:

Do you hold the rights to everything in it? Your own photographs, yes. Scraped images, no. Scans of other people’s work, no. A model’s outputs — contested, and depends on the model’s licence, which you should actually read.

Is anyone identifiable in it? Faces, voices, names, locations, handwriting. If so, you need their consent for this specific use, and “they agreed to be in the project” is not consent to be in a training set.

Can you say what the licence is? Not “it’s free” — the actual licence, by name. If you cannot, the dataset is not ready.

A dataset that answers these three clearly is more valuable than a larger one that does not, because the second is unusable by anyone who cares about provenance.

Step 1 — Install, and authenticate

pip install -U datasets huggingface_hub
hf auth login

A write-scoped token from your account settings. The older huggingface-cli login still works if you have an earlier huggingface_hub.

Step 2 — Pick a layout the loader already understands

This is where most self-published datasets fail, and the fix is free. The library recognises conventional structures and will build a loader for you if you use one. Deviate and every user has to write their own parsing code, which most will not do.

For images with labels, use folder-per-class:

data/
  train/
    cyanotype/   img001.jpg  img002.jpg
    gum-bichromate/   img101.jpg
  test/
    cyanotype/   img900.jpg

For anything with per-item metadata, put a metadata.csv or metadata.jsonl next to the files, with a file_name column pointing at each one:

{"file_name": "img001.jpg", "process": "cyanotype", "year": 2019, "exposure_min": 12}
{"file_name": "img002.jpg", "process": "cyanotype", "year": 2019, "exposure_min": 18}

For audio, the same pattern with file_name pointing at the wav or flac.

For tabular or text, a single train.parquet / test.parquet pair.

The rule to internalise: one row per item, one column per fact about it, and the file path as a column rather than as the only information. A folder of loose images with meaning encoded in the filenames is the thing to avoid — IMG_2019_cyano_12min_v3.jpg is data trapped in a string.

Step 3 — Convert to parquet if it is tabular or large

from datasets import Dataset
import pandas as pd

df = pd.read_csv("my_records.csv")
ds = Dataset.from_pandas(df)
ds.to_parquet("data/train.parquet")

Parquet is columnar and compressed, so a reader pulling one column does not read the whole file, and the Hub’s dataset viewer and streaming both depend on it. A 2 GB CSV becomes a few hundred megabytes and loads in a fraction of the time.

For image and audio datasets you generally leave the media as files and let parquet or the metadata file carry the records — embedding large binaries in parquet works but makes the files unwieldy.

Step 4 — Push it

from datasets import load_dataset

ds = load_dataset("imagefolder", data_dir="data")
ds.push_to_hub("your-username/alt-process-prints")

push_to_hub handles the repo creation, the upload, the parquet conversion and the LFS tracking. Pass private=True while you are still working on it.

For a large or awkward upload, the CLI is more robust than a Python session that might die two hours in:

hf upload your-username/alt-process-prints ./data --repo-type=dataset

Step 5 — Write the card, and write it for someone who was not there

The README.md in a dataset repo is the dataset card, and it is the whole difference between a resource and a pile. It starts with YAML frontmatter that drives the Hub’s indexing:

---
license: cc-by-4.0
task_categories:
  - image-classification
language:
  - en
size_categories:
  - 1K<n<10K
tags:
  - photography
  - alternative-process
pretty_name: Alternative Process Prints
---

Then prose. The questions people will actually have, in rough order of how often they go unanswered:

  • What is in it, concretely — how many items, what they are, what the columns mean. Units.
  • How it was collected — by whom, when, with what equipment, under what conditions. This is the section that lets someone judge whether your data suits their question.
  • What is not in it — the gaps, the biases, the processes you happened to shoot more of. Every dataset is a sample of something, and saying what it is a sample of is the single most useful paragraph in a card.
  • Known problems — mislabelled items, inconsistent lighting, the batch where the scanner profile changed.
  • Licence and consent — stated plainly, including who appears in it and on what basis.
  • How to cite it.

Writing the gaps and the problems section feels like undermining your own work and does the opposite: a card that names its limitations is one a careful person can use, and a card that does not is one they will not trust.

Step 6 — Version it, because people will pin to it

Any change to a published dataset silently changes other people’s results. Use tags:

cd alt-process-prints
git tag v1.0
git push origin v1.0

And tell users to pin:

ds = load_dataset("your-username/alt-process-prints", revision="v1.0")

A dataset repo is a git repo with LFS, so branches and tags work exactly as you would expect. Add a short changelog to the card and tag every substantive change.

The gotchas

git-lfs must be installed if you are using git directly rather than push_to_hub. Without it, large files get committed as pointers or as raw blobs and the repo breaks in ways that are confusing to diagnose.

git lfs install

There are per-file and per-repo limits, and they change — check the current Hub documentation rather than trusting a number from a tutorial. The practical approach for very large collections is sharded files of a few hundred megabytes each rather than one enormous archive: better for resumable downloads, streaming, and parallel reads.

Do not upload a zip of the whole thing. It defeats the viewer, streaming, partial downloads and the loader, and it is the single most common mistake in self-published datasets. Upload the files.

Check the dataset viewer after you push. If the Hub cannot render a preview, the loader cannot read your layout, and neither can your users. The viewer failing is the cheapest available signal that something is structurally wrong.

Streaming is the feature to test. load_dataset(..., streaming=True) lets someone work with a 500 GB dataset without downloading it. It only works if the layout and formats support it, and verifying it works is a two-line check that makes your dataset usable to people without the disk space.

Where to go next

  • Publish a Space alongside it. A small Gradio browser for your own dataset is an afternoon’s work and dramatically changes how many people look at it.
  • Think about the derived dataset. Captions, embeddings, segmentation masks over your own material are often more useful to others than the raw files, and are unambiguously yours to license.
  • Consider the archive as the work. Anna Ridler’s practice is the reference point here — her hand-made datasets are the artwork, not the input to one, and the labour of assembling and labelling is the subject rather than the overhead.