HuggingFace Datasets logo

HuggingFace Datasets

Visit

Hugging Face's datasets skill drives the Dataset Viewer API: list splits, page rows, search, filter, get parquet URLs, and upload data or agent traces.

Share:
View alternatives

HuggingFace Datasets is the huggingface-datasets skill in the official huggingface/skills repository (skill folder). Since the March 2026 reorganization it is a Dataset Viewer skill: it teaches a coding agent to explore and extract data from Hub datasets through read-only API calls, plus a few documented ways to upload data. (An older dataset-creation skill with a similar name was removed in that reorganization.)

Everything goes through https://datasets-server.huggingface.co with plain GET requests, so the agent can inspect a dataset without downloading it.

Key Features

  • Discover structure: /is-valid, /splits for subsets and splits, /first-rows for a preview.
  • Page through rows: /rows with a 0-based offset and length up to 100, using num_rows_total and partial to drive continuation.
  • Search and filter: /search for text in string columns; /filter with a where predicate and optional orderby.
  • Files and stats: /parquet for shard URLs, /size, /statistics per column, and /croissant metadata when available.
  • Gated data: private or gated datasets need Authorization: Bearer <HF_TOKEN>.
  • Uploads: through the Hub UI, or npx -y @huggingface/hub upload datasets/<ns>/<repo> ./folder data (add --private).
  • Agent traces: raw Claude Code, Codex and Pi session JSONL files uploaded to a dataset get a trace viewer; the skill recommends private repos.

Use Cases

  • Checking columns and sample rows before committing to a training dataset.
  • Pulling only the rows that match a filter for a quick analysis.
  • Publishing your own agent session traces for review.

Pricing

Free and open source (Apache-2.0 repository). The Dataset Viewer API is public for public datasets.

Getting Started

  1. Install HuggingFace CLI, then hf skills add huggingface-datasets.
  2. Ask: "show me the splits and first rows of stanfordnlp/imdb".
  3. For SQL instead of REST, the CLI skill's hf datasets sql runs DuckDB against the same parquet files.

Limitation: row endpoints cap at about 100 rows per call, so large pulls mean many requests or a parquet download. Agent traces can contain prompts, file paths and secrets; the skill says to keep those repos private.

FAQ

Does it write to my dataset?

The viewer calls are read-only. Uploads are separate, explicit commands.

Can it read private datasets?

Yes, with a token that has access.

Alternatives

Conclusion

A lightweight way to look inside Hub datasets before you download gigabytes. More in the skills hub and the hugging-face tag.

Comments

No comments yet. Be the first to comment!