HuggingFace Datasets is the huggingface-datasets skill in the official huggingface/skills repository (skill folder). Since the March 2026 reorganization it is a Dataset Viewer skill: it teaches a coding agent to explore and extract data from Hub datasets through read-only API calls, plus a few documented ways to upload data. (An older dataset-creation skill with a similar name was removed in that reorganization.)
Everything goes through https://datasets-server.huggingface.co with plain GET requests, so the agent can inspect a dataset without downloading it.
Key Features
- Discover structure:
/is-valid,/splitsfor subsets and splits,/first-rowsfor a preview. - Page through rows:
/rowswith a 0-basedoffsetandlengthup to 100, usingnum_rows_totalandpartialto drive continuation. - Search and filter:
/searchfor text in string columns;/filterwith awherepredicate and optionalorderby. - Files and stats:
/parquetfor shard URLs,/size,/statisticsper column, and/croissantmetadata when available. - Gated data: private or gated datasets need
Authorization: Bearer <HF_TOKEN>. - Uploads: through the Hub UI, or
npx -y @huggingface/hub upload datasets/<ns>/<repo> ./folder data(add--private). - Agent traces: raw Claude Code, Codex and Pi session JSONL files uploaded to a dataset get a trace viewer; the skill recommends private repos.
Use Cases
- Checking columns and sample rows before committing to a training dataset.
- Pulling only the rows that match a filter for a quick analysis.
- Publishing your own agent session traces for review.
Pricing
Free and open source (Apache-2.0 repository). The Dataset Viewer API is public for public datasets.
Getting Started
- Install HuggingFace CLI, then
hf skills add huggingface-datasets. - Ask: "show me the splits and first rows of stanfordnlp/imdb".
- For SQL instead of REST, the CLI skill's
hf datasets sqlruns DuckDB against the same parquet files.
Limitation: row endpoints cap at about 100 rows per call, so large pulls mean many requests or a parquet download. Agent traces can contain prompts, file paths and secrets; the skill says to keep those repos private.
FAQ
Does it write to my dataset?
The viewer calls are read-only. Uploads are separate, explicit commands.
Can it read private datasets?
Yes, with a token that has access.
Alternatives
- HuggingFace CLI:
hf datasets parquetandhf datasets sqlfrom the terminal. - HuggingFace Tool Builder: wraps repeated dataset queries in a script.
- Hugging Face: the platform that hosts the data.
Conclusion
A lightweight way to look inside Hub datasets before you download gigabytes. More in the skills hub and the hugging-face tag.
Comments
No comments yet. Be the first to comment!
Related Tools
Related Insights

Anthropic Subagent: The Multi-Agent Architecture Revolution
Deep dive into Anthropic multi-agent architecture design. Learn how Subagents break through context window limitations, achieve 90% performance improvements, and real-world applications in Claude Code.
How to pick an AI coding CLI for your repo
Match Claude Code, Codex CLI, OpenCode, Pi, omp, DeepSeek Harness, and Grok Build to the work you actually run. Docs rechecked 2026-10-06.
Skills + Hooks + Plugins: How Anthropic Redefined AI Coding Tool Extensibility
An in-depth analysis of Claude Code's trinity architecture of Skills, Hooks, and Plugins. Explore why this design is more advanced than GitHub Copilot and Cursor, and how it redefines AI coding tool extensibility through open standards.