HuggingFace Datasets
Manage, load, and process datasets from HuggingFace Hub for machine learning training and evaluation.
Key Features
- Dataset loading and caching
- Data preprocessing
- Format conversion
- Dataset splitting
- Streaming support
Use Cases
ML training data preparation, data analysis, dataset curation
Comments
No comments yet. Be the first to comment!
Related Tools
HuggingFace Experiment Tracking
github.com/huggingface/skills
Track experiments, metrics, and model performance across training runs for reproducible AI research.
HuggingFace Evaluation
github.com/huggingface/skills
Model evaluation tools with standard metrics, benchmarks, and comprehensive performance analysis for AI models.
HuggingFace Jobs
github.com/huggingface/skills
Manage and orchestrate ML training jobs, deployments, and compute resources on HuggingFace infrastructure.
Related Insights
Hook OpenCode Go into Codex on Windows. Do Not Open a Second Toolkit.
A ChatGPT-signed Codex desktop app still shows mostly GPT in the picker. On Windows, enable only OpenCode Go and the Grok, GLM, Kimi, DeepSeek, and MiniMax models you already pay for appear in the same selector. Keys stay local. Native GPT stays put.
What Locks Codex Is the Picker, Not the Models
You already pay for OpenCode Go, Grok, and Z.ai, but the Codex picker still shows mostly GPT. The community project codex-router does not teach another install ritual. It puts subscriptions you already bought back into the selector, keeps keys on the machine, and leaves native GPT alone.
Seven AI Coding CLIs, Six Months: No Matter How Strong the Model, Work Needs Supervision
Claude Code, Codex, opencode, pi, omp and DeepSeek Harness all have personalities. After six months of deep use I run a division of labor: pi for the fastest cheapest reviews, omp for complex PRs, DeepSeek Harness on V4 Flash for high-frequency low-cost review, and Claude Code, Qoder and Cursor for writing. No matter how strong the model, work needs supervision — ideally from an independent third party.