HuggingFace Evaluation now maps to the huggingface-community-evals skill in the official huggingface/skills repository (skill folder). The folder was renamed from hugging-face-evaluation in March 2026. Its scope is deliberately narrow: run evaluations of Hugging Face Hub models on local hardware, choose the right framework and inference backend, and smoke-test before a full run.
It does not edit model cards, publish .eval_results, open pull requests, or orchestrate remote jobs. Those steps are handed off elsewhere.
Key Features
- Two frameworks:
inspect-aifor explicit task control,lightevalfor leaderboard-style task strings such asleaderboard|mmlu|5. - Three bundled scripts:
inspect_eval_uv.py(Inference Providers, no local GPU),inspect_vllm_uv.py(local GPU with vLLM or Transformers) andlighteval_vllm_uv.py(local GPU with vLLM or accelerate). - Backend fallbacks: prefer vLLM for throughput; switch to
--backend hfor--backend acceleratewhen vLLM does not support the architecture. - Smoke tests:
--limit 10for inspect-ai,--max-samples 10for lighteval. - Hardware guidance: under 3B on a consumer GPU or Apple Silicon, 3B to 13B on a stronger GPU, larger models on high-memory GPUs or remote compute.
Use Cases
- Checking a fine-tuned checkpoint on MMLU or GSM8K before publishing it.
- Comparing two small models on the same task set on your own GPU.
- Reproducing a leaderboard-style number locally.
Pricing
The skill is free and open source (Apache-2.0 repository). Local runs cost your own hardware; the Inference Providers path may incur provider usage on your Hugging Face account.
Getting Started
- Install HuggingFace CLI, then
hf skills add huggingface-community-evals. - Check prerequisites:
uv --version,HF_TOKENfor gated models, andnvidia-smifor local GPU runs. - Smoke test:
uv run scripts/inspect_eval_uv.py --model meta-llama/Llama-3.2-1B --task mmlu --limit 20. - Scale up only after the smoke test passes.
Limitation: for remote execution the skill still tells agents to hand off to a hugging-face-jobs skill, which Hugging Face removed in April 2026. Use hf jobs uv run from the CLI skill instead. Custom model code needs --trust-remote-code, which runs code from the model repo.
FAQ
Does it publish results to the Hub?
No. It stops after the evaluation run.
What if I have no GPU?
Use the Inference Providers script, or run the same script remotely with hf jobs.
Alternatives
- HuggingFace Model Trainer: trains the checkpoint you evaluate here.
- Garak: probes models for security failures rather than benchmark accuracy.
- Langfuse: evaluation for deployed LLM apps rather than raw checkpoints.
Conclusion
A focused, practical skill for local benchmarking with sane fallbacks. Pair it with the CLI skill for anything remote. More in the skills hub.
Comments
No comments yet. Be the first to comment!
Related Tools
Related Insights

Anthropic Subagent: The Multi-Agent Architecture Revolution
Deep dive into Anthropic multi-agent architecture design. Learn how Subagents break through context window limitations, achieve 90% performance improvements, and real-world applications in Claude Code.
How to pick an AI coding CLI for your repo
Match Claude Code, Codex CLI, OpenCode, Pi, omp, DeepSeek Harness, and Grok Build to the work you actually run. Docs rechecked 2026-10-06.
Skills + Hooks + Plugins: How Anthropic Redefined AI Coding Tool Extensibility
An in-depth analysis of Claude Code's trinity architecture of Skills, Hooks, and Plugins. Explore why this design is more advanced than GitHub Copilot and Cursor, and how it redefines AI coding tool extensibility through open standards.