HuggingFace Model Trainer logo

HuggingFace Model Trainer

Visit

Hugging Face's LLM trainer skill fine-tunes models with TRL or Unsloth on Hugging Face Jobs GPUs: SFT, DPO, GRPO, reward models and GGUF export.

Share:
View alternatives

HuggingFace Model Trainer corresponds to the huggingface-llm-trainer skill in the official huggingface/skills repository (skill folder), renamed from hugging-face-model-trainer in March 2026. It teaches a coding agent to fine-tune language and vision-language models with TRL, or with Unsloth, on Hugging Face Jobs cloud GPUs, so you do not need a local GPU. Results are pushed to the Hub.

It is one of the most detailed skills in the repository, with about ten reference documents and example scripts.

Key Features

  • Training methods: SFT, DPO, GRPO and reward modeling, with references/training_methods.md explaining when to use which.
  • Unsloth path: the skill recommends Unsloth when GPU memory is tight, models are larger than 13B, or you train vision-language models.
  • Example scripts: train_sft_example.py, train_dpo_example.py, train_grpo_example.py and unsloth_sft_example.py, all UV scripts with inline dependencies.
  • Helpers: dataset_inspector.py validates a dataset's format before paying for GPUs, estimate_cost.py estimates cost, and convert_to_gguf.py exports for llama.cpp, Ollama or LM Studio.
  • Monitoring: every training script should include Trackio.
  • Guardrails: push_to_hub=True plus an HF_TOKEN secret, because the job environment is ephemeral; TRL configs use max_length, not max_seq_length.

Use Cases

  • Instruction-tuning a small open model on your own chat data.
  • Preference tuning with DPO after a first SFT pass.
  • Producing a GGUF file to run the result locally.

Pricing

The skill is free (Apache-2.0 repository). Training runs on Hugging Face Jobs, which the skill says requires a Pro, Team or Enterprise plan and bills by GPU type and time.

Getting Started

  1. Install HuggingFace CLI, then hf skills add huggingface-llm-trainer.
  2. Log in and make sure your token has write access.
  3. Validate the dataset first: uv run scripts/dataset_inspector.py --help.
  4. Ask for a small demo run (the skill suggests 50 to 100 examples on t4-small) before a real one.

Limitations and risks: the default 30 minute job timeout is too short for most training, and a job that times out loses all progress. The skill tells agents to submit jobs through a Hugging Face MCP hf_jobs() tool; without that server, run the same UV script with hf jobs uv run.

FAQ

Can it train on my own GPU?

The focus is Hugging Face Jobs; a macOS local training reference exists for small experiments.

Does it cover vision detection models?

No. Hugging Face ships a separate huggingface-vision-trainer skill for that.

Alternatives

  • Unsloth: the faster fine-tuning library, usable on your own hardware.
  • Modal: run your own training code on serverless GPUs.
  • HuggingFace Evaluation: benchmark the checkpoint afterward.

Conclusion

The most complete way to let an agent run fine-tuning on Hugging Face infrastructure, with real guardrails around cost and data loss. Pair with experiment tracking. More in the skills hub.

Comments

No comments yet. Be the first to comment!