Hf Jobs
Submit and manage Hugging Face Jobs, cloud GPU or CPU runs for training, fine-tuning, and batch inference. Use before any hf-jobs run, for hardware selection, cost estimates, GPU sandbox smoke tests, or when a job fails.
Dataset Audit
Audit a Hugging Face dataset before training or evaluation. Use before any training job, when choosing between datasets, or when a KeyError or format mismatch appears during training.
Hf Doctor
Diagnose the ML Intern environment, Hugging Face auth, uv, tokens, job pricing catalog, and data directory. Use at setup, when any HF call returns an auth error, or when jobs and sandboxes fail unexpectedly.
Ml Intern
ML engineering for the Hugging Face ecosystem. Use for fine-tuning, training, LoRA/SFT/DPO/GRPO, dataset prep, evaluation, inference, HF Jobs or GPU work, trackio monitoring, or any task touching TRL, Transformers, PEFT, diffusers, or Hub resources.
Ml Autopilot
Arm autonomous mode for hands-off or budgeted ML work. Use when the user asks for unattended training runs, gives a time budget, or wants iteration to continue without supervision.
Trackio
Configure trackio experiment monitoring for training runs, add alert callbacks, read alerts back between runs, and derive the next run's config from prior alerts. Use for any training job monitoring or when iterating on hyperparameters.