HU

hugging-face-evaluation

Manages model evaluation data, benchmarks, and metadata for Hugging Face model cards.

Install

mkdir -p .claude/skills/hugging-face-evaluation-fiscfed9 && curl -L -o skill.zip "https://agentskills.codes/api/skills/download/14769" && unzip -o skill.zip -d .claude/skills/hugging-face-evaluation-fiscfed9 && rm skill.zip

Installs to .claude/skills/hugging-face-evaluation-fiscfed9

Activation

This is the description your AI agent reads to decide when to run this skill — the better it matches your request, the more reliably it fires.

Add and manage evaluation results in Hugging Face model cards. Supports extracting eval tables from README content, importing scores from Artificial Analysis API, and running custom model evaluations with vLLM/lighteval. Works with the model-index metadata format.
264 charsno explicit “when” triggerlonger than Claude Code's old 250-char listing cap (fine on current versions)
Intermediate

Key capabilities

  • Extract evaluation tables from README content
  • Import benchmark scores from Artificial Analysis API
  • Run custom model evaluations with vLLM or lighteval
  • Generate leaderboard-compatible model-index metadata
  • Manage evaluation results in Hugging Face model cards

How it works

This skill provides tools to add structured evaluation results to Hugging Face model cards by extracting tables from READMEs, importing scores from Artificial Analysis, or running custom evaluations.

Inputs & outputs

You give it
Hugging Face model card README content, Artificial Analysis API key, model evaluation results
You get back
structured evaluation results in Hugging Face model cards, model-index YAML format, pull requests with evaluation updates

When to use hugging-face-evaluation

  • Add benchmarks to model card
  • Import scores from Artificial Analysis
  • Run custom model evaluation pipelines
  • Generate leaderboard-compatible metadata

About this skill

Overview

This skill provides tools to add structured evaluation results to Hugging Face model cards. It supports multiple methods for adding evaluation data:

  • Extracting existing evaluation tables from README content
  • Importing benchmark scores from Artificial Analysis
  • Running custom model evaluations with vLLM or accelerate backends (lighteval/inspect-ai)

When to Use

  • You need to add structured evaluation results to a Hugging Face model card.
  • You want to import benchmark data or run custom evaluations with vLLM, lighteval, or inspect-ai.
  • You are preparing leaderboard-compatible model-index metadata for a model release.

Integration with HF Ecosystem

  • Model Cards: Updates model-index metadata for leaderboard integration
  • Artificial Analysis: Direct API integration for benchmark imports
  • Papers with Code: Compatible with their model-index specification
  • Jobs: Run evaluations directly on Hugging Face Jobs with uv integration
  • vLLM: Efficient GPU inference for custom model evaluation
  • lighteval: HuggingFace's evaluation library with vLLM/accelerate backends
  • inspect-ai: UK AI Safety Institute's evaluation framework

Version

1.3.0

Dependencies

Core Dependencies

  • huggingface_hub>=0.26.0
  • markdown-it-py>=3.0.0
  • python-dotenv>=1.2.1
  • pyyaml>=6.0.3
  • requests>=2.32.5
  • re (built-in)

Inference Provider Evaluation

  • inspect-ai>=0.3.0
  • inspect-evals
  • openai

vLLM Custom Model Evaluation (GPU required)

  • lighteval[accelerate,vllm]>=0.6.0
  • vllm>=0.4.0
  • torch>=2.0.0
  • transformers>=4.40.0
  • accelerate>=0.30.0

Note: vLLM dependencies are installed automatically via PEP 723 script headers when using uv run.

IMPORTANT: Using This Skill

⚠️ CRITICAL: Check for Existing PRs Before Creating New Ones

Before creating ANY pull request with --create-pr, you MUST check for existing open PRs:

uv run scripts/evaluation_manager.py get-prs --repo-id "username/model-name"

If open PRs exist:

  1. DO NOT create a new PR - this creates duplicate work for maintainers
  2. Warn the user that open PRs already exist
  3. Show the user the existing PR URLs so they can review them
  4. Only proceed if the user explicitly confirms they want to create another PR

This prevents spamming model repositories with duplicate evaluation PRs.


All paths are relative to the directory containing this SKILL.md file. Before running any script, first cd to that directory or use the full path.

Use --help for the latest workflow guidance. Works with plain Python or uv run:

uv run scripts/evaluation_manager.py --help
uv run scripts/evaluation_manager.py inspect-tables --help
uv run scripts/evaluation_manager.py extract-readme --help

Key workflow (matches CLI help):

  1. get-prs → check for existing open PRs first
  2. inspect-tables → find table numbers/columns
  3. extract-readme --table N → prints YAML by default
  4. add --apply (push) or --create-pr to write changes

Core Capabilities

1. Inspect and Extract Evaluation Tables from README

  • Inspect Tables: Use inspect-tables to see all tables in a README with structure, columns, and sample rows
  • Parse Markdown Tables: Accurate parsing using markdown-it-py (ignores code blocks and examples)
  • Table Selection: Use --table N to extract from a specific table (required when multiple tables exist)
  • Format Detection: Recognize common formats (benchmarks as rows, columns, or comparison tables with multiple models)
  • Column Matching: Automatically identify model columns/rows; prefer --model-column-index (index from inspect output). Use --model-name-override only with exact column header text.
  • YAML Generation: Convert selected table to model-index YAML format
  • Task Typing: --task-type sets the task.type field in model-index output (e.g., text-generation, summarization)

2. Import from Artificial Analysis

  • API Integration: Fetch benchmark scores directly from Artificial Analysis
  • Automatic Formatting: Convert API responses to model-index format
  • Metadata Preservation: Maintain source attribution and URLs
  • PR Creation: Automatically create pull requests with evaluation updates

3. Model-Index Management

  • YAML Generation: Create properly formatted model-index entries
  • Merge Support: Add evaluations to existing model cards without overwriting
  • Validation: Ensure compliance with Papers with Code specification
  • Batch Operations: Process multiple models efficiently

4. Run Evaluations on HF Jobs (Inference Providers)

  • Inspect-AI Integration: Run standard evaluations using the inspect-ai library
  • UV Integration: Seamlessly run Python scripts with ephemeral dependencies on HF infrastructure
  • Zero-Config: No Dockerfiles or Space management required
  • Hardware Selection: Configure CPU or GPU hardware for the evaluation job
  • Secure Execution: Handles API tokens safely via secrets passed through the CLI

5. Run Custom Model Evaluations with vLLM (NEW)

⚠️ Important: This approach is only possible on devices with uv installed and sufficient GPU memory. Benefits: No need to use hf_jobs() MCP tool, can run scripts directly in terminal When to use: User working in local device directly when GPU is available

Before running the script

  • check the script path
  • check uv is installed
  • check gpu is available with nvidia-smi

Running the script

uv run scripts/train_sft_example.py

Features

  • vLLM Backend: High-performance GPU inference (5-10x faster than standard HF methods)
  • lighteval Framework: HuggingFace's evaluation library with Open LLM Leaderboard tasks
  • inspect-ai Framework: UK AI Safety Institute's evaluation library
  • Standalone or Jobs: Run locally or submit to HF Jobs infrastructure

Usage Instructions

The skill includes Python scripts in scripts/ to perform operations.

Prerequisites

  • Preferred: use uv run (PEP 723 header auto-installs deps)
  • Or install manually: pip install huggingface-hub markdown-it-py python-dotenv pyyaml requests
  • Set HF_TOKEN environment variable with Write-access token
  • For Artificial Analysis: Set AA_API_KEY environment variable
  • .env is loaded automatically if python-dotenv is installed

Method 1: Extract from README (CLI workflow)

Recommended flow (matches --help):

# 1) Inspect tables to get table numbers and column hints
uv run scripts/evaluation_manager.py inspect-tables --repo-id "username/model"

# 2) Extract a specific table (prints YAML by default)
uv run scripts/evaluation_manager.py extract-readme \
  --repo-id "username/model" \
  --table 1 \
  [--model-column-index <column index shown by inspect-tables>] \
  [--model-name-override "<column header/model name>"]  # use exact header text if you can't use the index

# 3) Apply changes (push or PR)
uv run scripts/evaluation_manager.py extract-readme \
  --repo-id "username/model" \
  --table 1 \
  --apply       # push directly
# or
uv run scripts/evaluation_manager.py extract-readme \
  --repo-id "username/model" \
  --table 1 \
  --create-pr   # open a PR

Validation checklist:

  • YAML is printed by default; compare against the README table before applying.
  • Prefer --model-column-index; if using --model-name-override, the column header text must be exact.
  • For transposed tables (models as rows), ensure only one row is extracted.

Method 2: Import from Artificial Analysis

Fetch benchmark scores from Artificial Analysis API and add them to a model card.

Basic Usage:

AA_API_KEY="your-api-key" uv run scripts/evaluation_manager.py import-aa \
  --creator-slug "anthropic" \
  --model-name "claude-sonnet-4" \
  --repo-id "username/model-name"

With Environment File:

# Create .env file
echo "AA_API_KEY=your-api-key" >> .env
echo "HF_TOKEN=your-hf-token" >> .env

# Run import
uv run scripts/evaluation_manager.py import-aa \
  --creator-slug "anthropic" \
  --model-name "claude-sonnet-4" \
  --repo-id "username/model-name"

Create Pull Request:

uv run scripts/evaluation_manager.py import-aa \
  --creator-slug "anthropic" \
  --model-name "claude-sonnet-4" \
  --repo-id "username/model-name" \
  --create-pr

Method 3: Run Evaluation Job

Submit an evaluation job on Hugging Face infrastructure using the hf jobs uv run CLI.

Direct CLI Usage:

HF_TOKEN=$HF_TOKEN \
hf jobs uv run hf-evaluation/scripts/inspect_eval_uv.py \
  --flavor cpu-basic \
  --secret HF_TOKEN=$HF_TOKEN \
  -- --model "meta-llama/Llama-2-7b-hf" \
     --task "mmlu"

GPU Example (A10G):

HF_TOKEN=$HF_TOKEN \
hf jobs uv run hf-evaluation/scripts/inspect_eval_uv.py \
  --flavor a10g-small \
  --secret HF_TOKEN=$HF_TOKEN \
  -- --model "meta-llama/Llama-2-7b-hf" \
     --task "gsm8k"

Python Helper (optional):

uv run scripts/run_eval_job.py \
  --model "meta-llama/Llama-2-7b-hf" \
  --task "mmlu" \
  --hardware "t4-small"

Method 4: Run Custom Model Evaluation with vLLM

Evaluate custom HuggingFace models directly on GPU using vLLM or accelerate backends. These scripts are separate from inference provider scripts and run models locally on the job's hardware.

When to Use vLLM Evaluation (vs Inference Providers)

FeaturevLLM ScriptsInference Provider Scripts
Model accessAny HF modelModels with API endpoints
HardwareYour GPU (or HF Jobs GPU)Provider's infrastructure
CostHF Jobs compute costAPI usage fees
SpeedvLLM optimizedDepends on provider
OfflineYes (after download)No

Option A: lighteval with vLLM Backend

lighteval is HuggingFace's evaluation library, supporting Open LLM Leaderboard tasks.

Standalone (local GPU):

# Run MMLU 5-shot with vLLM
uv 

---

*Content truncated.*

When not to use it

  • When existing open pull requests for evaluation updates are not reviewed first
  • When the README does not contain markdown tables with numeric scores

Prerequisites

huggingface_hub>=0.26.0markdown-it-py>=3.0.0python-dotenv>=1.2.1pyyaml>=6.0.3

Limitations

  • Requires checking for existing PRs before creating new ones to avoid duplicates
  • Custom model evaluations with vLLM require a GPU and `uv` installed
  • Does not support model architectures not supported by vLLM

How it compares

This skill automates the process of adding structured evaluation results to Hugging Face model cards, integrating with external APIs and evaluation libraries, which is more efficient than manual data entry.

Compared to similar skills

hugging-face-evaluation side by side with the closest alternatives in the catalog.

SkillInstallsUpdatedSafetyDifficulty
hugging-face-evaluation (this skill)02moReviewIntermediate
quant-analyst1033moNo flagsAdvanced
umap-learn62moReviewIntermediate
embedding-strategies82moNo flagsIntermediate

Try saying

Example prompts that trigger this skill in your AI assistant.

You might also like

quant-analyst

zenobi-us

Expert quantitative analyst specializing in financial modeling, algorithmic trading, and risk analytics. Masters statistical methods, derivatives pricing, and high-frequency trading with focus on mathematical rigor, performance optimization, and profitable strategy development.

103355

umap-learn

K-Dense-AI

UMAP dimensionality reduction. Fast nonlinear manifold learning for 2D/3D visualization, clustering preprocessing (HDBSCAN), supervised/parametric UMAP, for high-dimensional data.

6100

embedding-strategies

wshobson

Select and optimize embedding models for semantic search and RAG applications. Use when choosing embedding models, implementing chunking strategies, or optimizing embedding quality for specific domains.

890

building-automl-pipelines

jeremylongshore

Build automated machine learning pipelines, including feature engineering, model selection, and performance evaluation.

688

model-compare

rawwerks

Compare 3D CAD models using boolean operations (IoU, Dice, precision/recall). Use when evaluating generated models against gold references, diffing CAD revisions, or computing similarity metrics for ML training. Triggers on: model diff, compare models, IoU, intersection over union, model similarity, CAD comparison, STEP diff, 3D evaluation, gold reference, generated model, precision recall 3D.

783

matchms

davila7

Mass spectrometry analysis. Process mzML/MGF/MSP, spectral similarity (cosine, modified cosine), metadata harmonization, compound ID, for metabolomics and MS data processing.

674

Search skills

Search the agent skills registry