rebuild-leaderboard
Resets and re-runs model evaluations on the LatamBoard benchmark suite after data loss.
Install
mkdir -p .claude/skills/rebuild-leaderboard && curl -L -o skill.zip "https://agentskills.codes/api/skills/download/13796" && unzip -o skill.zip -d .claude/skills/rebuild-leaderboard && rm skill.zipInstalls to .claude/skills/rebuild-leaderboard
Activation
This is the description your AI agent reads to decide when to run this skill — the better it matches your request, the more reliably it fires.
Re-run all models on the LatamBoard leaderboard from scratch after data loss. Identifies which model configs exist, runs each one with the full latam_board task suite on the cluster, and publishes results to HuggingFace after each model so progress is never lost.Key capabilities
- →Confirm HF_TOKEN is set in the environment
- →Check the dataset name in the configuration file
- →Identify models with existing configurations in `configs/models/`
- →Run model evaluations one at a time using `benchy eval`
- →Publish results to HuggingFace after each model run
- →Create new model configurations for models without existing configs
How it works
The skill first checks environment and configuration, then iteratively runs model evaluations one by one, publishing results to HuggingFace after each model to ensure progress is saved.
Inputs & outputs
When to use rebuild-leaderboard
- →Restoring benchmark data
- →Running model evaluation suites
- →Publishing results to HuggingFace
About rebuild-leaderboard
This skill automates the re-evaluation of models on the LatamBoard leaderboard. It manages SSH-based cluster tasks and publishes updated summaries to HuggingFace after each individual model completion.
Re-run all models on the LatamBoard leaderboard from scratch after data loss. Identifies which model configs exist, runs each one with the full latam_board task suite on the cluster, and publishes results to HuggingFace after each model so progress is never lost.
When not to use it
- →When raw benchmark outputs are not gone
- →When the goal is not to re-evaluate all models on latamboard.surus.lat
- →When the HuggingFace dataset `mauroibz/leaderboard-results` is not the target
Prerequisites
Limitations
- →The skill is used when raw benchmark outputs are gone (e.g., cluster wipe).
- →It requires SSH access to `cluster.surus.ddns.net`.
- →Each model run takes a full GPU card and 30-90 minutes.
How it compares
This skill automates the entire re-evaluation process for the LatamBoard leaderboard, including environment checks, iterative model execution, and incremental publishing to HuggingFace, which is more reliable than a manual, all-at-once approa
Compared to similar skills
rebuild-leaderboard side by side with the closest alternatives in the catalog.
| Skill | Installs | Updated | Safety | Difficulty |
|---|---|---|---|---|
| rebuild-leaderboard (this skill) | 0 | 3mo | Review | Advanced |
| weights-and-biases | 3 | 8mo | Review | Intermediate |
| querying-mlflow-metrics | 0 | 7mo | Review | Beginner |
| quant-analyst | 103 | 4mo | No flags | Advanced |
Try saying
Example prompts that trigger this skill in your AI assistant.
More by surus-lat
View all by surus-lat →You might also like
weights-and-biases
davila7
Track ML experiments with automatic logging, visualize training in real-time, optimize hyperparameters with sweeps, and manage model registry with W&B - collaborative MLOps platform
querying-mlflow-metrics
alessandro9110
Fetches aggregated trace metrics (token usage, latency, trace counts, quality evaluations) from MLflow tracking servers. Triggers on requests to show metrics, analyze token usage, view LLM costs, check usage trends, or query trace statistics.
quant-analyst
zenobi-us
Expert quantitative analyst specializing in financial modeling, algorithmic trading, and risk analytics. Masters statistical methods, derivatives pricing, and high-frequency trading with focus on mathematical rigor, performance optimization, and profitable strategy development.
umap-learn
K-Dense-AI
UMAP dimensionality reduction. Fast nonlinear manifold learning for 2D/3D visualization, clustering preprocessing (HDBSCAN), supervised/parametric UMAP, for high-dimensional data.
embedding-strategies
wshobson
Select and optimize embedding models for semantic search and RAG applications. Use when choosing embedding models, implementing chunking strategies, or optimizing embedding quality for specific domains.
building-automl-pipelines
jeremylongshore
Build automated machine learning pipelines, including feature engineering, model selection, and performance evaluation.