MI

mistral-rate-limits

Tools for managing Mistral AI workspace limits (RPM/TPM) and implementing intelligent retry logic.

Install

mkdir -p .claude/skills/mistral-rate-limits && curl -L -o skill.zip "https://agentskills.codes/api/skills/download/7586" && unzip -o skill.zip -d .claude/skills/mistral-rate-limits && rm skill.zip

Installs to .claude/skills/mistral-rate-limits

Activation

This is the description your AI agent reads to decide when to run this skill — the better it matches your request, the more reliably it fires.

Implement Mistral AI rate limiting, backoff, and request management.
68 charsno explicit “when” trigger
Advanced

Key capabilities

  • →Manage Mistral AI RPM and TPM limits
  • →Implement token-aware rate limiting
  • →Retry API calls with `Retry-After` headers
  • →Wrap client for rate-limited requests
  • →Route requests with model fallback
  • →Batch embeddings with rate awareness

How it works

This skill manages Mistral AI API rate limits by tracking requests and tokens, implementing retry logic with `Retry-After` headers, and providing model fallback for throughput.

Inputs & outputs

You give it
Mistral AI API requests
You get back
Mistral AI API responses with rate limit management

When to use mistral-rate-limits

  • →Handling 429 rate limit errors
  • →Implementing token-aware request queuing
  • →Managing workspace-level API budgets
  • →Optimizing request throughput

About this skill

Mistral Rate and Backpressure Control

Overview

Treat provider limits as shared workspace capacity, not constants. Admit work against measured demand, preserve deadlines and fairness, and shed load before retry storms consume remaining budget.

Prerequisites

  • Current workspace request, token, and spend limits from the Admin Panel.
  • Per-operation deadline, priority, idempotency, and maximum-attempt policy.
  • Metrics for admitted, queued, throttled, retried, completed, and abandoned work.

Current Contract

Mistral documents workspace-shared limits across API keys and reports current dimensions in the Admin Panel. Limits vary; never encode a numeric example as a default.

Authentication

Limit telemetry excludes keys, prompts, responses, and files. Admin inspection requires separate authorized identity.

Instructions

  1. Capture current workspace limits and evidence time as observations.
  2. Measure demand by operation, model, requests, tokens, and concurrency.
  3. Define shared admission buckets and bounded queues with tenant fairness.
  4. Honor explicit retry timing; otherwise use capped jitter only for safe transient failures.
  5. Stop when deadline, attempt, token, or spend budget ends; return typed overload.
  6. Load-test locally, then canary approved traffic and compare queue/SLO evidence.

Tool Discipline

Use Read, Glob, and Grep to inspect code, locks, configuration, tests, and evidence. Use Write and Edit only for approved repository changes. Invocation alone does not authorize network calls, paid usage, uploads, stateful resources, admin mutations, deployments, or deletion.

Approval Boundaries

Require approval for live load, limit increases, capacity changes, or fallback models. Additional keys do not create independent capacity.

Error Handling

  • Per-process limiters oversubscribe shared capacity across replicas.
  • Retries after the user deadline waste tokens and worsen overload.
  • Fallback changes quality, residency, context, and cost.

Output

Return dated observed limits, demand profile, admission/retry policy, queue bounds, shed behavior, SLO evidence, and rollback.

Examples

  • Prioritize interactive work over approved batch preparation.
  • Return overload once the deadline cannot survive the queue.

Validation

Simulate bursts, replicas, 429, missing retry metadata, cancellation, deadline expiry, and budget exhaustion offline. Prove that admission resumes gradually after recovery.

Resources

  • Current first-party evidence map — recheck dated sources before relying on mutable endpoints, models, limits, prices, preview status, or retention.
  • Record live account observations as environment-specific evidence, not universal Mistral guarantees.

Prerequisites

Mistral API key configuredUnderstanding of workspace tier (Experiment vs Scale)Application with retry infrastructure

Limitations

  • →429 errors due to exceeded RPM or TPM
  • →Inconsistent limits because all keys share workspace budget
  • →Batch failures from too many tokens per batch

How it compares

This skill provides a token-aware rate limiter and model fallback specifically for Mistral AI, addressing both requests per minute and tokens per minute limits.

Compared to similar skills

mistral-rate-limits side by side with the closest alternatives in the catalog.

SkillInstallsUpdatedSafetyDifficulty
mistral-rate-limits (this skill)12moReviewAdvanced
fastapi-templates5204moNo flagsIntermediate
android-kotlin-development2687moReviewAdvanced
mcp-builder1365moReviewAdvanced

Try saying

Example prompts that trigger this skill in your AI assistant.

More by jeremylongshore

View all by jeremylongshore →

analyzing-logs

jeremylongshore

Analyze application logs to detect performance issues, identify error patterns, and improve stability by extracting key insights.

14123

ollama-setup

jeremylongshore

Configure auto-configure Ollama when user needs local LLM deployment, free AI alternatives, or wants to eliminate hosted API costs. Trigger phrases: "install ollama", "local AI", "free LLM", "self-hosted AI", "replace OpenAI", "no API costs". Use when appropriate context detected. Trigger with relevant phrases based on skill purpose.

1167

backtesting-trading-strategies

jeremylongshore

Backtest crypto and traditional trading strategies against historical data. Calculates performance metrics (Sharpe, Sortino, max drawdown), generates equity curves, and optimizes strategy parameters. Use when user wants to test a trading strategy, validate signals, or compare approaches. Trigger with phrases like "backtest strategy", "test trading strategy", "historical performance", "simulate trades", "optimize parameters", or "validate signals".

1071

generating-database-seed-data

jeremylongshore

Process this skill enables AI assistant to generate realistic test data and database seed scripts for development and testing environments. it uses faker libraries to create realistic data, maintains relational integrity, and allows configurable data volumes. u... Use when working with databases or data models. Trigger with phrases like 'database', 'query', or 'schema'.

1033

cursor-codebase-indexing

jeremylongshore

Execute set up and optimize Cursor codebase indexing. Triggers on "cursor index setup", "codebase indexing", "index codebase", "cursor semantic search". Use when working with cursor codebase indexing functionality. Trigger with phrases like "cursor codebase indexing", "cursor indexing", "cursor".

885

testing-mobile-apps

jeremylongshore

Execute mobile app testing on iOS and Android devices/simulators. Use when performing specialized testing. Trigger with phrases like "test mobile app", "run iOS tests", or "validate Android functionality".

810

You might also like

fastapi-templates

wshobson

Create production-ready FastAPI projects with async patterns, dependency injection, and comprehensive error handling. Use when building new FastAPI applications or setting up backend API projects.

5201,086

android-kotlin-development

aj-geddes

Develop native Android apps with Kotlin. Covers MVVM with Jetpack, Compose for modern UI, Retrofit for API calls, Room for local storage, and navigation architecture.

268679

mcp-builder

anthropics

Guide for creating high-quality MCP (Model Context Protocol) servers that enable LLMs to interact with external services through well-designed tools. Use when building MCP servers to integrate external APIs or services, whether in Python (FastMCP) or Node/TypeScript (MCP SDK).

136215

fastapi-pro

sickn33

Build high-performance async APIs with FastAPI, SQLAlchemy 2.0, and Pydantic V2. Master microservices, WebSockets, and modern Python async patterns. Use PROACTIVELY for FastAPI development, async optimization, or API architecture.

79181

api-design-principles

wshobson

Master REST and GraphQL API design principles to build intuitive, scalable, and maintainable APIs that delight developers. Use when designing new APIs, reviewing API specifications, or establishing API design standards.

72170

telegram-bot-builder

davila7

Expert in building Telegram bots that solve real problems - from simple automation to complex AI-powered bots. Covers bot architecture, the Telegram Bot API, user experience, monetization strategies, and scaling bots to thousands of users. Use when: telegram bot, bot api, telegram automation, chat bot telegram, tg bot.

106130

Search skills

Search the agent skills registry