mistral-rate-limits
Tools for managing Mistral AI workspace limits (RPM/TPM) and implementing intelligent retry logic.
Install
mkdir -p .claude/skills/mistral-rate-limits && curl -L -o skill.zip "https://agentskills.codes/api/skills/download/7586" && unzip -o skill.zip -d .claude/skills/mistral-rate-limits && rm skill.zipInstalls to .claude/skills/mistral-rate-limits
Activation
This is the description your AI agent reads to decide when to run this skill — the better it matches your request, the more reliably it fires.
Implement Mistral AI rate limiting, backoff, and request management.Key capabilities
- →Manage Mistral AI RPM and TPM limits
- →Implement token-aware rate limiting
- →Retry API calls with `Retry-After` headers
- →Wrap client for rate-limited requests
- →Route requests with model fallback
- →Batch embeddings with rate awareness
How it works
This skill manages Mistral AI API rate limits by tracking requests and tokens, implementing retry logic with `Retry-After` headers, and providing model fallback for throughput.
Inputs & outputs
When to use mistral-rate-limits
- →Handling 429 rate limit errors
- →Implementing token-aware request queuing
- →Managing workspace-level API budgets
- →Optimizing request throughput
About this skill
Mistral Rate and Backpressure Control
Overview
Treat provider limits as shared workspace capacity, not constants. Admit work against measured demand, preserve deadlines and fairness, and shed load before retry storms consume remaining budget.
Prerequisites
- Current workspace request, token, and spend limits from the Admin Panel.
- Per-operation deadline, priority, idempotency, and maximum-attempt policy.
- Metrics for admitted, queued, throttled, retried, completed, and abandoned work.
Current Contract
Mistral documents workspace-shared limits across API keys and reports current dimensions in the Admin Panel. Limits vary; never encode a numeric example as a default.
Authentication
Limit telemetry excludes keys, prompts, responses, and files. Admin inspection requires separate authorized identity.
Instructions
- Capture current workspace limits and evidence time as observations.
- Measure demand by operation, model, requests, tokens, and concurrency.
- Define shared admission buckets and bounded queues with tenant fairness.
- Honor explicit retry timing; otherwise use capped jitter only for safe transient failures.
- Stop when deadline, attempt, token, or spend budget ends; return typed overload.
- Load-test locally, then canary approved traffic and compare queue/SLO evidence.
Tool Discipline
Use Read, Glob, and Grep to inspect code, locks, configuration, tests, and evidence. Use Write and Edit only for approved repository changes. Invocation alone does not authorize network calls, paid usage, uploads, stateful resources, admin mutations, deployments, or deletion.
Approval Boundaries
Require approval for live load, limit increases, capacity changes, or fallback models. Additional keys do not create independent capacity.
Error Handling
- Per-process limiters oversubscribe shared capacity across replicas.
- Retries after the user deadline waste tokens and worsen overload.
- Fallback changes quality, residency, context, and cost.
Output
Return dated observed limits, demand profile, admission/retry policy, queue bounds, shed behavior, SLO evidence, and rollback.
Examples
- Prioritize interactive work over approved batch preparation.
- Return overload once the deadline cannot survive the queue.
Validation
Simulate bursts, replicas, 429, missing retry metadata, cancellation, deadline expiry, and budget exhaustion offline. Prove that admission resumes gradually after recovery.
Resources
- Current first-party evidence map — recheck dated sources before relying on mutable endpoints, models, limits, prices, preview status, or retention.
- Record live account observations as environment-specific evidence, not universal Mistral guarantees.
Prerequisites
Limitations
- →429 errors due to exceeded RPM or TPM
- →Inconsistent limits because all keys share workspace budget
- →Batch failures from too many tokens per batch
How it compares
This skill provides a token-aware rate limiter and model fallback specifically for Mistral AI, addressing both requests per minute and tokens per minute limits.
Compared to similar skills
mistral-rate-limits side by side with the closest alternatives in the catalog.
| Skill | Installs | Updated | Safety | Difficulty |
|---|---|---|---|---|
| mistral-rate-limits (this skill) | 1 | 2mo | Review | Advanced |
| fastapi-templates | 520 | 4mo | No flags | Intermediate |
| android-kotlin-development | 268 | 7mo | Review | Advanced |
| mcp-builder | 136 | 5mo | Review | Advanced |
Try saying
Example prompts that trigger this skill in your AI assistant.
More by jeremylongshore
View all by jeremylongshore →You might also like
fastapi-templates
wshobson
Create production-ready FastAPI projects with async patterns, dependency injection, and comprehensive error handling. Use when building new FastAPI applications or setting up backend API projects.
android-kotlin-development
aj-geddes
Develop native Android apps with Kotlin. Covers MVVM with Jetpack, Compose for modern UI, Retrofit for API calls, Room for local storage, and navigation architecture.
mcp-builder
anthropics
Guide for creating high-quality MCP (Model Context Protocol) servers that enable LLMs to interact with external services through well-designed tools. Use when building MCP servers to integrate external APIs or services, whether in Python (FastMCP) or Node/TypeScript (MCP SDK).
fastapi-pro
sickn33
Build high-performance async APIs with FastAPI, SQLAlchemy 2.0, and Pydantic V2. Master microservices, WebSockets, and modern Python async patterns. Use PROACTIVELY for FastAPI development, async optimization, or API architecture.
api-design-principles
wshobson
Master REST and GraphQL API design principles to build intuitive, scalable, and maintainable APIs that delight developers. Use when designing new APIs, reviewing API specifications, or establishing API design standards.
telegram-bot-builder
davila7
Expert in building Telegram bots that solve real problems - from simple automation to complex AI-powered bots. Covers bot architecture, the Telegram Bot API, user experience, monetization strategies, and scaling bots to thousands of users. Use when: telegram bot, bot api, telegram automation, chat bot telegram, tg bot.