OP

openrouter-rate-limits

Strategies and implementation patterns for managing OpenRouter rate limits and handling 429 throttle errors.

Install

mkdir -p .claude/skills/openrouter-rate-limits && curl -L -o skill.zip "https://agentskills.codes/api/skills/download/8782" && unzip -o skill.zip -d .claude/skills/openrouter-rate-limits && rm skill.zip

Installs to .claude/skills/openrouter-rate-limits

Activation

This is the description your AI agent reads to decide when to run this skill — the better it matches your request, the more reliably it fires.

Understand and handle OpenRouter rate limits. Use when hitting 429 errors,
74 chars✓ has a “when” trigger
Intermediate

Key capabilities

  • Query OpenRouter API key limits
  • Inspect rate limit headers from API responses
  • Configure OpenAI SDK for automatic retry with exponential backoff
  • Implement a client-side token bucket rate limiter
  • Process batch requests with rate awareness

How it works

The skill describes how to query OpenRouter for key limits and read rate limit headers. It provides Python code for configuring the OpenAI SDK to handle 429 errors with exponential backoff and implements a client-side token bucket for proactive rate limiting.

Inputs & outputs

You give it
OpenRouter API key and API requests
You get back
Managed API requests within rate limits and processed responses

When to use openrouter-rate-limits

  • Handle 429 rate limit errors
  • Implement retry logic for OpenRouter
  • Monitor API throughput
  • Configure high-throughput API systems

About this skill

OpenRouter Rate Limits

Overview

OpenRouter rate limits are per-key, not per-account. Free tier keys get lower limits; paid keys get higher limits that scale with credit balance. The OpenAI SDK has built-in retry with exponential backoff for 429 responses. Check your current limits via GET /api/v1/auth/key. Rate limit headers are returned on every response.

Prerequisites

  • An OpenRouter API key (sk-or-v1-...) exported as OPENROUTER_API_KEY — see the openrouter-install-auth skill for setup
  • curl and jq for querying your key's limits from GET /api/v1/auth/key
  • Python 3.8+ with the OpenAI SDK (sync OpenAI and AsyncOpenAI) plus the requests package for reading rate-limit headers directly
  • Awareness of your tier: free keys get 20 req/10s, keys with any credits 200 req/10s (see Rate Limit Tiers)

Instructions

  1. Query your key's limits via GET /api/v1/auth/key per Check Your Rate Limits — note rate_limit.requests and rate_limit.interval.
  2. Place yourself in the Rate Limit Tiers table, remembering free models carry separate daily caps (50 req/day free, 1000 req/day with $10+ credits).
  3. Inspect live headroom with check_rate_headers() per Read Rate Limit Headers — watch x-ratelimit-remaining and retry-after.
  4. Configure SDK retries per Retry Strategy with OpenAI SDK: max_retries=5, timeout=60.0; the SDK catches 429s and backs off with jitter automatically.
  5. Add the client-side TokenBucket limiter from Custom Rate Limiter, set below the server limit (e.g. 150 per 10s under a 200/10s cap) so you rarely hit 429 at all.
  6. For bulk jobs, use batch_with_rate_limit() per Batch Processing with Rate Awareness — staggered starts plus semaphore-capped concurrency instead of bursts.

Check Your Rate Limits

# Query current rate limit configuration for your key
curl -s https://openrouter.ai/api/v1/auth/key \
  -H "Authorization: Bearer $OPENROUTER_API_KEY" | jq '{
    label: .data.label,
    rate_limit: .data.rate_limit,
    is_free_tier: .data.is_free_tier,
    credits_used: .data.usage,
    credit_limit: .data.limit
  }'
# Example output:
# {
#   "label": "my-app-prod",
#   "rate_limit": {"requests": 200, "interval": "10s"},
#   "is_free_tier": false,
#   "credits_used": 12.34,
#   "credit_limit": 100
# }

Rate Limit Tiers

TierRequestsIntervalWho
Free (no credits)2010sNew accounts
Free (with credits)20010sAccounts with any credits
PaidHigherVariesBased on credit balance

Free models have separate limits: 50 req/day (free users), 1000 req/day (with $10+ credits).

Read Rate Limit Headers

import os
from openai import OpenAI
import requests as http_requests

# The OpenAI SDK abstracts headers, so use requests for direct access
def check_rate_headers():
    """Make a request and inspect rate limit headers."""
    resp = http_requests.post(
        "https://openrouter.ai/api/v1/chat/completions",
        headers={
            "Authorization": f"Bearer {os.environ['OPENROUTER_API_KEY']}",
            "Content-Type": "application/json",
            "HTTP-Referer": "https://my-app.com",
        },
        json={
            "model": "openai/gpt-4o-mini",
            "messages": [{"role": "user", "content": "hi"}],
            "max_tokens": 1,
        },
    )
    return {
        "status": resp.status_code,
        "x-ratelimit-limit": resp.headers.get("x-ratelimit-limit"),
        "x-ratelimit-remaining": resp.headers.get("x-ratelimit-remaining"),
        "x-ratelimit-reset": resp.headers.get("x-ratelimit-reset"),
        "retry-after": resp.headers.get("retry-after"),
    }

Retry Strategy with OpenAI SDK

from openai import OpenAI

# The SDK handles 429 retries automatically with exponential backoff
client = OpenAI(
    base_url="https://openrouter.ai/api/v1",
    api_key=os.environ["OPENROUTER_API_KEY"],
    max_retries=5,           # Default is 2; increase for high-throughput
    timeout=60.0,            # Per-request timeout
    default_headers={"HTTP-Referer": "https://my-app.com", "X-Title": "my-app"},
)

# The SDK will:
# 1. Catch 429 responses
# 2. Read Retry-After header
# 3. Wait with exponential backoff (+ jitter)
# 4. Retry up to max_retries times
response = client.chat.completions.create(
    model="anthropic/claude-3.5-sonnet",
    messages=[{"role": "user", "content": "Hello"}],
    max_tokens=200,
)

Custom Rate Limiter (Client-Side)

import time, threading
from collections import deque

class TokenBucket:
    """Client-side rate limiter to prevent hitting server limits."""

    def __init__(self, rate: int = 200, interval: float = 10.0):
        self.rate = rate           # Max requests per interval
        self.interval = interval
        self._timestamps = deque()
        self._lock = threading.Lock()

    def acquire(self, timeout: float = 30.0) -> bool:
        """Block until a request slot is available."""
        deadline = time.monotonic() + timeout
        while time.monotonic() < deadline:
            with self._lock:
                now = time.monotonic()
                # Remove timestamps outside the window
                while self._timestamps and now - self._timestamps[0] > self.interval:
                    self._timestamps.popleft()

                if len(self._timestamps) < self.rate:
                    self._timestamps.append(now)
                    return True

            time.sleep(0.1)  # Wait and retry
        return False  # Timed out

limiter = TokenBucket(rate=150, interval=10.0)  # Stay under 200 limit

def rate_limited_completion(messages, **kwargs):
    """Completion with client-side rate limiting."""
    if not limiter.acquire(timeout=30):
        raise TimeoutError("Rate limiter timeout")
    return client.chat.completions.create(messages=messages, **kwargs)

Batch Processing with Rate Awareness

import asyncio
from openai import AsyncOpenAI

async def batch_with_rate_limit(prompts: list[str], model="openai/gpt-4o-mini",
                                 max_concurrent=10, delay_between=0.05):
    """Process a batch of prompts with rate-aware concurrency."""
    semaphore = asyncio.Semaphore(max_concurrent)
    aclient = AsyncOpenAI(
        base_url="https://openrouter.ai/api/v1",
        api_key=os.environ["OPENROUTER_API_KEY"],
        max_retries=5,
        default_headers={"HTTP-Referer": "https://my-app.com", "X-Title": "my-app"},
    )

    async def process(prompt, idx):
        await asyncio.sleep(idx * delay_between)  # Stagger requests
        async with semaphore:
            response = await aclient.chat.completions.create(
                model=model,
                messages=[{"role": "user", "content": prompt}],
                max_tokens=200,
            )
            return response.choices[0].message.content

    return await asyncio.gather(*[process(p, i) for i, p in enumerate(prompts)])

Output

  • A key-limit snapshot from /api/v1/auth/key: label, rate_limit (requests + interval), is_free_tier, and credit usage
  • Per-request header readings from check_rate_headers(): x-ratelimit-limit, x-ratelimit-remaining, x-ratelimit-reset, retry-after
  • A rate-limited client: SDK auto-retry on 429 plus a TokenBucket that blocks (up to a timeout) instead of erroring
  • Ordered batch results from batch_with_rate_limit() produced without triggering a retry storm

Examples

Read your server-side limit, then size the client-side limiter under it:

curl -s https://openrouter.ai/api/v1/auth/key \
  -H "Authorization: Bearer $OPENROUTER_API_KEY" | jq '.data.rate_limit'
# {"requests": 200, "interval": "10s"}

With that 200/10s ceiling, configure TokenBucket(rate=150, interval=10.0) so steady-state traffic stays ~25% below the limit, and let the SDK's max_retries=5 absorb whatever bursts through. More worked examples: references/examples.md.

Error Handling

ErrorCauseFix
429 Too Many RequestsExceeded requests per intervalSDK auto-retries; increase max_retries
Retry stormMultiple clients retrying simultaneouslyAdd random jitter (0-1s) to retry delay
Silent throttlingResponses slow down before 429Monitor latency; proactively reduce rate
Free tier limit hit50 req/day on free modelsAdd credits ($10+) for 1000 req/day limit

Enterprise Considerations

  • Rate limits are per-key: use multiple keys to multiply effective throughput
  • The OpenAI SDK handles 429 retries automatically -- configure max_retries (default 2)
  • Implement client-side rate limiting to stay under limits proactively (cheaper than retries)
  • Free models have daily limits separate from the per-key rate limit
  • Monitor x-ratelimit-remaining headers to detect approaching limits before hitting 429
  • For batch workloads, use staggered concurrent requests rather than burst patterns

References

When not to use it

  • When OpenRouter API key is not available
  • When `curl` or `jq` are not installed for querying key limits

Prerequisites

An OpenRouter API key (`sk-or-v1-...`) exported as `OPENROUTER_API_KEY``curl` and `jq` for querying your key's limitsPython 3.8+ with the OpenAI SDK and `requests` package

Limitations

  • Rate limits are per-key, not per-account
  • Free tier keys get lower limits
  • Free models have separate daily limits

How it compares

This skill provides specific strategies for managing OpenRouter's per-key rate limits, including querying limits and using the OpenAI SDK's built-in retry mechanism, which differs from a generic rate limiting approach.

Compared to similar skills

openrouter-rate-limits side by side with the closest alternatives in the catalog.

SkillInstallsUpdatedSafetyDifficulty
openrouter-rate-limits (this skill)026dCautionIntermediate
openrouter-streaming-setup126dReviewIntermediate
serving-llms-vllm67moReviewAdvanced
exa-performance-tuning326dReviewIntermediate

Try saying

Example prompts that trigger this skill in your AI assistant.

More by jeremylongshore

View all by jeremylongshore

analyzing-logs

jeremylongshore

Analyze application logs to detect performance issues, identify error patterns, and improve stability by extracting key insights.

14123

ollama-setup

jeremylongshore

Configure auto-configure Ollama when user needs local LLM deployment, free AI alternatives, or wants to eliminate hosted API costs. Trigger phrases: "install ollama", "local AI", "free LLM", "self-hosted AI", "replace OpenAI", "no API costs". Use when appropriate context detected. Trigger with relevant phrases based on skill purpose.

1167

backtesting-trading-strategies

jeremylongshore

Backtest crypto and traditional trading strategies against historical data. Calculates performance metrics (Sharpe, Sortino, max drawdown), generates equity curves, and optimizes strategy parameters. Use when user wants to test a trading strategy, validate signals, or compare approaches. Trigger with phrases like "backtest strategy", "test trading strategy", "historical performance", "simulate trades", "optimize parameters", or "validate signals".

1071

generating-database-seed-data

jeremylongshore

Process this skill enables AI assistant to generate realistic test data and database seed scripts for development and testing environments. it uses faker libraries to create realistic data, maintains relational integrity, and allows configurable data volumes. u... Use when working with databases or data models. Trigger with phrases like 'database', 'query', or 'schema'.

1033

cursor-codebase-indexing

jeremylongshore

Execute set up and optimize Cursor codebase indexing. Triggers on "cursor index setup", "codebase indexing", "index codebase", "cursor semantic search". Use when working with cursor codebase indexing functionality. Trigger with phrases like "cursor codebase indexing", "cursor indexing", "cursor".

885

testing-mobile-apps

jeremylongshore

Execute mobile app testing on iOS and Android devices/simulators. Use when performing specialized testing. Trigger with phrases like "test mobile app", "run iOS tests", or "validate Android functionality".

810

You might also like

openrouter-streaming-setup

jeremylongshore

Implement streaming responses with OpenRouter. Use when building real-time chat interfaces or reducing time-to-first-token. Trigger with phrases like 'openrouter streaming', 'openrouter sse', 'stream response', 'real-time openrouter'.

111

serving-llms-vllm

davila7

Serves LLMs with high throughput using vLLM's PagedAttention and continuous batching. Use when deploying production LLM APIs, optimizing inference latency/throughput, or serving models with limited GPU memory. Supports OpenAI-compatible endpoints, quantization (GPTQ/AWQ/FP8), and tensor parallelism.

66

exa-performance-tuning

jeremylongshore

Optimize Exa API performance with caching, batching, and connection pooling. Use when experiencing slow API responses, implementing caching strategies, or optimizing request throughput for Exa integrations. Trigger with phrases like "exa performance", "optimize exa", "exa latency", "exa caching", "exa slow", "exa batch".

32

generating-grpc-services

jeremylongshore

Generate gRPC service definitions, stubs, and implementations from Protocol Buffers. Use when creating high-performance gRPC services. Trigger with phrases like "generate gRPC service", "create gRPC API", or "build gRPC server".

13

openrouter-caching-strategy

jeremylongshore

Implement response caching for OpenRouter efficiency. Use when optimizing costs or reducing latency for repeated queries. Trigger with phrases like 'openrouter cache', 'cache llm responses', 'openrouter redis', 'semantic caching'.

12

etag

aalmada

Use this skill for any request involving HTTP ETags, conditional requests, or optimistic concurrency in REST APIs: implementing/explaining ETag headers, preventing lost updates, designing cache validation or conditional GET/PUT/DELETE, explaining If-Match, If-None-Match, 304 Not Modified, or 412 Pre

00

Search skills

Search the agent skills registry