openrouter-cost-controls
Configures and monitors cost limits, budgets, and spending alerts for OpenRouter API keys.
Install
mkdir -p .claude/skills/openrouter-cost-controls && curl -L -o skill.zip "https://agentskills.codes/api/skills/download/8514" && unzip -o skill.zip -d .claude/skills/openrouter-cost-controls && rm skill.zipInstalls to .claude/skills/openrouter-cost-controls
Activation
This is the description your AI agent reads to decide when to run this skill — the better it matches your request, the more reliably it fires.
Implement cost controls for OpenRouter API usage. Use when setting budgets,Key capabilities
- →Query current OpenRouter credit balance and limits for an API key
- →Provision new OpenRouter API keys with specific credit limits
- →Track per-request costs using pre-flight estimates and actual generation costs
- →Implement client-side budget enforcement middleware to reject over-budget requests
- →Schedule budget alert scripts to notify when credit balance is low
- →Optimize costs by using cost-saving model variants like :floor or :free
How it works
This skill uses OpenRouter's API to manage per-key credit limits and query balances, then implements client-side middleware to enforce spending caps and track costs.
Inputs & outputs
When to use openrouter-cost-controls
- →Set per-key spending limits
- →Monitor API credit balance
- →Implement budget alerts for API usage
- →Prevent API overspend in production
About this skill
OpenRouter Cost Controls
Overview
OpenRouter provides per-key credit limits, a credit balance API, and per-generation cost queries. Combined with client-side budget middleware, you can enforce hard spending caps at the key level and soft caps in your application. This skill covers key-level limits, per-request cost tracking, budget enforcement middleware, and alert systems.
Prerequisites
- An OpenRouter API key (
sk-or-v1-...) exported asOPENROUTER_API_KEY— see theopenrouter-install-authskill for setup - Python 3.8+ with the OpenAI SDK and
requests; Node.js 18+ for the TypeScript per-request cost logger in the references curl,jq, andbcfor the balance check and the Budget Alert Script- An OpenRouter management key exported as
OPENROUTER_MGMT_KEYif you provision per-key credit limits viaPOST /api/v1/keys
Instructions
- Query
GET /api/v1/auth/keyper Check Credit Balance to see credits used, the key's limit, remaining balance, free-tier status, and rate limit. - Provision scoped keys per Per-Key Credit Limits —
POST /api/v1/keyswith a dollarlimit(e.g. $50 forbackend-prod) using the management key, then list keys to review usage against limits per service. - Deploy the
BudgetEnforcerfrom Budget Enforcement Middleware:check_budget()rejects any request whose pre-flight estimate exceeds the per-request or daily cap, andrecord_cost()books the exact spend fromGET /api/v1/generation?id=. - Cut unit cost with Cost-Saving Model Variants:
:floorpicks the cheapest provider,:freecosts nothing where available, and the task-basedROUTINGtable sends classification to gpt-4o-mini and simple Q&A to Llama 3.1 8B. - Schedule the Budget Alert Script (cron) so a balance drop below the threshold fires an alert to Slack or PagerDuty.
- Set
max_tokenson every request and enable auto-topup per Enterprise Considerations to cap completion cost without risking an outage.
Check Credit Balance
# Current balance and limits
curl -s https://openrouter.ai/api/v1/auth/key \
-H "Authorization: Bearer $OPENROUTER_API_KEY" | jq '{
credits_used: .data.usage,
credit_limit: .data.limit,
remaining: ((.data.limit // 0) - .data.usage),
is_free_tier: .data.is_free_tier,
rate_limit: .data.rate_limit
}'
Per-Key Credit Limits
import os, requests
MGMT_KEY = os.environ["OPENROUTER_MGMT_KEY"] # Management key
# Create a key with a $50 credit limit
resp = requests.post(
"https://openrouter.ai/api/v1/keys",
headers={"Authorization": f"Bearer {MGMT_KEY}"},
json={"name": "backend-prod", "limit": 50.0},
)
new_key = resp.json()["data"]["key"] # sk-or-v1-...
# List all keys with their limits and usage
keys = requests.get(
"https://openrouter.ai/api/v1/keys",
headers={"Authorization": f"Bearer {MGMT_KEY}"},
).json()
for k in keys.get("data", []):
print(f"{k['name']}: ${k.get('usage', 0):.4f} / ${k.get('limit', 'unlimited')}")
Budget Enforcement Middleware
import os, time, requests
from openai import OpenAI
client = OpenAI(
base_url="https://openrouter.ai/api/v1",
api_key=os.environ["OPENROUTER_API_KEY"],
default_headers={"HTTP-Referer": "https://my-app.com", "X-Title": "my-app"},
)
class BudgetEnforcer:
"""Client-side budget enforcement with server-side cost verification."""
def __init__(self, daily_limit: float = 10.0, per_request_limit: float = 0.50):
self.daily_limit = daily_limit
self.per_request_limit = per_request_limit
self._daily_spend = 0.0
self._day = time.strftime("%Y-%m-%d")
def _reset_if_new_day(self):
today = time.strftime("%Y-%m-%d")
if today != self._day:
self._daily_spend = 0.0
self._day = today
def estimate_cost(self, model: str, prompt_tokens: int, max_tokens: int) -> float:
"""Pre-flight cost estimate using cached pricing."""
# Representative rates (fetch from /models in production)
RATES = {
"anthropic/claude-3.5-sonnet": (3.0, 15.0), # per 1M tokens
"openai/gpt-4o": (2.50, 10.0),
"openai/gpt-4o-mini": (0.15, 0.60),
"meta-llama/llama-3.1-8b-instruct": (0.06, 0.06),
}
prompt_rate, comp_rate = RATES.get(model, (3.0, 15.0))
return (prompt_tokens * prompt_rate / 1_000_000) + (max_tokens * comp_rate / 1_000_000)
def check_budget(self, model: str, prompt_tokens: int, max_tokens: int):
"""Raise if request would exceed budget."""
self._reset_if_new_day()
estimated = self.estimate_cost(model, prompt_tokens, max_tokens)
if estimated > self.per_request_limit:
raise ValueError(
f"Request estimated at ${estimated:.4f} exceeds per-request limit ${self.per_request_limit}"
)
if self._daily_spend + estimated > self.daily_limit:
raise ValueError(
f"Daily spend ${self._daily_spend:.4f} + request ${estimated:.4f} "
f"exceeds daily limit ${self.daily_limit}"
)
def record_cost(self, generation_id: str):
"""Record actual cost from generation endpoint."""
try:
gen = requests.get(
f"https://openrouter.ai/api/v1/generation?id={generation_id}",
headers={"Authorization": f"Bearer {os.environ['OPENROUTER_API_KEY']}"},
timeout=5,
).json()
cost = float(gen.get("data", {}).get("total_cost", 0))
self._daily_spend += cost
return cost
except Exception:
return 0.0
budget = BudgetEnforcer(daily_limit=25.0, per_request_limit=1.0)
Cost-Saving Model Variants
# :floor variant -- cheapest provider for a model
response = client.chat.completions.create(
model="anthropic/claude-3.5-sonnet:floor", # Cheapest provider
messages=[{"role": "user", "content": "Summarize this..."}],
max_tokens=500,
)
# :free variant -- free providers (where available)
response = client.chat.completions.create(
model="google/gemma-2-9b-it:free",
messages=[{"role": "user", "content": "Hello"}],
max_tokens=100,
)
# Route simple tasks to cheap models
ROUTING = {
"classification": "openai/gpt-4o-mini", # $0.15/$0.60 per 1M
"summarization": "anthropic/claude-3-haiku", # $0.25/$1.25 per 1M
"code_generation": "anthropic/claude-3.5-sonnet", # $3/$15 per 1M
"simple_qa": "meta-llama/llama-3.1-8b-instruct", # $0.06/$0.06 per 1M
}
Budget Alert Script
#!/bin/bash
# Alert when credits drop below threshold
THRESHOLD=5.0
REMAINING=$(curl -s https://openrouter.ai/api/v1/auth/key \
-H "Authorization: Bearer $OPENROUTER_API_KEY" | \
jq '((.data.limit // 0) - .data.usage)')
if (( $(echo "$REMAINING < $THRESHOLD" | bc -l) )); then
echo "ALERT: OpenRouter credits low: \$$REMAINING remaining"
# Send to Slack, PagerDuty, etc.
fi
Output
- A credit-balance JSON summary:
credits_used,credit_limit,remaining,is_free_tier, andrate_limitfor the active key - Newly provisioned API keys with hard dollar limits, plus a per-key
usage / limitlisting from the management API ValueErrorrejections from the middleware when a request would exceed the per-request or daily budget, and a running daily-spend total booked from actual generation costs- Alert lines such as
ALERT: OpenRouter credits low: $4.87 remainingwhenever the balance crosses the configured threshold
Examples
Check what's left on a key before turning on traffic:
curl -s https://openrouter.ai/api/v1/auth/key \
-H "Authorization: Bearer $OPENROUTER_API_KEY" \
| jq '{used: .data.usage, limit: .data.limit, remaining: (.data.limit - .data.usage)}'
Expected output:
{"used": 3.42, "limit": 50, "remaining": 46.58}
More worked examples, including thread-safe budget middleware and a TypeScript cost logger: references/examples.md.
Error Handling
| Error | Cause | Fix |
|---|---|---|
| 402 Payment Required | Credits exhausted | Top up at openrouter.ai/credits or use :free model |
| 402 Key limit reached | Per-key credit limit hit | Increase key limit or create new key |
| Budget middleware rejects | Client-side limit exceeded | Increase limit or optimize prompt tokens |
| Stale pricing data | Cached rates outdated | Refresh from /api/v1/models daily |
Enterprise Considerations
- Set per-key credit limits via management API to isolate blast radius per service/team
- Query
/api/v1/generation?id=after each request for exact cost auditing - Use
:floorvariant to automatically pick the cheapest provider for a model - Route simple tasks to budget models ($0.06/1M) and reserve premium models for complex tasks
- Set
max_tokenson every request to cap completion cost - Enable auto-topup in the dashboard to prevent production service interruptions
References
- Examples | Errors
- Credits | Key Provisioning
Prerequisites
Limitations
- →402 Payment Required error if credits are exhausted
- →402 Key limit reached error if per-key credit limit is hit
- →Stale pricing data if cached rates are not refreshed from /api/v1/models
How it compares
This skill provides a framework for both hard key-level and soft application-level budget enforcement for OpenRouter API usage, unlike relying solely on manual monitoring.
Compared to similar skills
openrouter-cost-controls side by side with the closest alternatives in the catalog.
| Skill | Installs | Updated | Safety | Difficulty |
|---|---|---|---|---|
| openrouter-cost-controls (this skill) | 0 | 26d | Caution | Intermediate |
| plaid-fintech | 3 | 6mo | No flags | Advanced |
| kisautotrade-kis-api | 0 | 1mo | No flags | Intermediate |
| mcp-builder | 136 | 3mo | Review | Advanced |
Try saying
Example prompts that trigger this skill in your AI assistant.
More by jeremylongshore
View all by jeremylongshore →You might also like
plaid-fintech
davila7
Expert patterns for Plaid API integration including Link token flows, transactions sync, identity verification, Auth for ACH, balance checks, webhook handling, and fintech compliance best practices. Use when: plaid, bank account linking, bank connection, ach, account aggregation.
kisautotrade-kis-api
CountJung
Korea Investment Securities KIS Open API bridge for Codex in the repository-owned trading app. Use for authentication, TR-ID, REST/WebSocket endpoints, order, balance, execution, overseas stock, paper trading, API errors, and official sample verification.
mcp-builder
anthropics
Guide for creating high-quality MCP (Model Context Protocol) servers that enable LLMs to interact with external services through well-designed tools. Use when building MCP servers to integrate external APIs or services, whether in Python (FastMCP) or Node/TypeScript (MCP SDK).
telegram-bot-builder
davila7
Expert in building Telegram bots that solve real problems - from simple automation to complex AI-powered bots. Covers bot architecture, the Telegram Bot API, user experience, monetization strategies, and scaling bots to thousands of users. Use when: telegram bot, bot api, telegram automation, chat bot telegram, tg bot.
stripe-integration
wshobson
Implement Stripe payment processing for robust, PCI-compliant payment flows including checkout, subscriptions, and webhooks. Use when integrating Stripe payments, building subscription systems, or implementing secure checkout flows.
langchain-architecture
wshobson
Design LLM applications using the LangChain framework with agents, memory, and tool integration patterns. Use when building LangChain applications, implementing AI agents, or creating complex LLM workflows.