exa-rate-limits
Implement exponential backoff and retry patterns to handle Exa rate limiting effectively.
Install
mkdir -p .claude/skills/exa-rate-limits && curl -L -o skill.zip "https://agentskills.codes/api/skills/download/8158" && unzip -o skill.zip -d .claude/skills/exa-rate-limits && rm skill.zipInstalls to .claude/skills/exa-rate-limits
Activation
This is the description your AI agent reads to decide when to run this skill — the better it matches your request, the more reliably it fires.
Implement Exa rate limiting, exponential backoff, and request queuing.Key capabilities
- →Implement exponential backoff with jitter for 429 errors
- →Manage request concurrency using p-queue
- →Execute adaptive rate limiting based on success rates
- →Process batch requests with configurable delays
- →Retry logic for 429 and 5xx server errors
How it works
The skill provides wrappers and queue configurations to manage request frequency, ensuring calls stay within the 10 QPS limit. It uses exponential backoff with jitter to handle 429 errors and concurrency control to prevent burst rejections.
Inputs & outputs
When to use exa-rate-limits
- →Handle API rate limits and 429 errors
- →Implement exponential backoff retries
- →Optimize request throughput
- →Monitor API QPS consumption
About this skill
Exa Endpoint Rate-limit Control
Overview
Budget Exa throughput per endpoint and distinguish caller rate limiting from vendor overload and billing exhaustion. Treat credentials, queries, retrieved content, generated output, spend, and destructive state as separately governed boundaries.
Prerequisites
- The target repository, environment, Exa team, product surface, and accountable owner.
- The workload's data classification, latency and freshness promise, cost ceiling, and retention policy.
- Current first-party documentation plus credentials only for a narrowly approved live check.
Current Contract
Published defaults are 10 QPS for /search, 100 QPS for /contents, and 10 QPS for /answer, while enterprise limits may differ. A 429 can reflect API-key, team, or network limits and should honor Retry-After when present. A 503 SERVICE_OVERLOADED is a separate capacity signal.
Authentication
For normal REST work, inject EXA_API_KEY from an approved server-side secret manager and send it only as Authorization: Bearer to the configured first-party Exa API host. Team Management service keys, hosted MCP OAuth or enterprise managed authorization, and payment-protocol calls are separate trust models. Never print, commit, place in a URL, or expose a credential to an untrusted client.
Instructions
- Inventory endpoint mix, team limits, key-specific controls, concurrency, and burst shape.
- Set independent token buckets for Search, Contents, Answer, Agent, and administrative traffic.
- Bound queues and propagate deadlines instead of allowing hidden backlog growth.
- Honor Retry-After and use capped jitter only for retryable classes.
- Measure attempts, completions, throttles, overloads, queue age, and cost together.
- Request a limit change only with measured demand and an owner-approved capacity plan.
Tool Discipline
Use Read, Glob, and Grep to inspect repository code, configuration, fixtures, and evidence. Use Write and Edit only for approved implementation or documentation changes. Do not call Exa, run paid research, create or alter a Monitor, Webset, Agent run, Batch, team, member, API key, budget, webhook, or deployment merely because this skill was invoked.
Approval Boundaries
Require an accountable owner before live queries involving sensitive intent, production credentials, spend or rate-limit changes, forced live crawling, generated summaries, external delivery, deployment, member or key changes, schedule creation, or destructive cancellation, stopping, deletion, or revocation. Read-only repository inspection and synthetic offline validation do not authorize live vendor actions.
Failure Modes
- Adding keys does not prove that team or network capacity increased.
- Retrying 402 or invalid 400 requests wastes capacity and money.
- Global throttling can let a high-volume Contents job starve interactive Search.
Output
Return the operation scope, environment, team and product surface, authorization class, contract and policy decisions, deterministic validation results, content-free identifiers, status and cost counts, risks, cleanup or rollback state, and a concise pass or fail receipt. Exclude credentials, raw queries, prompts, presigned URLs, retrieved content, generated output, and customer-derived data unless separately approved.
Example
- Give interactive Search its own 10-QPS budget and run Contents backfills through a separately bounded queue.
- Finish with request or resource IDs, assertion counts, cost and terminal state, rollback or deletion status, and the decision owner; never reproduce secrets or retrieved content.
Validation
Rerun the smallest relevant deterministic test, compare actual behavior with the requested outcome and current first-party contract, verify sensitive fields are absent from evidence, and confirm deadlines, terminal state, downstream retention, and rollback before reporting success.
References
Review the dated first-party evidence map before relying on any endpoint, parameter, search type, price, limit, beta, compliance, identity, retry, or lifecycle claim.
When not to use it
- →When the API is not Exa
- →When simple synchronous requests are sufficient without throttling
Prerequisites
Limitations
- →Default limit is 10 QPS across all endpoints
- →Research API has a concurrent task limit
How it compares
Unlike manual retry loops, this approach incorporates jitter to prevent thundering herd problems and uses adaptive logic to dynamically adjust delays based on API performance.
Compared to similar skills
exa-rate-limits side by side with the closest alternatives in the catalog.
| Skill | Installs | Updated | Safety | Difficulty |
|---|---|---|---|---|
| exa-rate-limits (this skill) | 0 | 2mo | Review | Intermediate |
| deepgram-performance-tuning | 3 | 2mo | Review | Intermediate |
| graphql | 6 | 8mo | No flags | Advanced |
| guidewire-sdk-patterns | 2 | 2mo | Review | Advanced |
Try saying
Example prompts that trigger this skill in your AI assistant.
More by jeremylongshore
View all by jeremylongshore →You might also like
deepgram-performance-tuning
jeremylongshore
Optimize Deepgram API performance for faster transcription and lower latency. Use when improving transcription speed, reducing latency, or optimizing audio processing pipelines. Trigger with phrases like "deepgram performance", "speed up deepgram", "optimize transcription", "deepgram latency", "deepgram faster".
graphql
davila7
GraphQL gives clients exactly the data they need - no more, no less. One endpoint, typed schema, introspection. But the flexibility that makes it powerful also makes it dangerous. Without proper controls, clients can craft queries that bring down your server. This skill covers schema design, resolvers, DataLoader for N+1 prevention, federation for microservices, and client integration with Apollo/urql. Key insight: GraphQL is a contract. The schema is the API documentation. Design it carefully.
guidewire-sdk-patterns
jeremylongshore
Master Guidewire SDK patterns including Digital SDK, REST API Client, and Gosu best practices. Use when implementing integrations, building frontends with Jutro, or writing server-side Gosu code. Trigger with phrases like "guidewire sdk", "digital sdk", "jutro sdk", "guidewire patterns", "gosu best practices", "rest api client".
groq-performance-tuning
jeremylongshore
Optimize Groq API performance with caching, batching, and connection pooling. Use when experiencing slow API responses, implementing caching strategies, or optimizing request throughput for Groq integrations. Trigger with phrases like "groq performance", "optimize groq", "groq latency", "groq caching", "groq slow", "groq batch".
openrouter-streaming-setup
jeremylongshore
Implement streaming responses with OpenRouter. Use when building real-time chat interfaces or reducing time-to-first-token. Trigger with phrases like 'openrouter streaming', 'openrouter sse', 'stream response', 'real-time openrouter'.
perplexity-multi-env-setup
jeremylongshore
Configure Perplexity across development, staging, and production environments. Use when setting up multi-environment deployments, configuring per-environment secrets, or implementing environment-specific Perplexity configurations. Trigger with phrases like "perplexity environments", "perplexity staging", "perplexity dev prod", "perplexity environment setup", "perplexity config by env".