ideogram-rate-limits
Implements exponential backoff and request queuing to manage Ideogram's concurrent request limits.
Install
mkdir -p .claude/skills/ideogram-rate-limits && curl -L -o skill.zip "https://agentskills.codes/api/skills/download/2738" && unzip -o skill.zip -d .claude/skills/ideogram-rate-limits && rm skill.zipInstalls to .claude/skills/ideogram-rate-limits
Activation
This is the description your AI agent reads to decide when to run this skill — the better it matches your request, the more reliably it fires.
Implement Ideogram rate limiting, backoff, and request queuing patterns.Key capabilities
- →Implement exponential backoff with jitter for API retries
- →Control concurrency of Ideogram API requests using a queue
- →Apply a token bucket algorithm for client-side rate limiting
- →Process batches of Ideogram requests with progress tracking
- →Handle HTTP 429 rate limit errors
How it works
The skill provides code for exponential backoff with jitter to retry failed requests, a concurrency-limited queue to manage in-flight requests, and a token bucket for client-side rate limiting. These mechanisms prevent exceeding Ideogram's API limits.
Inputs & outputs
When to use ideogram-rate-limits
- →Implement retry logic for Ideogram
- →Optimize image generation throughput
- →Handle API rate limit errors
- →Manage concurrent API requests
About this skill
Ideogram Concurrency and Queue Control
Overview
Keep paid image work inside a measurable concurrency envelope. Combine admission control, tenant fairness, operation deadlines, and endpoint-aware retry so a traffic burst does not become duplicate spend, an unbounded queue, or persistent 429 pressure.
Prerequisites
- Request arrival rate, latency objective, output count, endpoint mix, and cost ceiling.
- Queue ownership, per-tenant policy, cancellation behavior, and overload response.
- Current account capacity and an approved contact path for increased scale.
Current Contract
Ideogram's API overview documents a default limit of 10 in-flight requests and directs larger-capacity needs to [email protected]. This is a concurrency boundary, not permission to launch ten workers per process. Account-wide observed behavior remains authoritative.
Authentication
Workers inject IDEOGRAM_API_KEY server-side and send Api-Key only to https://api.ideogram.ai. Queue records and metrics must exclude keys, prompts, images, and response URLs.
Instructions
- Measure arrival rate, service time by endpoint, current in-flight count, queue age, timeout rate, and
429responses. - Place a shared account-level semaphore below the verified limit; start conservatively and reserve capacity for recovery probes.
- Add bounded per-tenant admission, queue length, age, output count, and total operation deadlines.
- Prefer asynchronous endpoints for work that cannot fit an interactive deadline and persist each
generation_id. - Retry only classified transient responses using server guidance when present, exponential backoff with jitter, and a strict attempt ceiling.
- Prevent ambiguous duplicate generation by reconciling known async identifiers before resubmission.
- Raise capacity only through the vendor path and a reviewed load plan; canary the new limit.
Tool Discipline
Use Read, Glob, and Grep for queue, worker, metric, and fixture inspection. Use Write and Edit for approved limit, test, or documentation changes. Invocation does not authorize load generation, quota negotiation, or increased production concurrency.
Approval Boundaries
Require owners for live load tests, spend, tenant-priority changes, queue dropping, capacity increases, and deployment. Document how queued work is cancelled or drained before changing the envelope.
Error Handling
- A
429is an admission-control signal; immediate retries amplify pressure. - Expired queue work should terminate before calling Ideogram, not after spending credit.
- Do not retry
400,401,422, unsafe output, or an already accepted async submission as if transient.
Output
Return verified limit source, chosen semaphore, queue and tenant bounds, deadline and retry policy, measured status counts, cost impact, test result, canary state, and rollback setting. Exclude content and credentials.
Examples
- Set eight shared worker permits, retain two for recovery probes, and cap each tenant below the account total.
- Report
peak_inflight=8; queue_p95=3s; 429=0; expired_before_submit=12; duplicates=0.
Validation
Use a deterministic queue simulator first, verify fairness and deadline expiration, inject 429 and timeout outcomes, and prove duplicate suppression. Run a paid load check only within the approved request and cost ceiling.
Resources
- Current first-party evidence map — use the dated endpoint, webhook, billing, team, and training links as the contract index for this workflow.
- Recheck the endpoint-specific page and current OpenAPI description before relying on an enum, limit, beta feature, or lifecycle claim.
- Record live observations as environment-specific evidence, not as universal vendor guarantees.
When not to use it
- →When Ideogram's default limit of 10 in-flight requests is sufficient
Prerequisites
Limitations
- →Ideogram enforces a default limit of 10 in-flight requests
- →Image generation takes 5-15 seconds per call
- →Higher limits require contacting Ideogram
How it compares
This skill provides specific code implementations for handling Ideogram's concurrent in-flight request rate limit, unlike a generic retry mechanism that might not account for the specific limit type.
Compared to similar skills
ideogram-rate-limits side by side with the closest alternatives in the catalog.
| Skill | Installs | Updated | Safety | Difficulty |
|---|---|---|---|---|
| ideogram-rate-limits (this skill) | 2 | 2mo | Review | Intermediate |
| groq-performance-tuning | 1 | 2mo | No flags | Intermediate |
| openrouter-streaming-setup | 1 | 2mo | Review | Intermediate |
| firecrawl-reliability-patterns | 3 | 2mo | Review | Advanced |
Try saying
Example prompts that trigger this skill in your AI assistant.
More by jeremylongshore
View all by jeremylongshore →You might also like
groq-performance-tuning
jeremylongshore
Optimize Groq API performance with caching, batching, and connection pooling. Use when experiencing slow API responses, implementing caching strategies, or optimizing request throughput for Groq integrations. Trigger with phrases like "groq performance", "optimize groq", "groq latency", "groq caching", "groq slow", "groq batch".
openrouter-streaming-setup
jeremylongshore
Implement streaming responses with OpenRouter. Use when building real-time chat interfaces or reducing time-to-first-token. Trigger with phrases like 'openrouter streaming', 'openrouter sse', 'stream response', 'real-time openrouter'.
firecrawl-reliability-patterns
jeremylongshore
Implement FireCrawl reliability patterns including circuit breakers, idempotency, and graceful degradation. Use when building fault-tolerant FireCrawl integrations, implementing retry strategies, or adding resilience to production FireCrawl services. Trigger with phrases like "firecrawl reliability", "firecrawl circuit breaker", "firecrawl idempotent", "firecrawl resilience", "firecrawl fallback", "firecrawl bulkhead".
linear-rate-limits
jeremylongshore
Handle Linear API rate limiting and quotas effectively. Use when dealing with rate limit errors, implementing throttling, or optimizing API usage patterns. Trigger with phrases like "linear rate limit", "linear throttling", "linear API quota", "linear 429 error", "linear request limits".
generating-grpc-services
jeremylongshore
Generate gRPC service definitions, stubs, and implementations from Protocol Buffers. Use when creating high-performance gRPC services. Trigger with phrases like "generate gRPC service", "create gRPC API", or "build gRPC server".
perplexity-performance-tuning
jeremylongshore
Optimize Perplexity API performance with caching, batching, and connection pooling. Use when experiencing slow API responses, implementing caching strategies, or optimizing request throughput for Perplexity integrations. Trigger with phrases like "perplexity performance", "optimize perplexity", "perplexity latency", "perplexity caching", "perplexity slow", "perplexity batch".