ID

ideogram-rate-limits

Implements exponential backoff and request queuing to manage Ideogram's concurrent request limits.

Install

mkdir -p .claude/skills/ideogram-rate-limits && curl -L -o skill.zip "https://agentskills.codes/api/skills/download/2738" && unzip -o skill.zip -d .claude/skills/ideogram-rate-limits && rm skill.zip

Installs to .claude/skills/ideogram-rate-limits

Activation

This is the description your AI agent reads to decide when to run this skill — the better it matches your request, the more reliably it fires.

Implement Ideogram rate limiting, backoff, and request queuing patterns.
72 charsno explicit “when” trigger
Intermediate

Key capabilities

  • →Implement exponential backoff with jitter for API retries
  • →Control concurrency of Ideogram API requests using a queue
  • →Apply a token bucket algorithm for client-side rate limiting
  • →Process batches of Ideogram requests with progress tracking
  • →Handle HTTP 429 rate limit errors

How it works

The skill provides code for exponential backoff with jitter to retry failed requests, a concurrency-limited queue to manage in-flight requests, and a token bucket for client-side rate limiting. These mechanisms prevent exceeding Ideogram's API limits.

Inputs & outputs

You give it
Ideogram API requests (e.g., image generation prompts)
You get back
Processed Ideogram API responses, managed within rate limits

When to use ideogram-rate-limits

  • →Implement retry logic for Ideogram
  • →Optimize image generation throughput
  • →Handle API rate limit errors
  • →Manage concurrent API requests

About this skill

Ideogram Concurrency and Queue Control

Overview

Keep paid image work inside a measurable concurrency envelope. Combine admission control, tenant fairness, operation deadlines, and endpoint-aware retry so a traffic burst does not become duplicate spend, an unbounded queue, or persistent 429 pressure.

Prerequisites

  • Request arrival rate, latency objective, output count, endpoint mix, and cost ceiling.
  • Queue ownership, per-tenant policy, cancellation behavior, and overload response.
  • Current account capacity and an approved contact path for increased scale.

Current Contract

Ideogram's API overview documents a default limit of 10 in-flight requests and directs larger-capacity needs to [email protected]. This is a concurrency boundary, not permission to launch ten workers per process. Account-wide observed behavior remains authoritative.

Authentication

Workers inject IDEOGRAM_API_KEY server-side and send Api-Key only to https://api.ideogram.ai. Queue records and metrics must exclude keys, prompts, images, and response URLs.

Instructions

  1. Measure arrival rate, service time by endpoint, current in-flight count, queue age, timeout rate, and 429 responses.
  2. Place a shared account-level semaphore below the verified limit; start conservatively and reserve capacity for recovery probes.
  3. Add bounded per-tenant admission, queue length, age, output count, and total operation deadlines.
  4. Prefer asynchronous endpoints for work that cannot fit an interactive deadline and persist each generation_id.
  5. Retry only classified transient responses using server guidance when present, exponential backoff with jitter, and a strict attempt ceiling.
  6. Prevent ambiguous duplicate generation by reconciling known async identifiers before resubmission.
  7. Raise capacity only through the vendor path and a reviewed load plan; canary the new limit.

Tool Discipline

Use Read, Glob, and Grep for queue, worker, metric, and fixture inspection. Use Write and Edit for approved limit, test, or documentation changes. Invocation does not authorize load generation, quota negotiation, or increased production concurrency.

Approval Boundaries

Require owners for live load tests, spend, tenant-priority changes, queue dropping, capacity increases, and deployment. Document how queued work is cancelled or drained before changing the envelope.

Error Handling

  • A 429 is an admission-control signal; immediate retries amplify pressure.
  • Expired queue work should terminate before calling Ideogram, not after spending credit.
  • Do not retry 400, 401, 422, unsafe output, or an already accepted async submission as if transient.

Output

Return verified limit source, chosen semaphore, queue and tenant bounds, deadline and retry policy, measured status counts, cost impact, test result, canary state, and rollback setting. Exclude content and credentials.

Examples

  • Set eight shared worker permits, retain two for recovery probes, and cap each tenant below the account total.
  • Report peak_inflight=8; queue_p95=3s; 429=0; expired_before_submit=12; duplicates=0.

Validation

Use a deterministic queue simulator first, verify fairness and deadline expiration, inject 429 and timeout outcomes, and prove duplicate suppression. Run a paid load check only within the approved request and cost ceiling.

Resources

  • Current first-party evidence map — use the dated endpoint, webhook, billing, team, and training links as the contract index for this workflow.
  • Recheck the endpoint-specific page and current OpenAPI description before relying on an enum, limit, beta feature, or lifecycle claim.
  • Record live observations as environment-specific evidence, not as universal vendor guarantees.

When not to use it

  • →When Ideogram's default limit of 10 in-flight requests is sufficient

Prerequisites

`IDEOGRAM_API_KEY` configured`p-queue` npm package (optional, for queue-based approach)

Limitations

  • →Ideogram enforces a default limit of 10 in-flight requests
  • →Image generation takes 5-15 seconds per call
  • →Higher limits require contacting Ideogram

How it compares

This skill provides specific code implementations for handling Ideogram's concurrent in-flight request rate limit, unlike a generic retry mechanism that might not account for the specific limit type.

Compared to similar skills

ideogram-rate-limits side by side with the closest alternatives in the catalog.

SkillInstallsUpdatedSafetyDifficulty
ideogram-rate-limits (this skill)22moReviewIntermediate
groq-performance-tuning12moNo flagsIntermediate
openrouter-streaming-setup12moReviewIntermediate
firecrawl-reliability-patterns32moReviewAdvanced

Try saying

Example prompts that trigger this skill in your AI assistant.

More by jeremylongshore

View all by jeremylongshore →

analyzing-logs

jeremylongshore

Analyze application logs to detect performance issues, identify error patterns, and improve stability by extracting key insights.

14123

ollama-setup

jeremylongshore

Configure auto-configure Ollama when user needs local LLM deployment, free AI alternatives, or wants to eliminate hosted API costs. Trigger phrases: "install ollama", "local AI", "free LLM", "self-hosted AI", "replace OpenAI", "no API costs". Use when appropriate context detected. Trigger with relevant phrases based on skill purpose.

1167

backtesting-trading-strategies

jeremylongshore

Backtest crypto and traditional trading strategies against historical data. Calculates performance metrics (Sharpe, Sortino, max drawdown), generates equity curves, and optimizes strategy parameters. Use when user wants to test a trading strategy, validate signals, or compare approaches. Trigger with phrases like "backtest strategy", "test trading strategy", "historical performance", "simulate trades", "optimize parameters", or "validate signals".

1071

generating-database-seed-data

jeremylongshore

Process this skill enables AI assistant to generate realistic test data and database seed scripts for development and testing environments. it uses faker libraries to create realistic data, maintains relational integrity, and allows configurable data volumes. u... Use when working with databases or data models. Trigger with phrases like 'database', 'query', or 'schema'.

1033

cursor-codebase-indexing

jeremylongshore

Execute set up and optimize Cursor codebase indexing. Triggers on "cursor index setup", "codebase indexing", "index codebase", "cursor semantic search". Use when working with cursor codebase indexing functionality. Trigger with phrases like "cursor codebase indexing", "cursor indexing", "cursor".

885

testing-mobile-apps

jeremylongshore

Execute mobile app testing on iOS and Android devices/simulators. Use when performing specialized testing. Trigger with phrases like "test mobile app", "run iOS tests", or "validate Android functionality".

810

You might also like

groq-performance-tuning

jeremylongshore

Optimize Groq API performance with caching, batching, and connection pooling. Use when experiencing slow API responses, implementing caching strategies, or optimizing request throughput for Groq integrations. Trigger with phrases like "groq performance", "optimize groq", "groq latency", "groq caching", "groq slow", "groq batch".

111

openrouter-streaming-setup

jeremylongshore

Implement streaming responses with OpenRouter. Use when building real-time chat interfaces or reducing time-to-first-token. Trigger with phrases like 'openrouter streaming', 'openrouter sse', 'stream response', 'real-time openrouter'.

111

firecrawl-reliability-patterns

jeremylongshore

Implement FireCrawl reliability patterns including circuit breakers, idempotency, and graceful degradation. Use when building fault-tolerant FireCrawl integrations, implementing retry strategies, or adding resilience to production FireCrawl services. Trigger with phrases like "firecrawl reliability", "firecrawl circuit breaker", "firecrawl idempotent", "firecrawl resilience", "firecrawl fallback", "firecrawl bulkhead".

36

linear-rate-limits

jeremylongshore

Handle Linear API rate limiting and quotas effectively. Use when dealing with rate limit errors, implementing throttling, or optimizing API usage patterns. Trigger with phrases like "linear rate limit", "linear throttling", "linear API quota", "linear 429 error", "linear request limits".

09

generating-grpc-services

jeremylongshore

Generate gRPC service definitions, stubs, and implementations from Protocol Buffers. Use when creating high-performance gRPC services. Trigger with phrases like "generate gRPC service", "create gRPC API", or "build gRPC server".

13

perplexity-performance-tuning

jeremylongshore

Optimize Perplexity API performance with caching, batching, and connection pooling. Use when experiencing slow API responses, implementing caching strategies, or optimizing request throughput for Perplexity integrations. Trigger with phrases like "perplexity performance", "optimize perplexity", "perplexity latency", "perplexity caching", "perplexity slow", "perplexity batch".

13

Search skills

Search the agent skills registry