PE

perplexity-prod-checklist

A deployment checklist covering security, rate limiting, and monitoring for Perplexity Sonar API launches.

Install

mkdir -p .claude/skills/perplexity-prod-checklist && curl -L -o skill.zip "https://agentskills.codes/api/skills/download/4108" && unzip -o skill.zip -d .claude/skills/perplexity-prod-checklist && rm skill.zip

Installs to .claude/skills/perplexity-prod-checklist

Activation

This is the description your AI agent reads to decide when to run this skill — the better it matches your request, the more reliably it fires.

Execute Perplexity production deployment checklist for Sonar API integrations.
78 charsno explicit “when” trigger
Advanced

Key capabilities

  • Implement retry logic with exponential backoff
  • Sanitize inputs to prevent PII exposure
  • Set up alerting rules for API failures
  • Implement graceful degradation between model tiers

How it works

The skill provides a checklist and code patterns for production deployment, focusing on secret management, error handling, and cost controls. It includes health check implementations to monitor API availability.

Inputs & outputs

You give it
Production deployment configuration
You get back
Deployment readiness status and monitoring alerts

When to use perplexity-prod-checklist

  • Prepare Perplexity integrations for production launch
  • Validate security practices for API keys
  • Implement robust retry logic for search calls
  • Check compliance for production search features

About this skill

Perplexity Production Checklist

Overview

Complete checklist for deploying Perplexity Sonar API integrations to production. Perplexity-specific concerns: every API call performs a live web search (variable latency), citations link to third-party sites (must validate), and costs scale per-request plus per-token.

Prerequisites

  • Staging environment tested
  • Production API key generated (separate from dev/staging)
  • Monitoring configured
  • Cost budget defined

Production Readiness Checklist

API Configuration

  • Production PERPLEXITY_API_KEY in secret manager (not env file)
  • Key starts with pplx- and has credits loaded
  • Separate API keys for dev/staging/prod
  • Base URL is https://api.perplexity.ai (not localhost/proxy)
  • Model selection configured: sonar for fast, sonar-pro for deep

Code Quality

  • All search calls wrapped in retry with exponential backoff
  • Rate limiting implemented (50 RPM default)
  • Query sanitization strips PII before sending to Perplexity
  • Citations parsed from response (not extracted from text)
  • max_tokens set on all requests (prevents runaway costs)
  • Timeouts configured: 15s for sonar, 30s for sonar-pro
  • Error handling covers 401, 402, 429, 500+ status codes
  • No hardcoded API keys in source code

Performance

  • Result caching implemented for repeated queries
  • Cache TTL appropriate: 30min for news, 4hrs for research, 24hrs for facts
  • Streaming enabled for user-facing search (reduces perceived latency)
  • Request queue prevents burst overload
  • search_domain_filter used where appropriate (reduces search time)

Monitoring

  • Latency tracked per model (sonar ~2s, sonar-pro ~5s, deep-research ~30s)
  • Error rate monitored (alert on >5% failure rate)
  • Token usage tracked for cost projection
  • Citation count per response logged (quality signal)
  • 429 rate limit errors tracked with alert

Cost Controls

  • Monthly budget cap set on API key
  • Model routing: simple queries to sonar, complex to sonar-pro
  • max_tokens capped per endpoint
  • Cache hit rate monitored (target >30%)
  • Cost per query tracked by model

Graceful Degradation

async function searchWithFallback(query: string) {
  try {
    // Primary: sonar-pro for deep answers
    return await perplexity.chat.completions.create({
      model: "sonar-pro",
      messages: [{ role: "user", content: query }],
      max_tokens: 2048,
    });
  } catch (err: any) {
    if (err.status === 429 || err.status >= 500) {
      // Fallback: sonar for faster, cheaper response
      return await perplexity.chat.completions.create({
        model: "sonar",
        messages: [{ role: "user", content: query }],
        max_tokens: 512,
      });
    }
    throw err;
  }
}

Health Check Endpoint

app.get("/health/perplexity", async (req, res) => {
  const start = Date.now();
  try {
    const response = await perplexity.chat.completions.create({
      model: "sonar",
      messages: [{ role: "user", content: "ping" }],
      max_tokens: 5,
    });
    res.json({
      status: "healthy",
      latencyMs: Date.now() - start,
      model: response.model,
    });
  } catch (err: any) {
    res.status(503).json({
      status: "unhealthy",
      error: err.status || err.message,
      latencyMs: Date.now() - start,
    });
  }
});

Alerting Rules

AlertConditionSeverity
API UnreachableHealth check fails 3xP1
High Error Rate429/5xx > 5% over 5minP2
High Latencyp95 > 15s for sonarP2
Budget ExceededMonthly cost > 80% capP2
Auth FailureAny 401/402 errorP1

Error Handling

IssueCauseSolution
Variable latencyWeb search per requestSet appropriate timeouts per model
Broken citationsSource pages changedValidate citation URLs before displaying
Cost overrunNo model routingRoute simple queries to sonar
Rate limit spikesBurst trafficQueue requests with p-queue

Output

  • Production-ready Perplexity integration with all checks passing
  • Health check endpoint for monitoring
  • Graceful degradation from sonar-pro to sonar
  • Alerting rules configured

Resources

Next Steps

For version upgrades, see perplexity-upgrade-migration.

Prerequisites

Staging environment testedProduction API keyMonitoring configured

Limitations

  • Requires separate API keys for environments
  • Dependent on monitoring infrastructure

How it compares

It provides a specific production-readiness framework tailored to the unique cost and latency characteristics of the Perplexity Sonar API.

Compared to similar skills

perplexity-prod-checklist side by side with the closest alternatives in the catalog.

SkillInstallsUpdatedSafetyDifficulty
perplexity-prod-checklist (this skill)127dReviewAdvanced
django-verification54moReviewIntermediate
deployment-validation-config-validate14moReviewAdvanced
documenso-prod-checklist127dCautionIntermediate

Try saying

Example prompts that trigger this skill in your AI assistant.

More by jeremylongshore

View all by jeremylongshore

analyzing-logs

jeremylongshore

Analyze application logs to detect performance issues, identify error patterns, and improve stability by extracting key insights.

14123

ollama-setup

jeremylongshore

Configure auto-configure Ollama when user needs local LLM deployment, free AI alternatives, or wants to eliminate hosted API costs. Trigger phrases: "install ollama", "local AI", "free LLM", "self-hosted AI", "replace OpenAI", "no API costs". Use when appropriate context detected. Trigger with relevant phrases based on skill purpose.

1167

backtesting-trading-strategies

jeremylongshore

Backtest crypto and traditional trading strategies against historical data. Calculates performance metrics (Sharpe, Sortino, max drawdown), generates equity curves, and optimizes strategy parameters. Use when user wants to test a trading strategy, validate signals, or compare approaches. Trigger with phrases like "backtest strategy", "test trading strategy", "historical performance", "simulate trades", "optimize parameters", or "validate signals".

1071

generating-database-seed-data

jeremylongshore

Process this skill enables AI assistant to generate realistic test data and database seed scripts for development and testing environments. it uses faker libraries to create realistic data, maintains relational integrity, and allows configurable data volumes. u... Use when working with databases or data models. Trigger with phrases like 'database', 'query', or 'schema'.

1033

cursor-codebase-indexing

jeremylongshore

Execute set up and optimize Cursor codebase indexing. Triggers on "cursor index setup", "codebase indexing", "index codebase", "cursor semantic search". Use when working with cursor codebase indexing functionality. Trigger with phrases like "cursor codebase indexing", "cursor indexing", "cursor".

885

testing-mobile-apps

jeremylongshore

Execute mobile app testing on iOS and Android devices/simulators. Use when performing specialized testing. Trigger with phrases like "test mobile app", "run iOS tests", or "validate Android functionality".

810

You might also like

django-verification

affaan-m

Verification loop for Django projects: migrations, linting, tests with coverage, security scans, and deployment readiness checks before release or PR.

521

deployment-validation-config-validate

sickn33

You are a configuration management expert specializing in validating, testing, and ensuring the correctness of application configurations. Create comprehensive validation schemas, implement configurat

13

documenso-prod-checklist

jeremylongshore

Execute Documenso production deployment checklist and rollback procedures. Use when deploying Documenso integrations to production, preparing for launch, or implementing go-live procedures. Trigger with phrases like "documenso production", "deploy documenso", "documenso go-live", "documenso launch checklist".

12

gh-actions-validator

jeremylongshore

Validate use when validating GitHub Actions workflows for Google Cloud and Vertex AI deployments. Trigger with phrases like "validate github actions", "setup workload identity federation", "github actions security", "deploy agent with ci/cd", or "automate vertex ai deployment". Enforces Workload Identity Federation (WIF), validates OIDC permissions, ensures least privilege IAM, and implements security best practices.

11

ideogram-multi-env-setup

jeremylongshore

Configure Ideogram across development, staging, and production environments. Use when setting up multi-environment deployments, configuring per-environment secrets, or implementing environment-specific Ideogram configurations. Trigger with phrases like "ideogram environments", "ideogram staging", "ideogram dev prod", "ideogram environment setup", "ideogram config by env".

11

apollo-prod-checklist

jeremylongshore

Execute Apollo.io production deployment checklist. Use when preparing to deploy Apollo integrations to production, doing pre-launch verification, or auditing production readiness. Trigger with phrases like "apollo production checklist", "deploy apollo", "apollo go-live", "apollo production ready", "apollo launch checklist".

10

Search skills

Search the agent skills registry