perplexity-prod-checklist
A deployment checklist covering security, rate limiting, and monitoring for Perplexity Sonar API launches.
Install
mkdir -p .claude/skills/perplexity-prod-checklist && curl -L -o skill.zip "https://agentskills.codes/api/skills/download/4108" && unzip -o skill.zip -d .claude/skills/perplexity-prod-checklist && rm skill.zipInstalls to .claude/skills/perplexity-prod-checklist
Activation
This is the description your AI agent reads to decide when to run this skill — the better it matches your request, the more reliably it fires.
Execute Perplexity production deployment checklist for Sonar API integrations.Key capabilities
- →Implement retry logic with exponential backoff
- →Sanitize inputs to prevent PII exposure
- →Set up alerting rules for API failures
- →Implement graceful degradation between model tiers
How it works
The skill provides a checklist and code patterns for production deployment, focusing on secret management, error handling, and cost controls. It includes health check implementations to monitor API availability.
Inputs & outputs
When to use perplexity-prod-checklist
- →Prepare Perplexity integrations for production launch
- →Validate security practices for API keys
- →Implement robust retry logic for search calls
- →Check compliance for production search features
About this skill
Perplexity Production Checklist
Overview
Complete checklist for deploying Perplexity Sonar API integrations to production. Perplexity-specific concerns: every API call performs a live web search (variable latency), citations link to third-party sites (must validate), and costs scale per-request plus per-token.
Prerequisites
- Staging environment tested
- Production API key generated (separate from dev/staging)
- Monitoring configured
- Cost budget defined
Production Readiness Checklist
API Configuration
- Production
PERPLEXITY_API_KEYin secret manager (not env file) - Key starts with
pplx-and has credits loaded - Separate API keys for dev/staging/prod
- Base URL is
https://api.perplexity.ai(not localhost/proxy) - Model selection configured:
sonarfor fast,sonar-profor deep
Code Quality
- All search calls wrapped in retry with exponential backoff
- Rate limiting implemented (50 RPM default)
- Query sanitization strips PII before sending to Perplexity
- Citations parsed from response (not extracted from text)
-
max_tokensset on all requests (prevents runaway costs) - Timeouts configured: 15s for sonar, 30s for sonar-pro
- Error handling covers 401, 402, 429, 500+ status codes
- No hardcoded API keys in source code
Performance
- Result caching implemented for repeated queries
- Cache TTL appropriate: 30min for news, 4hrs for research, 24hrs for facts
- Streaming enabled for user-facing search (reduces perceived latency)
- Request queue prevents burst overload
-
search_domain_filterused where appropriate (reduces search time)
Monitoring
- Latency tracked per model (sonar ~2s, sonar-pro ~5s, deep-research ~30s)
- Error rate monitored (alert on >5% failure rate)
- Token usage tracked for cost projection
- Citation count per response logged (quality signal)
- 429 rate limit errors tracked with alert
Cost Controls
- Monthly budget cap set on API key
- Model routing: simple queries to
sonar, complex tosonar-pro -
max_tokenscapped per endpoint - Cache hit rate monitored (target >30%)
- Cost per query tracked by model
Graceful Degradation
async function searchWithFallback(query: string) {
try {
// Primary: sonar-pro for deep answers
return await perplexity.chat.completions.create({
model: "sonar-pro",
messages: [{ role: "user", content: query }],
max_tokens: 2048,
});
} catch (err: any) {
if (err.status === 429 || err.status >= 500) {
// Fallback: sonar for faster, cheaper response
return await perplexity.chat.completions.create({
model: "sonar",
messages: [{ role: "user", content: query }],
max_tokens: 512,
});
}
throw err;
}
}
Health Check Endpoint
app.get("/health/perplexity", async (req, res) => {
const start = Date.now();
try {
const response = await perplexity.chat.completions.create({
model: "sonar",
messages: [{ role: "user", content: "ping" }],
max_tokens: 5,
});
res.json({
status: "healthy",
latencyMs: Date.now() - start,
model: response.model,
});
} catch (err: any) {
res.status(503).json({
status: "unhealthy",
error: err.status || err.message,
latencyMs: Date.now() - start,
});
}
});
Alerting Rules
| Alert | Condition | Severity |
|---|---|---|
| API Unreachable | Health check fails 3x | P1 |
| High Error Rate | 429/5xx > 5% over 5min | P2 |
| High Latency | p95 > 15s for sonar | P2 |
| Budget Exceeded | Monthly cost > 80% cap | P2 |
| Auth Failure | Any 401/402 error | P1 |
Error Handling
| Issue | Cause | Solution |
|---|---|---|
| Variable latency | Web search per request | Set appropriate timeouts per model |
| Broken citations | Source pages changed | Validate citation URLs before displaying |
| Cost overrun | No model routing | Route simple queries to sonar |
| Rate limit spikes | Burst traffic | Queue requests with p-queue |
Output
- Production-ready Perplexity integration with all checks passing
- Health check endpoint for monitoring
- Graceful degradation from sonar-pro to sonar
- Alerting rules configured
Resources
Next Steps
For version upgrades, see perplexity-upgrade-migration.
Prerequisites
Limitations
- →Requires separate API keys for environments
- →Dependent on monitoring infrastructure
How it compares
It provides a specific production-readiness framework tailored to the unique cost and latency characteristics of the Perplexity Sonar API.
Compared to similar skills
perplexity-prod-checklist side by side with the closest alternatives in the catalog.
| Skill | Installs | Updated | Safety | Difficulty |
|---|---|---|---|---|
| perplexity-prod-checklist (this skill) | 1 | 27d | Review | Advanced |
| django-verification | 5 | 4mo | Review | Intermediate |
| deployment-validation-config-validate | 1 | 4mo | Review | Advanced |
| documenso-prod-checklist | 1 | 27d | Caution | Intermediate |
Try saying
Example prompts that trigger this skill in your AI assistant.
More by jeremylongshore
View all by jeremylongshore →You might also like
django-verification
affaan-m
Verification loop for Django projects: migrations, linting, tests with coverage, security scans, and deployment readiness checks before release or PR.
deployment-validation-config-validate
sickn33
You are a configuration management expert specializing in validating, testing, and ensuring the correctness of application configurations. Create comprehensive validation schemas, implement configurat
documenso-prod-checklist
jeremylongshore
Execute Documenso production deployment checklist and rollback procedures. Use when deploying Documenso integrations to production, preparing for launch, or implementing go-live procedures. Trigger with phrases like "documenso production", "deploy documenso", "documenso go-live", "documenso launch checklist".
gh-actions-validator
jeremylongshore
Validate use when validating GitHub Actions workflows for Google Cloud and Vertex AI deployments. Trigger with phrases like "validate github actions", "setup workload identity federation", "github actions security", "deploy agent with ci/cd", or "automate vertex ai deployment". Enforces Workload Identity Federation (WIF), validates OIDC permissions, ensures least privilege IAM, and implements security best practices.
ideogram-multi-env-setup
jeremylongshore
Configure Ideogram across development, staging, and production environments. Use when setting up multi-environment deployments, configuring per-environment secrets, or implementing environment-specific Ideogram configurations. Trigger with phrases like "ideogram environments", "ideogram staging", "ideogram dev prod", "ideogram environment setup", "ideogram config by env".
apollo-prod-checklist
jeremylongshore
Execute Apollo.io production deployment checklist. Use when preparing to deploy Apollo integrations to production, doing pre-launch verification, or auditing production readiness. Trigger with phrases like "apollo production checklist", "deploy apollo", "apollo go-live", "apollo production ready", "apollo launch checklist".