MO

model-debugging

It isolates service failures from user-related errors by analyzing log data from Cloudflare and Tinybird.

Install

mkdir -p .claude/skills/model-debugging && curl -L -o skill.zip "https://agentskills.codes/api/skills/download/4062" && unzip -o skill.zip -d .claude/skills/model-debugging && rm skill.zip

Installs to .claude/skills/model-debugging

Activation

This is the description your AI agent reads to decide when to run this skill — the better it matches your request, the more reliably it fires.

Debug and diagnose model errors in Pollinations services. Analyze logs, find error patterns, identify affected users. For taking action on user tiers, see tier-management skill.
177 charsno explicit “when” trigger
Advanced

Key capabilities

  • →Isolate 500-level backend errors
  • →Identify 4xx authentication/billing patterns
  • →Analyze request ID flows in logs
  • →Filter Tinybird/Cloudflare error data
  • →Correlate IP/user patterns with failures

How it works

It cross-references log metadata and status codes to filter out expected client-side errors (401/402) from genuine infrastructure performance issues.

Inputs & outputs

You give it
RequestID or error code pattern.
You get back
Diagnostic breakdown of service failure origins.

When to use model-debugging

  • →Analyze backend failure patterns
  • →Identify users affected by errors
  • →Diagnose model service request failures

About this skill

Model Debugging Skill

Use this skill when:

  • Investigating model failures, high error rates, or service issues
  • Finding users affected by errors (402 billing, 403 permissions, 500 backend)
  • Analyzing Tinybird/Cloudflare logs for patterns
  • Diagnosing specific request failures

Understanding Model Monitor Error Rates

Why does the Model Monitor show high error rates when models work fine manually?

The Model Monitor at https://monitor.pollinations.ai shows all real-world traffic, including:

  • 401 errors: Anonymous users without API keys (most common)
  • 402 errors: Users with insufficient pollen balance or exhausted API key budget
  • 403 errors: Users denied access to specific models (API key restrictions)
  • 400 errors: Invalid request parameters (e.g., openai-audio without modalities param)
  • 429 errors: Rate-limited requests
  • 500/504 errors: Actual backend failures (investigate these)

When you test manually with a valid secret key (sk_), you bypass auth/quota issues, so models appear to work fine.

Key insight: High 401/402/403/400 rates are expected from real-world usage. Focus investigation on 500/504 errors.

Also watch for slow-but-200 failures: a 200 that arrives after the client's own timeout never shows in error rates or p95 — it only shows in the per-user latency tail. See "Slow-but-200 / Client-Side Timeout" below.


Data Flow Architecture

User Request → enter.pollinations.ai (Cloudflare Worker)
                    ↓
              Logs to Cloudflare Workers Observability
                    ↓
              Events stored in D1 database
                    ↓
              Batched to Tinybird (async, 100-500 events)
                    ↓
              Model Monitor queries Tinybird (model_health.pipe)

Structured Logging: enter.pollinations.ai uses LogTape with:

  • requestId: Unique per request (passed to downstream via x-request-id header)
  • status, body: Full error response from downstream services
  • Context: method, routePath, userAgent, ipAddress

Quick Diagnostics

1. Check Model Monitor

View current model health at: https://monitor.pollinations.ai

2. Query Recent Errors from D1 Database

# Via enter.pollinations.ai worker (requires wrangler)
cd enter.pollinations.ai
npx wrangler d1 execute pollinations-db --remote --command "SELECT model_requested, response_status, error_message, COUNT(*) as count FROM event WHERE response_status >= 400 AND created_at > datetime('now', '-1 hour') GROUP BY model_requested, response_status, error_message ORDER BY count DESC LIMIT 20"

3. Capture Live Logs

enter.pollinations.ai (Cloudflare Worker)

cd enter.pollinations.ai
wrangler tail --format json | tee logs.jsonl
# Or with formatting:
wrangler tail --format json | npx tsx scripts/format-logs.ts

gen.pollinations.ai (image + text gateway)

Image and text generation now run inside the gen Cloudflare Worker (the legacy EC2 image-pollinations and text-pollinations services are decommissioned). Use wrangler tail from gen.pollinations.ai/:

cd gen.pollinations.ai
wrangler tail --format json | tee gen-logs.jsonl

Legacy anonymous image (OVH)

Anonymous traffic to image.pollinations.ai still terminates on the OVH host:

# Real-time logs
ssh -i ~/.ssh/id_rsa_ovh [email protected] "sudo journalctl -u image-pollinations -f"

# Last 3 minutes
ssh -i ~/.ssh/id_rsa_ovh [email protected] "sudo journalctl -u image-pollinations --since '3 minutes ago' --no-pager" > legacy-image-logs.txt

Common Error Patterns

Azure Content Safety DNS Failure

Error: getaddrinfo ENOTFOUND gptimagemain1-resource.cognitiveservices.azure.com Cause: Azure Content Safety resource deleted or misconfigured Impact: Fail-open (content proceeds without safety check) Fix: Create new Azure Content Safety resource and update .env:

AZURE_CONTENT_SAFETY_ENDPOINT=https://<new-resource>.cognitiveservices.azure.com/
AZURE_CONTENT_SAFETY_API_KEY=<new-key>

Azure Kontext Content Filter

Error: Content rejected due to sexual/hate/violence content detection Cause: Azure's content moderation blocking prompts/images Impact: 400 error returned to user Fix: User error - prompt violates content policy

Vertex AI Invalid Image

Error: Provided image is not valid Cause: User passing unsupported image URL (e.g., Google Drive links) Impact: 400 error returned to user Fix: User error - need direct image URL

Translation Service Down

Error: No active translate servers available Cause: Translation service unavailable Impact: Prompts not translated (non-fatal) Fix: Check translation service status

OpenAI Audio Invalid Voice

Error: Invalid value for audio.voice Cause: User requesting unsupported voice name Impact: 400 error returned to user Fix: User error - use supported voices: alloy, echo, fable, onyx, nova, shimmer, coral, verse, ballad, ash, sage, etc.

Oversized Text Seed Surfaced as 500

Error: 'seed' must be Integer, invalid request error, or a generic upstream 500 Cause: A client sent a seed above signed INT32 max (2147483647) to a strict provider Impact: The provider may misclassify invalid client input as 500, inflating model health errors Fix: Reject oversized seeds as 400 at gateway validation; group incidents by user, API key, and request shape before treating them as a model outage

Veo No Video Data

Error: No video data in response Cause: Vertex AI returned empty video response Impact: 500 error Fix: Check Vertex AI quota/status, may be transient

Slow-but-200 / Client-Side Timeout

Error: None server-side — 200, but users see a timeout or "no response" Cause: The latency tail is above the client's own timeout (e.g. Roblox HttpService ~30s) while status rates and p95 look fine Impact: Invisible on error dashboards; usually one user or one request shape (e.g. huge prompts) Fix: Query the tail per user — countIf(response_time>30000), max(response_time) (query in "Raw SQL Queries" below)

Instant Rejection from a Per-Key Queue/Rate Limit

Error: Queue full, or a 402/429 on the user's very first request Cause: An intended throttle, or leaked in-memory state (a slot never freed) Impact: User blocked while the backend is idle Fix: Compare the limiter's state with the backend's own load counters. Wait out the time window and send one request: if it clears, it was a live throttle. If a fix you predicted (e.g. a restart to clear memory) changes nothing, the diagnosis was wrong — re-check live state before trying a second fix. For IP-keyed limits, check how many distinct IPs the limiter actually sees: a proxy can collapse many clients onto one IP, which looks exactly like a leaked slot.

Stream/Usage Errors (missing usage, dropped terminal SSE event)

Error: Missing token usage (usage_missing) or a missing terminal event (e.g. [DONE]) Cause: Usually an upstream disconnect or transport error hidden downstream, not the provider omitting data Impact: A transport/parsing bug gets blamed on the provider Fix: Look at the last SSE chunks and the original transport error across provider → adapter → validator before blaming the provider; truncated logs often drop the end. If two components parse the same stream (e.g. validator and tracker), make sure they agree on end-of-stream — a stream ending with one trailing newline, not a blank line, is valid; a disagreement shows up as a phantom provider failure.

Health-Check False Exclusions (402/403 During Automated Probing)

Error: A monitor excuses a failing model as "client noise" (402/403) Cause: The same status can come from the probe's own wallet/auth or from the provider's credits, quota, or suspension Impact: Real outages get vetoed, or probe problems look like outages Fix: Before tuning thresholds, follow a few excluded cases through the decision log and find who actually failed — probe or provider


Environment Variables to Check

Image and text env vars now live in the gen Worker secrets (gen.pollinations.ai/secrets/{dev,staging,prod}.vars.json, SOPS-encrypted). Decrypt to inspect:

sops -d gen.pollinations.ai/secrets/prod.vars.json | jq 'keys[] | select(test("AZURE|GOOGLE|CLOUDFLARE|OPENAI"))'

Key variables:

  • AZURE_CONTENT_SAFETY_ENDPOINT - Azure Content Safety API endpoint
  • AZURE_CONTENT_SAFETY_API_KEY - Azure Content Safety API key
  • GOOGLE_PROJECT_ID - Google Cloud project for Vertex AI
  • AZURE_MYCELI_PROD_SWEDEN_API_KEY - Shared Azure API key (Kontext, GPT Image, GPT Image 1.5)

Updating Secrets

Secrets are stored encrypted with SOPS:

  • gen.pollinations.ai/secrets/{dev,staging,prod}.vars.json
  • enter.pollinations.ai/secrets/{dev,staging,prod}.vars.json

To update:

# Decrypt, edit, re-encrypt
sops gen.pollinations.ai/secrets/prod.vars.json

# Deploy to the gen Worker (secrets ship with the deploy)
cd gen.pollinations.ai && npm run deploy

Log Analysis Commands

# Count errors by type (against captured wrangler-tail JSON)
jq -r '.logs[]?.message[]? // .message? // empty' gen-logs.jsonl | grep -oE "(Azure Flux Kontext|Vertex AI|No active translate|getaddrinfo ENOTFOUND)" | sort | uniq -c | sort -rn

# Find content filter rejections
jq -r '.logs[]?.message[]? // .message? // empty' gen-logs.jsonl | grep -i "Content rejected" | sort | uniq -c

Model-Specific Debugging

ModelBackendCommon Issues
fluxAzure/ReplicateRate limits, content filter
kontextAzure Flux KontextContent filter (strict)
nanobananaVertex AI GeminiInvalid image URLs, content filter
seedream-proByteDance ARKNSFW filter, API key issues
veoVertex AIQuota, empty respons

Content truncated.

When not to use it

  • →Tier-management or billing adjustments
  • →Direct user support

Limitations

  • →Dependent on log visibility in Tinybird/Cloudflare
  • →Cannot resolve billing/tier issues directly

How it compares

It distinguishes between real-world usage noise and actionable backend system failures.

Compared to similar skills

model-debugging side by side with the closest alternatives in the catalog.

SkillInstallsUpdatedSafetyDifficulty
model-debugging (this skill)14moCautionAdvanced
langsmith-observability48moReviewIntermediate
error-diagnostics-smart-debug55moNo flagsAdvanced
debugging-toolkit-smart-debug45moNo flagsIntermediate

Try saying

Example prompts that trigger this skill in your AI assistant.

Search skills

Search the agent skills registry