groq-observability
Set up monitoring for Groq API calls using Prometheus metrics and latency tracking.
Install
mkdir -p .claude/skills/groq-observability && curl -L -o skill.zip "https://agentskills.codes/api/skills/download/8060" && unzip -o skill.zip -d .claude/skills/groq-observability && rm skill.zipInstalls to .claude/skills/groq-observability
Activation
This is the description your AI agent reads to decide when to run this skill — the better it matches your request, the more reliably it fires.
Set up observability for Groq integrations: latency histograms, tokenKey capabilities
- →Instrument Groq API calls to capture latency, tokens, queue time, and estimated cost
- →Register Prometheus metrics for latency, tokens, cost, errors, throughput, and rate limits
- →Parse rate limit headers from Groq responses into a gauge
- →Configure Prometheus alert rules for latency, rate limits, throughput, error rate, and cost
- →Emit structured JSON log lines per request for log aggregation
How it works
The skill instruments Groq API calls to capture performance metrics. It then feeds these metrics to Prometheus, tracks rate limits, and sets up alerts and structured logging.
Inputs & outputs
When to use groq-observability
- →Track Groq inference latency histograms
- →Monitor token throughput and rate limit gauges
- →Configure Prometheus alerts for latency degradation
- →Build Grafana dashboards for Groq health
About this skill
Groq Observability
Overview
Monitor Groq LPU inference for latency, token throughput, rate limit utilization, and cost. Groq's defining advantage is speed (280-560 tok/s), so latency degradation is the highest-priority signal. The API returns rich timing metadata (queue_time, prompt_time, completion_time) and rate limit headers on every response.
Prerequisites
- A Groq account with an API key exported as the
GROQ_API_KEYenvironment variable — thegroq-sdkclient reads it automatically (new Groq()). - Node.js with
groq-sdkandprom-clientinstalled (npm install groq-sdk prom-client). - A Prometheus scrape target and (optionally) Grafana for the dashboard panels.
Key Metrics to Track
| Metric | Type | Source | Why |
|---|---|---|---|
| TTFT (time to first token) | Histogram | Client-side timing | Groq's main value prop |
| Tokens/second | Gauge | usage.completion_time | Throughput degradation |
| Total latency | Histogram | Client-side timing | End-to-end performance |
| Rate limit remaining | Gauge | x-ratelimit-remaining-* headers | Prevent 429s |
| Token usage | Counter | usage.total_tokens | Cost attribution |
| Error rate by code | Counter | Error handler | Availability |
| Estimated cost | Counter | Tokens * model price | Budget tracking |
Instructions
Apply these six steps in order. Steps 1-2 are the core instrumentation loop — wrap the client, then feed a Prometheus instrument set from each call. Steps 3-6 add rate-limit tracking, alerting, structured logs, and dashboards on top. The lean client skeleton is below; the full code for every step lives in references/implementation.md.
- Instrumented client — wrap
groq.chat.completions.createso latency, tokens, queue time, and estimated cost are captured on the same path as the request (trackedCompletion). - Prometheus metrics — register a histogram (latency), counters (tokens, cost, errors), and gauges (throughput, rate-limit remaining), then feed them from
emitMetrics. - Rate limit header tracking — parse
x-ratelimit-remaining-*off every response into a gauge so you alert before a 429, not after. - Prometheus alert rules — ship latency/rate-limit/throughput/error/cost alerts tuned to Groq's sub-200ms, 280+ tok/s baseline.
- Structured request logging — emit one JSON line per request for log aggregation, preserving per-request detail metrics roll up.
- Dashboard panels — TTFT distribution, tokens/sec, rate-limit utilization, request volume, error rate, cost, and queue time.
import Groq from "groq-sdk";
const groq = new Groq(); // reads GROQ_API_KEY
async function trackedCompletion(model: string, messages: any[]) {
const start = performance.now();
const result = await groq.chat.completions.create({ model, messages });
const latencyMs = performance.now() - start;
const usage = result.usage!;
const metrics = {
model,
latencyMs: Math.round(latencyMs),
tokensPerSec: Math.round(usage.completion_tokens / ((usage as any).completion_time || latencyMs / 1000)),
totalTokens: usage.total_tokens,
};
emitMetrics(metrics); // -> Prometheus (Step 2)
return { result, metrics };
}
See references/implementation.md for the complete
GroqMetrics shape, pricing table, Prometheus instruments, rate-limit tracking,
alert rules, structured logging, and dashboard panel list.
Output
Applying the workflow produces:
- A
trackedCompletionwrapper that returns{ result, metrics }, wheremetricsis aGroqMetricsobject (latency, TTFT, tokens/sec, token counts, queue time, estimated cost). - A Prometheus metric set —
groq_latency_ms(histogram),groq_tokens_total/groq_cost_usd/groq_errors_total(counters), andgroq_tokens_per_second/groq_ratelimit_remaining(gauges). - Five alert rules (
GroqLatencyHigh,GroqRateLimitCritical,GroqThroughputDrop,GroqErrorRateHigh,GroqCostSpike). - A structured JSON log line per request and a 7-panel dashboard spec.
Examples
Instrument a single completion and emit a structured log line:
const { result, metrics } = await trackedCompletion(
"llama-3.3-70b-versatile",
[{ role: "user", content: "Summarize this incident report in two sentences." }]
);
logGroqRequest(metrics, result.id);
// metrics.tokensPerSec -> 310, metrics.estimatedCostUsd -> 0.000404
For a 429-guard using rate-limit headers and a dashboard health-reading table, see references/examples.md.
Error Handling
| Issue | Cause | Solution |
|---|---|---|
| 429 with high retry-after | RPM or TPM exhausted | Implement request queuing |
| Latency spike > 2s | Model overloaded or large prompt | Reduce prompt size or switch to lighter model |
| 503 Service Unavailable | Groq capacity issue | Enable fallback to alternative provider |
| Tokens/sec drop | Streaming disabled or large prompts | Enable streaming for better perceived performance |
Resources
- references/implementation.md — full code for all six observability steps.
- references/examples.md — worked instrumentation, 429-guard, and dashboard-reading examples.
- Groq API Reference (usage fields)
- Groq Rate Limits
- prom-client on npm
- For incident response procedures, see the
groq-incident-runbookskill.
Prerequisites
Limitations
- →429 errors with high retry-after can occur if RPM or TPM are exhausted
- →Latency spikes > 2s can indicate model overload or large prompts
- →503 Service Unavailable can occur due to Groq capacity issues
How it compares
This skill provides a dedicated observability framework for Groq, offering detailed performance insights beyond basic API logging.
Compared to similar skills
groq-observability side by side with the closest alternatives in the catalog.
| Skill | Installs | Updated | Safety | Difficulty |
|---|---|---|---|---|
| groq-observability (this skill) | 0 | 27d | No flags | Advanced |
| langfuse | 7 | 6mo | No flags | Intermediate |
| appinsights-instrumentation | 6 | 7mo | Review | Beginner |
| instantly-observability | 2 | 27d | Caution | Intermediate |
Try saying
Example prompts that trigger this skill in your AI assistant.
More by jeremylongshore
View all by jeremylongshore →You might also like
langfuse
davila7
Expert in Langfuse - the open-source LLM observability platform. Covers tracing, prompt management, evaluation, datasets, and integration with LangChain, LlamaIndex, and OpenAI. Essential for debugging, monitoring, and improving LLM applications in production. Use when: langfuse, llm observability, llm tracing, prompt management, llm evaluation.
appinsights-instrumentation
github
Instrument a webapp to send useful telemetry data to Azure App Insights
instantly-observability
jeremylongshore
Set up comprehensive observability for Instantly integrations with metrics, traces, and alerts. Use when implementing monitoring for Instantly operations, setting up dashboards, or configuring alerting for Instantly integration health. Trigger with phrases like "instantly monitoring", "instantly metrics", "instantly observability", "monitor instantly", "instantly alerts", "instantly tracing".
apollo-observability
jeremylongshore
Set up Apollo.io monitoring and observability. Use when implementing logging, metrics, tracing, and alerting for Apollo integrations. Trigger with phrases like "apollo monitoring", "apollo metrics", "apollo observability", "apollo logging", "apollo alerts".
fireflies-observability
jeremylongshore
Set up comprehensive observability for Fireflies.ai integrations with metrics, traces, and alerts. Use when implementing monitoring for Fireflies.ai operations, setting up dashboards, or configuring alerting for Fireflies.ai integration health. Trigger with phrases like "fireflies monitoring", "fireflies metrics", "fireflies observability", "monitor fireflies", "fireflies alerts", "fireflies tracing".
evernote-observability
jeremylongshore
Implement observability for Evernote integrations. Use when setting up monitoring, logging, tracing, or alerting for Evernote applications. Trigger with phrases like "evernote monitoring", "evernote logging", "evernote metrics", "evernote observability".