mistral-observability
Provides instrumentation patterns for tracking Mistral AI costs, token usage, and latency.
Install
mkdir -p .claude/skills/mistral-observability && curl -L -o skill.zip "https://agentskills.codes/api/skills/download/4750" && unzip -o skill.zip -d .claude/skills/mistral-observability && rm skill.zipInstalls to .claude/skills/mistral-observability
Activation
This is the description your AI agent reads to decide when to run this skill — the better it matches your request, the more reliably it fires.
Set up comprehensive observability for Mistral AI with metrics, traces,Key capabilities
- →Instrument Mistral API client calls
- →Emit Prometheus metrics for requests and tokens
- →Configure alerting rules for latency and cost
- →Generate Grafana dashboard panels
- →Implement structured logging for observability
How it works
The skill provides an instrumented wrapper for the Mistral client that calculates duration, token usage, and costs, then pushes these metrics to a backend. It also includes YAML configurations for Prometheus alerting and Grafana dashboard definitions.
Inputs & outputs
When to use mistral-observability
- →Track token consumption
- →Monitor API latency
- →Calculate integration costs
- →Setup dashboard metrics
About this skill
Mistral AI Observability
Overview
Monitor Mistral AI API usage, latency, token consumption, error rates, and costs. Covers instrumented client wrapper, Prometheus metrics, Grafana dashboard panels, alerting rules, and structured logging.
Prerequisites
- Mistral API integration in production
- Prometheus or OpenTelemetry-compatible metrics backend
- Alerting system (Alertmanager, PagerDuty, or similar)
Instructions
Step 1: Instrumented Client Wrapper
import { Mistral } from '@mistralai/mistralai';
const PRICING: Record<string, { input: number; output: number }> = {
'mistral-small-latest': { input: 0.10, output: 0.30 },
'mistral-large-latest': { input: 0.50, output: 1.50 },
'codestral-latest': { input: 0.30, output: 0.90 },
'mistral-embed': { input: 0.10, output: 0 },
};
interface MetricsEvent {
model: string;
endpoint: string;
durationMs: number;
status: 'success' | 'error';
statusCode?: number;
inputTokens?: number;
outputTokens?: number;
costUsd?: number;
}
function emitMetrics(event: MetricsEvent): void {
// Push to your metrics backend (Prometheus, Datadog, etc.)
console.log(JSON.stringify({ type: 'mistral_metric', ...event }));
}
async function instrumentedChat(
client: Mistral,
model: string,
messages: any[],
options?: any,
) {
const start = performance.now();
try {
const response = await client.chat.complete({ model, messages, ...options });
const duration = Math.round(performance.now() - start);
const pricing = PRICING[model] ?? PRICING['mistral-small-latest'];
const pt = response.usage?.promptTokens ?? 0;
const ct = response.usage?.completionTokens ?? 0;
emitMetrics({
model,
endpoint: 'chat.complete',
durationMs: duration,
status: 'success',
inputTokens: pt,
outputTokens: ct,
costUsd: (pt / 1e6) * pricing.input + (ct / 1e6) * pricing.output,
});
return response;
} catch (error: any) {
emitMetrics({
model,
endpoint: 'chat.complete',
durationMs: Math.round(performance.now() - start),
status: 'error',
statusCode: error.status,
});
throw error;
}
}
Step 2: Prometheus Metrics
// Using prom-client
import { Counter, Histogram, Gauge } from 'prom-client';
const mistralRequests = new Counter({
name: 'mistral_requests_total',
help: 'Total Mistral API requests',
labelNames: ['model', 'endpoint', 'status'],
});
const mistralDuration = new Histogram({
name: 'mistral_request_duration_ms',
help: 'Mistral request duration in milliseconds',
labelNames: ['model', 'endpoint'],
buckets: [100, 250, 500, 1000, 2500, 5000, 10000],
});
const mistralTokens = new Counter({
name: 'mistral_tokens_total',
help: 'Total tokens consumed',
labelNames: ['model', 'direction'], // direction: input | output
});
const mistralCost = new Counter({
name: 'mistral_cost_usd_total',
help: 'Estimated cost in USD',
labelNames: ['model'],
});
const mistralErrors = new Counter({
name: 'mistral_errors_total',
help: 'Total Mistral errors',
labelNames: ['model', 'status_code'],
});
// Record metrics from instrumented wrapper
function recordPrometheusMetrics(event: MetricsEvent): void {
mistralRequests.inc({ model: event.model, endpoint: event.endpoint, status: event.status });
mistralDuration.observe({ model: event.model, endpoint: event.endpoint }, event.durationMs);
if (event.status === 'success') {
if (event.inputTokens) mistralTokens.inc({ model: event.model, direction: 'input' }, event.inputTokens);
if (event.outputTokens) mistralTokens.inc({ model: event.model, direction: 'output' }, event.outputTokens);
if (event.costUsd) mistralCost.inc({ model: event.model }, event.costUsd);
} else {
mistralErrors.inc({ model: event.model, status_code: String(event.statusCode ?? 'unknown') });
}
}
Step 3: Alerting Rules
# prometheus/mistral-alerts.yaml
groups:
- name: mistral
rules:
- alert: MistralHighErrorRate
expr: rate(mistral_errors_total[5m]) / rate(mistral_requests_total[5m]) > 0.05
for: 5m
labels: { severity: critical }
annotations:
summary: "Mistral error rate exceeds 5%"
runbook: "See mistral-incident-runbook skill"
- alert: MistralHighLatency
expr: histogram_quantile(0.95, rate(mistral_request_duration_ms_bucket[5m])) > 5000
for: 5m
labels: { severity: warning }
annotations:
summary: "Mistral P95 latency exceeds 5 seconds"
- alert: MistralRateLimited
expr: rate(mistral_errors_total{status_code="429"}[5m]) > 0
for: 2m
labels: { severity: warning }
annotations:
summary: "Mistral rate limiting detected"
- alert: MistralCostSpike
expr: increase(mistral_cost_usd_total[1h]) > 10
labels: { severity: warning }
annotations:
summary: "Mistral spend exceeds $10/hour"
- alert: MistralAuthFailure
expr: increase(mistral_errors_total{status_code="401"}[5m]) > 0
labels: { severity: critical }
annotations:
summary: "Mistral authentication failing — API key may be revoked"
Step 4: Grafana Dashboard Panels
Key panels to create:
| Panel | Query | Type |
|---|---|---|
| Request Rate | rate(mistral_requests_total[5m]) | Time series |
| P50/P95/P99 Latency | histogram_quantile(0.95, rate(..._bucket[5m])) | Time series |
| Token Velocity | rate(mistral_tokens_total{direction="output"}[5m]) | Time series |
| Hourly Cost | increase(mistral_cost_usd_total[1h]) | Stat |
| Error Rate | rate(mistral_errors_total[5m]) by status_code | Time series |
| Model Distribution | sum by (model) (rate(mistral_requests_total[5m])) | Pie chart |
Step 5: Structured Log Format
interface MistralLogEntry {
ts: string;
level: 'info' | 'warn' | 'error';
model: string;
endpoint: string;
durationMs: number;
inputTokens?: number;
outputTokens?: number;
costUsd?: number;
status: string;
statusCode?: number;
requestId?: string;
}
function logMistralRequest(entry: MistralLogEntry): void {
// Ship to SIEM, CloudWatch, or log aggregator
// NEVER log message content — PII risk
console.log(JSON.stringify(entry));
}
Error Handling
| Issue | Cause | Solution |
|---|---|---|
| Missing token counts | Streaming not aggregated | Sum tokens from stream chunks |
| Cost drift from bill | Pricing table outdated | Update PRICING map when rates change |
| Alert storm on 429s | Rate limit burst | Tune alert threshold, add request queue |
| High cardinality | Per-request labels | Never label by request ID or user ID |
Resources
Output
- Instrumented client wrapper with timing and cost tracking
- Prometheus metrics (requests, duration, tokens, cost, errors)
- Alerting rules for error rate, latency, rate limits, cost, auth
- Grafana dashboard panel specifications
- Structured logging format for SIEM integration
When not to use it
- →Logging message content containing PII
- →Labeling metrics by high-cardinality data like request IDs
Prerequisites
Limitations
- →Streaming responses require manual aggregation for accurate token counts
- →Pricing table in the wrapper must be updated manually when rates change
How it compares
Unlike manual logging, this approach provides pre-configured alerting rules and dashboard panels specifically mapped to Mistral API metrics.
Compared to similar skills
mistral-observability side by side with the closest alternatives in the catalog.
| Skill | Installs | Updated | Safety | Difficulty |
|---|---|---|---|---|
| mistral-observability (this skill) | 1 | 27d | No flags | Intermediate |
| analyzing-logs | 14 | 27d | Review | Beginner |
| obsidian-observability | 5 | 27d | Review | Intermediate |
| instruments-profiling | 3 | 2mo | No flags | Advanced |
Try saying
Example prompts that trigger this skill in your AI assistant.
More by jeremylongshore
View all by jeremylongshore →You might also like
analyzing-logs
jeremylongshore
Analyze application logs to detect performance issues, identify error patterns, and improve stability by extracting key insights.
obsidian-observability
jeremylongshore
Set up comprehensive logging and monitoring for Obsidian plugins. Use when implementing debug logging, tracking plugin performance, or setting up error reporting for your Obsidian plugin. Trigger with phrases like "obsidian logging", "obsidian monitoring", "obsidian debug", "track obsidian plugin".
instruments-profiling
steipete
Use when profiling native macOS or iOS apps with Instruments/xctrace. Covers correct binary selection, CLI arguments, exports, and common gotchas.
worker-benchmarks
ruvnet
Run comprehensive worker system benchmarks and performance analysis
sentry-rate-limits
jeremylongshore
Manage Sentry rate limits and quota optimization. Use when hitting rate limits, optimizing event volume, or managing Sentry costs. Trigger with phrases like "sentry rate limit", "sentry quota", "reduce sentry events", "sentry 429".
agent-benchmark-suite
ruvnet
Agent skill for benchmark-suite - invoke with $agent-benchmark-suite