MI

mistral-observability

Provides instrumentation patterns for tracking Mistral AI costs, token usage, and latency.

Install

mkdir -p .claude/skills/mistral-observability && curl -L -o skill.zip "https://agentskills.codes/api/skills/download/4750" && unzip -o skill.zip -d .claude/skills/mistral-observability && rm skill.zip

Installs to .claude/skills/mistral-observability

Activation

This is the description your AI agent reads to decide when to run this skill — the better it matches your request, the more reliably it fires.

Set up comprehensive observability for Mistral AI with metrics, traces,
71 charsno explicit “when” trigger
Intermediate

Key capabilities

  • Instrument Mistral API client calls
  • Emit Prometheus metrics for requests and tokens
  • Configure alerting rules for latency and cost
  • Generate Grafana dashboard panels
  • Implement structured logging for observability

How it works

The skill provides an instrumented wrapper for the Mistral client that calculates duration, token usage, and costs, then pushes these metrics to a backend. It also includes YAML configurations for Prometheus alerting and Grafana dashboard definitions.

Inputs & outputs

You give it
Mistral API client call parameters
You get back
Prometheus metrics and structured log entries

When to use mistral-observability

  • Track token consumption
  • Monitor API latency
  • Calculate integration costs
  • Setup dashboard metrics

About this skill

Mistral AI Observability

Overview

Monitor Mistral AI API usage, latency, token consumption, error rates, and costs. Covers instrumented client wrapper, Prometheus metrics, Grafana dashboard panels, alerting rules, and structured logging.

Prerequisites

  • Mistral API integration in production
  • Prometheus or OpenTelemetry-compatible metrics backend
  • Alerting system (Alertmanager, PagerDuty, or similar)

Instructions

Step 1: Instrumented Client Wrapper

import { Mistral } from '@mistralai/mistralai';

const PRICING: Record<string, { input: number; output: number }> = {
  'mistral-small-latest':  { input: 0.10, output: 0.30 },
  'mistral-large-latest':  { input: 0.50, output: 1.50 },
  'codestral-latest':      { input: 0.30, output: 0.90 },
  'mistral-embed':         { input: 0.10, output: 0 },
};

interface MetricsEvent {
  model: string;
  endpoint: string;
  durationMs: number;
  status: 'success' | 'error';
  statusCode?: number;
  inputTokens?: number;
  outputTokens?: number;
  costUsd?: number;
}

function emitMetrics(event: MetricsEvent): void {
  // Push to your metrics backend (Prometheus, Datadog, etc.)
  console.log(JSON.stringify({ type: 'mistral_metric', ...event }));
}

async function instrumentedChat(
  client: Mistral,
  model: string,
  messages: any[],
  options?: any,
) {
  const start = performance.now();
  try {
    const response = await client.chat.complete({ model, messages, ...options });
    const duration = Math.round(performance.now() - start);
    const pricing = PRICING[model] ?? PRICING['mistral-small-latest'];
    const pt = response.usage?.promptTokens ?? 0;
    const ct = response.usage?.completionTokens ?? 0;

    emitMetrics({
      model,
      endpoint: 'chat.complete',
      durationMs: duration,
      status: 'success',
      inputTokens: pt,
      outputTokens: ct,
      costUsd: (pt / 1e6) * pricing.input + (ct / 1e6) * pricing.output,
    });

    return response;
  } catch (error: any) {
    emitMetrics({
      model,
      endpoint: 'chat.complete',
      durationMs: Math.round(performance.now() - start),
      status: 'error',
      statusCode: error.status,
    });
    throw error;
  }
}

Step 2: Prometheus Metrics

// Using prom-client
import { Counter, Histogram, Gauge } from 'prom-client';

const mistralRequests = new Counter({
  name: 'mistral_requests_total',
  help: 'Total Mistral API requests',
  labelNames: ['model', 'endpoint', 'status'],
});

const mistralDuration = new Histogram({
  name: 'mistral_request_duration_ms',
  help: 'Mistral request duration in milliseconds',
  labelNames: ['model', 'endpoint'],
  buckets: [100, 250, 500, 1000, 2500, 5000, 10000],
});

const mistralTokens = new Counter({
  name: 'mistral_tokens_total',
  help: 'Total tokens consumed',
  labelNames: ['model', 'direction'], // direction: input | output
});

const mistralCost = new Counter({
  name: 'mistral_cost_usd_total',
  help: 'Estimated cost in USD',
  labelNames: ['model'],
});

const mistralErrors = new Counter({
  name: 'mistral_errors_total',
  help: 'Total Mistral errors',
  labelNames: ['model', 'status_code'],
});

// Record metrics from instrumented wrapper
function recordPrometheusMetrics(event: MetricsEvent): void {
  mistralRequests.inc({ model: event.model, endpoint: event.endpoint, status: event.status });
  mistralDuration.observe({ model: event.model, endpoint: event.endpoint }, event.durationMs);

  if (event.status === 'success') {
    if (event.inputTokens) mistralTokens.inc({ model: event.model, direction: 'input' }, event.inputTokens);
    if (event.outputTokens) mistralTokens.inc({ model: event.model, direction: 'output' }, event.outputTokens);
    if (event.costUsd) mistralCost.inc({ model: event.model }, event.costUsd);
  } else {
    mistralErrors.inc({ model: event.model, status_code: String(event.statusCode ?? 'unknown') });
  }
}

Step 3: Alerting Rules

# prometheus/mistral-alerts.yaml
groups:
  - name: mistral
    rules:
      - alert: MistralHighErrorRate
        expr: rate(mistral_errors_total[5m]) / rate(mistral_requests_total[5m]) > 0.05
        for: 5m
        labels: { severity: critical }
        annotations:
          summary: "Mistral error rate exceeds 5%"
          runbook: "See mistral-incident-runbook skill"

      - alert: MistralHighLatency
        expr: histogram_quantile(0.95, rate(mistral_request_duration_ms_bucket[5m])) > 5000
        for: 5m
        labels: { severity: warning }
        annotations:
          summary: "Mistral P95 latency exceeds 5 seconds"

      - alert: MistralRateLimited
        expr: rate(mistral_errors_total{status_code="429"}[5m]) > 0
        for: 2m
        labels: { severity: warning }
        annotations:
          summary: "Mistral rate limiting detected"

      - alert: MistralCostSpike
        expr: increase(mistral_cost_usd_total[1h]) > 10
        labels: { severity: warning }
        annotations:
          summary: "Mistral spend exceeds $10/hour"

      - alert: MistralAuthFailure
        expr: increase(mistral_errors_total{status_code="401"}[5m]) > 0
        labels: { severity: critical }
        annotations:
          summary: "Mistral authentication failing — API key may be revoked"

Step 4: Grafana Dashboard Panels

Key panels to create:

PanelQueryType
Request Raterate(mistral_requests_total[5m])Time series
P50/P95/P99 Latencyhistogram_quantile(0.95, rate(..._bucket[5m]))Time series
Token Velocityrate(mistral_tokens_total{direction="output"}[5m])Time series
Hourly Costincrease(mistral_cost_usd_total[1h])Stat
Error Raterate(mistral_errors_total[5m]) by status_codeTime series
Model Distributionsum by (model) (rate(mistral_requests_total[5m]))Pie chart

Step 5: Structured Log Format

interface MistralLogEntry {
  ts: string;
  level: 'info' | 'warn' | 'error';
  model: string;
  endpoint: string;
  durationMs: number;
  inputTokens?: number;
  outputTokens?: number;
  costUsd?: number;
  status: string;
  statusCode?: number;
  requestId?: string;
}

function logMistralRequest(entry: MistralLogEntry): void {
  // Ship to SIEM, CloudWatch, or log aggregator
  // NEVER log message content — PII risk
  console.log(JSON.stringify(entry));
}

Error Handling

IssueCauseSolution
Missing token countsStreaming not aggregatedSum tokens from stream chunks
Cost drift from billPricing table outdatedUpdate PRICING map when rates change
Alert storm on 429sRate limit burstTune alert threshold, add request queue
High cardinalityPer-request labelsNever label by request ID or user ID

Resources

Output

  • Instrumented client wrapper with timing and cost tracking
  • Prometheus metrics (requests, duration, tokens, cost, errors)
  • Alerting rules for error rate, latency, rate limits, cost, auth
  • Grafana dashboard panel specifications
  • Structured logging format for SIEM integration

When not to use it

  • Logging message content containing PII
  • Labeling metrics by high-cardinality data like request IDs

Prerequisites

Mistral API integration in productionPrometheus or OpenTelemetry-compatible metrics backendAlerting system like Alertmanager or PagerDuty

Limitations

  • Streaming responses require manual aggregation for accurate token counts
  • Pricing table in the wrapper must be updated manually when rates change

How it compares

Unlike manual logging, this approach provides pre-configured alerting rules and dashboard panels specifically mapped to Mistral API metrics.

Compared to similar skills

mistral-observability side by side with the closest alternatives in the catalog.

SkillInstallsUpdatedSafetyDifficulty
mistral-observability (this skill)127dNo flagsIntermediate
analyzing-logs1427dReviewBeginner
obsidian-observability527dReviewIntermediate
instruments-profiling32moNo flagsAdvanced

Try saying

Example prompts that trigger this skill in your AI assistant.

More by jeremylongshore

View all by jeremylongshore

analyzing-logs

jeremylongshore

Analyze application logs to detect performance issues, identify error patterns, and improve stability by extracting key insights.

14123

ollama-setup

jeremylongshore

Configure auto-configure Ollama when user needs local LLM deployment, free AI alternatives, or wants to eliminate hosted API costs. Trigger phrases: "install ollama", "local AI", "free LLM", "self-hosted AI", "replace OpenAI", "no API costs". Use when appropriate context detected. Trigger with relevant phrases based on skill purpose.

1167

backtesting-trading-strategies

jeremylongshore

Backtest crypto and traditional trading strategies against historical data. Calculates performance metrics (Sharpe, Sortino, max drawdown), generates equity curves, and optimizes strategy parameters. Use when user wants to test a trading strategy, validate signals, or compare approaches. Trigger with phrases like "backtest strategy", "test trading strategy", "historical performance", "simulate trades", "optimize parameters", or "validate signals".

1071

generating-database-seed-data

jeremylongshore

Process this skill enables AI assistant to generate realistic test data and database seed scripts for development and testing environments. it uses faker libraries to create realistic data, maintains relational integrity, and allows configurable data volumes. u... Use when working with databases or data models. Trigger with phrases like 'database', 'query', or 'schema'.

1033

cursor-codebase-indexing

jeremylongshore

Execute set up and optimize Cursor codebase indexing. Triggers on "cursor index setup", "codebase indexing", "index codebase", "cursor semantic search". Use when working with cursor codebase indexing functionality. Trigger with phrases like "cursor codebase indexing", "cursor indexing", "cursor".

885

testing-mobile-apps

jeremylongshore

Execute mobile app testing on iOS and Android devices/simulators. Use when performing specialized testing. Trigger with phrases like "test mobile app", "run iOS tests", or "validate Android functionality".

810

Search skills

Search the agent skills registry