langfuse-cost-tuning
Tracks and analyzes LLM token usage and costs via Langfuse dashboards and metrics APIs.
Install
mkdir -p .claude/skills/langfuse-cost-tuning && curl -L -o skill.zip "https://agentskills.codes/api/skills/download/6585" && unzip -o skill.zip -d .claude/skills/langfuse-cost-tuning && rm skill.zipInstalls to .claude/skills/langfuse-cost-tuning
Activation
This is the description your AI agent reads to decide when to run this skill — the better it matches your request, the more reliably it fires.
Monitor and optimize LLM costs using Langfuse analytics and dashboards.Key capabilities
- →Track LLM costs for supported models
- →Analyze token usage patterns
- →Implement budget-aware model routing
- →Set up automated budget alerts
- →Query costs programmatically via Metrics API
- →Configure pricing for custom models
How it works
Langfuse automatically calculates costs for supported models when token usage is captured in generation and embedding observations. For custom models, pricing can be configured in the Langfuse UI.
Inputs & outputs
When to use langfuse-cost-tuning
- →Track total LLM API spending
- →Identify high-cost prompts or models
- →Set up automated budget alerts
- →Analyze token usage patterns
About this skill
Langfuse Cost Tuning
Overview
Track, analyze, and optimize LLM costs using Langfuse's built-in token/cost tracking, the Metrics API for programmatic cost analysis, model routing for cost reduction, and automated budget alerts.
Prerequisites
- Langfuse tracing with token usage captured (via
observeOpenAIor manualusagefields) - For Metrics API:
@langfuse/clientinstalled - Understanding of LLM pricing models
How Langfuse Tracks Costs
Langfuse automatically calculates costs for supported models (OpenAI, Anthropic, Google) when token usage is captured. For custom models, you can configure pricing in the Langfuse UI under Settings > Model Definitions.
Cost tracking works on observations of type generation and embedding. The observeOpenAI wrapper captures usage automatically; for manual tracing, include usage in your observation updates.
Instructions
Step 1: Ensure Token Usage is Captured
// Automatic: observeOpenAI captures everything
import { observeOpenAI } from "@langfuse/openai";
const openai = observeOpenAI(new OpenAI());
// Tokens, model, latency, and cost are all auto-tracked
// Manual: include usage in generation observations
import { startActiveObservation, updateActiveObservation } from "@langfuse/tracing";
await startActiveObservation(
{ name: "llm-call", asType: "generation" },
async () => {
updateActiveObservation({ model: "gpt-4o" }); // Model required for cost calc
const response = await openai.chat.completions.create({
model: "gpt-4o",
messages: [{ role: "user", content: prompt }],
});
updateActiveObservation({
output: response.choices[0].message.content,
usage: {
promptTokens: response.usage?.prompt_tokens,
completionTokens: response.usage?.completion_tokens,
totalTokens: response.usage?.total_tokens,
},
// Optional: override inferred cost (in USD)
// costInUsd: 0.0015,
});
}
);
Step 2: Query Costs via Metrics API
import { LangfuseClient } from "@langfuse/client";
const langfuse = new LangfuseClient();
// Fetch aggregated cost metrics
async function getCostReport(days: number) {
const fromTimestamp = new Date(Date.now() - days * 86400000).toISOString();
// Use the API to list traces with cost data
const traces = await langfuse.api.traces.list({
fromTimestamp,
limit: 1000,
orderBy: "timestamp",
});
const costByModel = new Map<string, { cost: number; tokens: number; count: number }>();
for (const trace of traces.data) {
const observations = await langfuse.api.observations.list({
traceId: trace.id,
type: "GENERATION",
});
for (const obs of observations.data) {
const model = obs.model || "unknown";
const existing = costByModel.get(model) || { cost: 0, tokens: 0, count: 0 };
existing.cost += obs.calculatedTotalCost || 0;
existing.tokens += obs.totalTokens || 0;
existing.count += 1;
costByModel.set(model, existing);
}
}
console.log("\n=== LLM Cost Report ===");
console.log(`Period: Last ${days} days\n`);
let totalCost = 0;
for (const [model, data] of costByModel.entries()) {
console.log(`${model}:`);
console.log(` Calls: ${data.count}`);
console.log(` Tokens: ${data.tokens.toLocaleString()}`);
console.log(` Cost: $${data.cost.toFixed(4)}`);
totalCost += data.cost;
}
console.log(`\nTotal: $${totalCost.toFixed(4)}`);
}
getCostReport(7);
Step 3: Implement Smart Model Routing
Route requests to cheaper models when appropriate:
import { observe, updateActiveObservation } from "@langfuse/tracing";
interface ModelConfig {
model: string;
costPer1MInput: number;
costPer1MOutput: number;
maxComplexity: "simple" | "moderate" | "complex";
}
const MODELS: ModelConfig[] = [
{ model: "gpt-4o-mini", costPer1MInput: 0.15, costPer1MOutput: 0.60, maxComplexity: "simple" },
{ model: "gpt-4o", costPer1MInput: 2.50, costPer1MOutput: 10.00, maxComplexity: "moderate" },
{ model: "claude-sonnet-4-20250514", costPer1MInput: 3.00, costPer1MOutput: 15.00, maxComplexity: "complex" },
];
function selectModel(task: string, inputLength: number): ModelConfig {
const simpleTasks = ["classify", "extract", "summarize-short", "translate"];
const isSimple = simpleTasks.some((t) => task.includes(t));
const isShort = inputLength < 500;
if (isSimple && isShort) return MODELS[0]; // gpt-4o-mini
if (isSimple || inputLength < 2000) return MODELS[1]; // gpt-4o
return MODELS[2]; // claude-sonnet-4
}
const costOptimizedLLM = observe(
{ name: "cost-optimized-llm", asType: "generation" },
async (task: string, input: string) => {
const config = selectModel(task, input.length);
updateActiveObservation({
model: config.model,
metadata: {
task,
selectedReason: `${config.maxComplexity} tier`,
estimatedCostPer1M: config.costPer1MInput,
},
});
const response = await callModel(config.model, input);
updateActiveObservation({
output: response.content,
usage: response.usage,
});
return response;
}
);
Step 4: Budget Alerts
// scripts/cost-alert.ts -- run as cron job
import { LangfuseClient } from "@langfuse/client";
const langfuse = new LangfuseClient();
const ALERT_THRESHOLDS = {
dailyWarn: 50, // $50/day warning
dailyCritical: 200, // $200/day critical
perRequestWarn: 1, // $1/request warning
};
async function checkCostAlerts() {
const since = new Date(Date.now() - 86400000).toISOString(); // Last 24h
const traces = await langfuse.api.traces.list({
fromTimestamp: since,
limit: 500,
});
let dailyCost = 0;
let maxRequestCost = 0;
for (const trace of traces.data) {
const observations = await langfuse.api.observations.list({
traceId: trace.id,
type: "GENERATION",
});
const traceCost = observations.data.reduce(
(sum, obs) => sum + (obs.calculatedTotalCost || 0), 0
);
dailyCost += traceCost;
maxRequestCost = Math.max(maxRequestCost, traceCost);
}
console.log(`Daily cost: $${dailyCost.toFixed(2)}`);
console.log(`Max request cost: $${maxRequestCost.toFixed(4)}`);
if (dailyCost > ALERT_THRESHOLDS.dailyCritical) {
await sendAlert("CRITICAL", `Daily LLM cost: $${dailyCost.toFixed(2)}`);
} else if (dailyCost > ALERT_THRESHOLDS.dailyWarn) {
await sendAlert("WARNING", `Daily LLM cost: $${dailyCost.toFixed(2)}`);
}
}
checkCostAlerts();
Langfuse Dashboard Features
Langfuse provides built-in cost analytics in the UI:
- Cost Dashboard: Tracks token usage and costs over time by model, user, and session
- Latency Dashboard: Response times across models and user segments
- Custom Dashboards: Build custom views with multi-level aggregations
- Pricing Tiers: Supports complex pricing (cached tokens, audio tokens, per-model tiers)
Cost Optimization Strategies
| Strategy | Savings | Effort | How |
|---|---|---|---|
| Model downgrade | 50-95% | Low | Route simple tasks to gpt-4o-mini |
| Prompt optimization | 10-30% | Low | Remove filler words, use structured prompts |
| Response caching | 20-80% | Medium | Cache identical prompts with TTL |
| Batch processing | 50% | Medium | Use OpenAI Batch API for offline tasks |
| Token limits | 10-40% | Low | Set max_tokens on all calls |
Error Handling
| Issue | Cause | Solution |
|---|---|---|
| Missing cost data | No usage in generation | Ensure usage is included with promptTokens/completionTokens |
| Wrong cost calculation | Model name mismatch | Use exact model ID (e.g., gpt-4o-2024-08-06) |
| Custom model no cost | No pricing configured | Add model pricing in Langfuse Settings > Model Definitions |
| Stale pricing | Model prices changed | Update model definitions periodically |
Resources
When not to use it
- →When token usage is not captured in Langfuse tracing
- →When LLM pricing models are not understood
Prerequisites
Limitations
- →Cost calculation requires token usage to be captured
- →Model ID must be exact for correct cost calculation
- →Custom models require pricing configuration in Langfuse UI
How it compares
This skill provides programmatic access and automated controls for LLM cost management, unlike manual cost tracking.
Compared to similar skills
langfuse-cost-tuning side by side with the closest alternatives in the catalog.
| Skill | Installs | Updated | Safety | Difficulty |
|---|---|---|---|---|
| langfuse-cost-tuning (this skill) | 1 | 27d | No flags | Intermediate |
| databuddy | 1 | 3mo | Caution | Intermediate |
| analytics-pipeline | 1 | 6mo | No flags | Intermediate |
| logging-best-practices | 0 | 6mo | No flags | Intermediate |
Try saying
Example prompts that trigger this skill in your AI assistant.
More by jeremylongshore
View all by jeremylongshore →You might also like
databuddy
databuddy-analytics
Integrate Databuddy analytics into applications using the SDK or REST API. Use when implementing analytics tracking, feature flags, custom events, Web Vitals, error tracking, LLM observability, or querying analytics data programmatically.
analytics-pipeline
dadbodgeoff
Real-time analytics with Redis counters, periodic PostgreSQL flush, and time-series aggregation. High-performance event tracking without database bottlenecks.
logging-best-practices
neondatabase
Logging best practices focused on wide events (canonical log lines) for powerful debugging and analytics
firecrawl-cost-tuning
jeremylongshore
Optimize FireCrawl costs through tier selection, sampling, and usage monitoring. Use when analyzing FireCrawl billing, reducing API costs, or implementing usage monitoring and budget alerts. Trigger with phrases like "firecrawl cost", "firecrawl billing", "reduce firecrawl costs", "firecrawl pricing", "firecrawl expensive", "firecrawl budget".
segment-cdp
davila7
Expert patterns for Segment Customer Data Platform including Analytics.js, server-side tracking, tracking plans with Protocols, identity resolution, destinations configuration, and data governance best practices. Use when: segment, analytics.js, customer data platform, cdp, tracking plan.
groq-cost-tuning
jeremylongshore
Optimize Groq costs through tier selection, sampling, and usage monitoring. Use when analyzing Groq billing, reducing API costs, or implementing usage monitoring and budget alerts. Trigger with phrases like "groq cost", "groq billing", "reduce groq costs", "groq pricing", "groq expensive", "groq budget".