AN

anthropic-api

Provides SDK patterns, model selection, and cost optimization strategies for the Claude API.

Install

mkdir -p .claude/skills/anthropic-api && curl -L -o skill.zip "https://agentskills.codes/api/skills/download/14358" && unzip -o skill.zip -d .claude/skills/anthropic-api && rm skill.zip

Installs to .claude/skills/anthropic-api

Activation

This is the description your AI agent reads to decide when to run this skill — the better it matches your request, the more reliably it fires.

Expert guidance for building applications with Anthropic''s Claude API.
71 charsno explicit “when” trigger
Intermediate

Key capabilities

  • Select appropriate Claude model based on use case and cost
  • Set up Python SDK for Anthropic API
  • Set up TypeScript SDK for Anthropic API
  • Implement prompt caching for cost optimization
  • Process large volumes of requests using the Batch API
  • Count tokens before sending requests for cost estimation

How it works

This skill guides users through API selection, SDK implementation, and cost optimization techniques like prompt caching and batch processing for Anthropic's Claude API.

Inputs & outputs

You give it
user prompt and desired Claude model interaction
You get back
Claude API responses, optimized for cost and performance

When to use anthropic-api

  • Implement Claude API streaming
  • Optimize prompt caching strategy
  • Choose the right Claude model

About this skill

Anthropic Claude API Expert Guide

Build production-grade applications with Anthropic's Claude API using best practices, cost optimization strategies, and proven patterns.

Model Selection Guide

ModelModel IDBest ForInput/Output Cost
Claude Opus 4.5claude-opus-4-5-20250514Most capable, complex reasoning$5 / $25 per MTok
Claude Sonnet 4.5claude-sonnet-4-5-20250514Balanced performance/cost$3 / $15 per MTok
Claude Haiku 4.5claude-haiku-4-5-20250514Fast, high-volume, cost-efficient$1 / $5 per MTok

Decision Framework:

  • Use Haiku for: Classification, extraction, simple Q&A, high-volume workloads
  • Use Sonnet for: Code generation, analysis, most production use cases (90% of Opus quality at 20% cost)
  • Use Opus for: Complex reasoning, research, when quality is paramount

Setup

Python SDK

pip install anthropic
import os
from anthropic import Anthropic

client = Anthropic(
    api_key=os.environ["ANTHROPIC_API_KEY"]  # Default, can be omitted
)

message = client.messages.create(
    model="claude-sonnet-4-5-20250514",
    max_tokens=1024,
    messages=[
        {"role": "user", "content": "Hello, Claude"}
    ]
)
print(message.content[0].text)

TypeScript SDK

npm install @anthropic-ai/sdk
import Anthropic from "@anthropic-ai/sdk";

const client = new Anthropic({
  apiKey: process.env.ANTHROPIC_API_KEY,
});

const message = await client.messages.create({
  model: "claude-sonnet-4-5-20250514",
  max_tokens: 1024,
  messages: [{ role: "user", content: "Hello, Claude" }],
});
console.log(message.content[0].text);

Cost Optimization Strategies

1. Prompt Caching (Up to 90% Savings)

Cache stable content like system prompts, documents, or tool definitions. Cache hits cost only 10% of base input price.

from anthropic import Anthropic

client = Anthropic()

# First request - writes to cache (1.25x input cost)
response = client.messages.create(
    model="claude-sonnet-4-5-20250514",
    max_tokens=1024,
    system=[
        {
            "type": "text",
            "text": "You are an expert legal assistant...",  # Large system prompt
            "cache_control": {"type": "ephemeral"}  # 5-minute cache
        }
    ],
    messages=[{"role": "user", "content": "Analyze this contract..."}]
)

# Subsequent requests within 5 minutes - cache hit (0.1x input cost)
response2 = client.messages.create(
    model="claude-sonnet-4-5-20250514",
    max_tokens=1024,
    system=[
        {
            "type": "text",
            "text": "You are an expert legal assistant...",  # Same content
            "cache_control": {"type": "ephemeral"}
        }
    ],
    messages=[{"role": "user", "content": "Different question..."}]
)

# Check cache usage
print(f"Cache read: {response2.usage.cache_read_input_tokens}")
print(f"Cache write: {response2.usage.cache_creation_input_tokens}")

Cache Duration Options:

  • ephemeral - 5-minute cache (1.25x write, 0.1x read)
  • For 1-hour cache, use extended caching (2x write, 0.1x read)

Minimum Token Requirements:

  • Claude 3.5+ models: 1,024 tokens minimum per cache checkpoint
  • Content below minimum won't be cached

2. Batch API (50% Discount)

Process large volumes asynchronously with guaranteed 24-hour completion.

import asyncio
from anthropic import AsyncAnthropic

client = AsyncAnthropic()

async def process_batch():
    # Create batch
    batch = await client.messages.batches.create(
        requests=[
            {
                "custom_id": "request-1",
                "params": {
                    "model": "claude-sonnet-4-5-20250514",
                    "max_tokens": 1024,
                    "messages": [{"role": "user", "content": "Summarize document 1"}]
                }
            },
            {
                "custom_id": "request-2",
                "params": {
                    "model": "claude-sonnet-4-5-20250514",
                    "max_tokens": 1024,
                    "messages": [{"role": "user", "content": "Summarize document 2"}]
                }
            }
        ]
    )

    print(f"Batch ID: {batch.id}")

    # Poll for results (or use webhooks)
    while True:
        batch = await client.messages.batches.retrieve(batch.id)
        if batch.processing_status == "ended":
            break
        await asyncio.sleep(60)

    # Get results
    async for entry in await client.messages.batches.results(batch.id):
        if entry.result.type == "succeeded":
            print(f"{entry.custom_id}: {entry.result.message.content[0].text}")

When to Use Batch:

  • Background processing (reports, analysis)
  • Bulk content generation
  • Data enrichment pipelines
  • Any non-real-time workload

3. Token Optimization

# Count tokens before sending (estimate costs)
token_count = client.messages.count_tokens(
    model="claude-sonnet-4-5-20250514",
    messages=[{"role": "user", "content": "Your message here"}]
)
print(f"Input tokens: {token_count.input_tokens}")

# Set appropriate max_tokens (don't over-reserve)
response = client.messages.create(
    model="claude-sonnet-4-5-20250514",
    max_tokens=500,  # Set to expected output, not maximum
    messages=[{"role": "user", "content": "Brief summary of..."}]
)

# Check actual usage
print(f"Used: {response.usage.output_tokens} tokens")

Streaming

Basic Streaming

from anthropic import Anthropic

client = Anthropic()

# Simple streaming
with client.messages.stream(
    model="claude-sonnet-4-5-20250514",
    max_tokens=1024,
    messages=[{"role": "user", "content": "Tell me a story"}]
) as stream:
    for text in stream.text_stream:
        print(text, end="", flush=True)
    print()

    # Get final message after stream completes
    message = stream.get_final_message()
    print(f"Total tokens: {message.usage.output_tokens}")

Async Streaming

import asyncio
from anthropic import AsyncAnthropic

client = AsyncAnthropic()

async def stream_response():
    async with client.messages.stream(
        model="claude-sonnet-4-5-20250514",
        max_tokens=1024,
        messages=[{"role": "user", "content": "Explain quantum computing"}]
    ) as stream:
        async for text in stream.text_stream:
            print(text, end="", flush=True)
        print()

asyncio.run(stream_response())

Handling All Event Types

async with client.messages.stream(
    model="claude-sonnet-4-5-20250514",
    max_tokens=1024,
    messages=[{"role": "user", "content": "Use a tool"}]
) as stream:
    async for event in stream:
        if event.type == "text":
            print(event.text, end="")
        elif event.type == "input_json":
            # Tool input being streamed
            print(f"Tool input delta: {event.partial_json}")
        elif event.type == "content_block_stop":
            print(f"\nBlock complete: {event.content_block}")
        elif event.type == "message_stop":
            print(f"\nFinal message: {event.message}")

Tool Use

Defining and Using Tools

from anthropic import Anthropic

client = Anthropic()

# Define tools
tools = [
    {
        "name": "get_weather",
        "description": "Get current weather for a location",
        "input_schema": {
            "type": "object",
            "properties": {
                "location": {
                    "type": "string",
                    "description": "City and state, e.g. 'San Francisco, CA'"
                },
                "unit": {
                    "type": "string",
                    "enum": ["celsius", "fahrenheit"],
                    "description": "Temperature unit"
                }
            },
            "required": ["location"]
        }
    }
]

# Initial request
response = client.messages.create(
    model="claude-sonnet-4-5-20250514",
    max_tokens=1024,
    tools=tools,
    messages=[{"role": "user", "content": "What's the weather in Paris?"}]
)

# Check if tool use requested
if response.stop_reason == "tool_use":
    tool_use = next(block for block in response.content if block.type == "tool_use")

    # Execute tool (your implementation)
    tool_result = execute_weather_lookup(tool_use.input)

    # Continue conversation with tool result
    final_response = client.messages.create(
        model="claude-sonnet-4-5-20250514",
        max_tokens=1024,
        tools=tools,
        messages=[
            {"role": "user", "content": "What's the weather in Paris?"},
            {"role": "assistant", "content": response.content},
            {
                "role": "user",
                "content": [
                    {
                        "type": "tool_result",
                        "tool_use_id": tool_use.id,
                        "content": tool_result
                    }
                ]
            }
        ]
    )

Tool Choice Options

# Let Claude decide (default)
tool_choice={"type": "auto"}

# Force tool use
tool_choice={"type": "any"}

# Force specific tool
tool_choice={"type": "tool", "name": "get_weather"}

# Disable tools for this request
tool_choice={"type": "none"}

Error Handling

import anthropic
from anthropic import Anthropic

client = Anthropic()

try:
    response = client.messages.create(
        model="claude-sonnet-4-5-20250514",
        max_tokens=1024,
        messages=[{"role": "user", "content": "Hello"}]
    )
except anthropic.APIConnectionError as e:
    # Network issues
    print(f"Connection failed: {e.__cause__}")
except anthropic.RateLimitError as e:
    # 429 - implement exponential backoff
    print(f"Rate limited. Retry after backoff.")
except anthropic.BadRequestError as e:
    # 400 - check your request
    print(f"Bad request: {e.message}")
except anthropic.AuthenticationError as e:
    # 401 - check

---

*Content truncated.*

When not to use it

  • When using models other than Anthropic's Claude
  • When the task does not involve API interaction
  • When real-time processing is required for batch workloads

Prerequisites

ANTHROPIC_API_KEY

Limitations

  • The skill focuses on Anthropic's Claude API.
  • The skill does not cover all possible API integrations.
  • The skill does not manage API key generation.

How it compares

This skill provides specific strategies for optimizing cost and performance with Anthropic's Claude API, unlike general API usage which may not consider these factors.

Compared to similar skills

anthropic-api side by side with the closest alternatives in the catalog.

SkillInstallsUpdatedSafetyDifficulty
anthropic-api (this skill)06moReviewIntermediate
stripe-integration482moNo flagsAdvanced
langchain268moReviewIntermediate
langfuse76moNo flagsIntermediate

Try saying

Example prompts that trigger this skill in your AI assistant.

You might also like

stripe-integration

wshobson

Implement Stripe payment processing for robust, PCI-compliant payment flows including checkout, subscriptions, and webhooks. Use when integrating Stripe payments, building subscription systems, or implementing secure checkout flows.

48165

langchain

zechenzhangAGI

Framework for building LLM-powered applications with agents, chains, and RAG. Supports multiple providers (OpenAI, Anthropic, Google), 500+ integrations, ReAct agents, tool calling, memory management, and vector store retrieval. Use for building chatbots, question-answering systems, autonomous agents, or RAG applications. Best for rapid prototyping and production deployments.

26138

langfuse

davila7

Expert in Langfuse - the open-source LLM observability platform. Covers tracing, prompt management, evaluation, datasets, and integration with LangChain, LlamaIndex, and OpenAI. Essential for debugging, monitoring, and improving LLM applications in production. Use when: langfuse, llm observability, llm tracing, prompt management, llm evaluation.

743

openrouter-hello-world

jeremylongshore

Create your first OpenRouter API request with a simple example. Use when learning OpenRouter or testing your setup. Trigger with phrases like 'openrouter hello world', 'openrouter first request', 'openrouter quickstart', 'test openrouter'.

733

telegram-dev

2025Emma

Telegram 生态开发全栈指南 - 涵盖 Bot API、Mini Apps (Web Apps)、MTProto 客户端开发。包括消息处理、支付、内联模式、Webhook、认证、存储、传感器 API 等完整开发资源。

232

llm-application-dev

skillcreatorai

Building applications with Large Language Models - prompt engineering, RAG patterns, and LLM integration. Use for AI-powered features, chatbots, or LLM-based automation.

323

Search skills

Search the agent skills registry