bedrock
This tool integrates AWS Bedrock foundation models to perform model invocation, embedding generation, and RAG configuration.
Install
mkdir -p .claude/skills/bedrock && curl -L -o skill.zip "https://agentskills.codes/api/skills/download/3405" && unzip -o skill.zip -d .claude/skills/bedrock && rm skill.zipInstalls to .claude/skills/bedrock
Activation
This is the description your AI agent reads to decide when to run this skill — the better it matches your request, the more reliably it fires.
AWS Bedrock foundation models for generative AI. Use when invoking foundation models, building AI applications, creating embeddings, configuring model access, or implementing RAG patterns.Key capabilities
- →Invoke foundation models
- →Generate vector embeddings
- →Stream model responses
- →Manage multi-turn conversations
- →List and query available models
How it works
The tool interfaces with the AWS Bedrock API to send prompts to foundation models and receive generated content or embeddings, supporting both synchronous and streaming modes.
Inputs & outputs
When to use bedrock
- →Invoking text models
- →Generating vector embeddings
- →Building RAG applications
About this skill
AWS Bedrock
Amazon Bedrock provides access to foundation models (FMs) from AI companies through a unified API. Build generative AI applications with text generation, embeddings, and image generation capabilities.
Table of Contents
Core Concepts
Foundation Models
Pre-trained models available through Bedrock:
- Claude (Anthropic): Text generation, analysis, coding
- Nova / Titan (Amazon): Text, multimodal, embeddings
- GPT / gpt-oss (OpenAI): Text generation, reasoning
- Llama (Meta): Open-weight text generation
- Mistral: Efficient text generation
- Stable Image (Stability AI): Image generation and editing
Model Access
In commercial Regions, access to all serverless models is enabled by default (no console opt-in). In GovCloud (US), models are still enabled manually on the Model access page (third-party models also in the linked commercial account):
- First invocation of a third-party model auto-subscribes via AWS Marketplace (up to 15 min); caller needs
aws-marketplace:Subscribe,Unsubscribe,ViewSubscriptions - Anthropic models on
bedrock-runtimeneed a one-time use case form per account/org (put-use-case-for-model-access) - Invoking implies EULA acceptance; to block a model, deny both
bedrock:InvokeModelandbedrock:InvokeModelWithResponseStreamon it (SCP/IAM); streaming APIs such asConverseStreamuse the latter. Denyingaws-marketplace:Subscribealone does not block first use
Endpoints
| Endpoint | APIs | Use for |
|---|---|---|
bedrock-runtime.{region}.amazonaws.com (recommended) | InvokeModel, Converse, Anthropic Messages (/anthropic), OpenAI Responses/Chat Completions (/openai/v1) | Guardrails, cross-Region inference, prompt routing, application inference profiles |
bedrock-mantle.{region}.api.aws | OpenAI Responses/Chat Completions (/openai/v1), Anthropic Messages | Server-side tools (Web Search), background=true async, Projects/Workspaces, single-Region access to CRIS-only models |
- Same per-token price on both; auth via SigV4 or Bedrock API key (
AWS_BEARER_TOKEN_BEDROCK) - IAM:
bedrock:InvokeModel(runtime) vsbedrock-mantle:CreateInference(mantle) - Responses API on
bedrock-runtimeis synchronous only and has no server-side tools
Inference Profiles and Model Lifecycle
- Newer models (e.g. Claude Sonnet 5) have no in-Region on-demand ID on
bedrock-runtime: use a geo (us.,eu.,au.) orglobal.inference profile ID asmodelId - Lifecycle is
Active->Legacy->EOL(seemodelLifecycleinget-foundation-model). Legacy: no new Provisioned Throughput, fine-tuning, or quota increases; EOL: requests fail - Model cards list an "EOL no sooner than" date; check before pinning a model ID
Knowledge Bases and Agents
- Managed knowledge bases (
type: MANAGED): Bedrock runs storage, indexing, and retrieval. Only type that supportsAgenticRetrieveStream(query decomposition, iterative retrieval, optional AgentCore Memory viamemoryConfiguration) - Native multimodal managed KBs embed video/audio/image directly with TwelveLabs Marengo Embed 3.0 (
twelvelabs.marengo-embed-3-0-v1:0); query with text viaRetrieveonly (noRetrieveAndGenerate) - Bedrock Agents Classic is in maintenance mode: closed to new accounts since July 30, 2026 (
CreateAgent/InvokeInlineAgentreturn 403 without prior 12-month usage), model catalog frozen. Build new agents on Amazon Bedrock AgentCore
Inference Types
| Type | Use Case | Pricing |
|---|---|---|
| On-Demand | Variable workloads | Per token |
| Provisioned Throughput | Consistent high-volume | Hourly commitment |
| Batch Inference | Async large-scale | Discounted per token |
Common Patterns
Invoke Model (Text Generation)
AWS CLI:
# Invoke Claude
aws bedrock-runtime invoke-model \
--model-id us.anthropic.claude-sonnet-5 \
--content-type application/json \
--accept application/json \
--cli-binary-format raw-in-base64-out \
--body '{
"anthropic_version": "bedrock-2023-05-31",
"max_tokens": 4096,
"messages": [
{"role": "user", "content": "Explain AWS Lambda in 3 sentences."}
]
}' \
response.json
# Claude Sonnet 5/Opus 5 think by default: content may start with a thinking block
cat response.json | jq -r '.content[] | select(.type=="text") | .text'
boto3:
import boto3
import json
bedrock = boto3.client('bedrock-runtime')
def invoke_claude(prompt, max_tokens=4096):
response = bedrock.invoke_model(
modelId='us.anthropic.claude-sonnet-5',
contentType='application/json',
accept='application/json',
body=json.dumps({
'anthropic_version': 'bedrock-2023-05-31',
'max_tokens': max_tokens,
'messages': [
{'role': 'user', 'content': prompt}
]
})
)
result = json.loads(response['body'].read())
# Skip thinking blocks (adaptive thinking is on by default for Sonnet 5).
# max_tokens caps thinking + text, so a truncated response may have no text block.
if result['stop_reason'] == 'max_tokens':
print('Truncated at max_tokens: raise it or lower output_config.effort')
return next((b['text'] for b in result['content'] if b['type'] == 'text'), '')
# Usage
response = invoke_claude('What is Amazon S3?')
print(response)
Streaming Response
import boto3
import json
bedrock = boto3.client('bedrock-runtime')
def stream_claude(prompt):
response = bedrock.invoke_model_with_response_stream(
modelId='us.anthropic.claude-sonnet-5',
contentType='application/json',
accept='application/json',
body=json.dumps({
'anthropic_version': 'bedrock-2023-05-31',
'max_tokens': 4096,
'messages': [
{'role': 'user', 'content': prompt}
]
})
)
for event in response['body']:
chunk = json.loads(event['chunk']['bytes'])
if chunk['type'] == 'content_block_delta':
yield chunk['delta'].get('text', '')
# Usage
for text in stream_claude('Write a haiku about cloud computing.'):
print(text, end='', flush=True)
Generate Embeddings
import boto3
import json
bedrock = boto3.client('bedrock-runtime')
def get_embedding(text):
response = bedrock.invoke_model(
modelId='amazon.titan-embed-text-v2:0',
contentType='application/json',
accept='application/json',
body=json.dumps({
'inputText': text,
'dimensions': 1024,
'normalize': True
})
)
result = json.loads(response['body'].read())
return result['embedding']
# Usage
embedding = get_embedding('AWS Lambda is a serverless compute service.')
print(f'Embedding dimension: {len(embedding)}')
Conversation with History
import boto3
import json
bedrock = boto3.client('bedrock-runtime')
class Conversation:
def __init__(self, system_prompt=None):
self.messages = []
self.system = system_prompt
def chat(self, user_message):
self.messages.append({
'role': 'user',
'content': user_message
})
body = {
'anthropic_version': 'bedrock-2023-05-31',
'max_tokens': 4096,
'messages': self.messages
}
if self.system:
body['system'] = self.system
response = bedrock.invoke_model(
modelId='us.anthropic.claude-sonnet-5',
contentType='application/json',
accept='application/json',
body=json.dumps(body)
)
result = json.loads(response['body'].read())
if result['stop_reason'] == 'max_tokens':
# max_tokens caps thinking + text; don't store a truncated/empty turn
self.messages.pop()
raise RuntimeError('Truncated at max_tokens: raise it or lower output_config.effort')
assistant_message = next(
(b['text'] for b in result['content'] if b['type'] == 'text'), ''
)
self.messages.append({
'role': 'assistant',
'content': assistant_message
})
return assistant_message
# Usage
conv = Conversation(system_prompt='You are an AWS solutions architect.')
print(conv.chat('What database should I use for a chat application?'))
print(conv.chat('What about for time-series data?'))
List Available Models
# List all foundation models
aws bedrock list-foundation-models \
--query 'modelSummaries[*].[modelId,modelName,providerName]' \
--output table
# Filter by provider
aws bedrock list-foundation-models \
--by-provider anthropic \
--query 'modelSummaries[*].modelId'
# Get model details (includes modelLifecycle.status)
aws bedrock get-foundation-model \
--model-identifier anthropic.claude-sonnet-5
Check Model Access
# agreementAvailability.status AVAILABLE / NOT_AVAILABLE, authorizationStatus
aws bedrock get-foundation-model-availability \
--model-id anthropic.claude-sonnet-5
# Anthropic one-time use case form (base64-encoded JSON:
# companyName, companyWebsite, intendedUsers, industryOption, otherIndustryOption, useCases)
aws bedrock put-use-case-for-model-access --form-data <base64-json>
# Programmatic agreement for third-party models
aws bedrock list-foundation-model-agreement-offers --model-id <model-id>
aws bedrock create-foundation-model-agreement --model-id <model-id> --offer-token <token>
Count Tokens
# Free; returns inputTokens. Not supported for every model (e.g. CRIS-only Claude models)
aws bedrock-runtime count-tokens \
--model-id anthropic
---
*Content truncated.*
When not to use it
- →When local model execution is required
- →When AWS region-specific model access is not configured
Prerequisites
Limitations
- →Model access is region-specific
- →Subject to AWS service quotas and throttling
- →Requires explicit IAM configuration
How it compares
It provides a unified interface for multiple foundation models via AWS infrastructure, rather than managing individual model provider APIs.
Compared to similar skills
bedrock side by side with the closest alternatives in the catalog.
| Skill | Installs | Updated | Safety | Difficulty |
|---|---|---|---|---|
| bedrock (this skill) | 1 | 8mo | Review | Intermediate |
| langchain | 26 | 10mo | Review | Intermediate |
| reasoningbank-with-agentdb | 5 | 11mo | Review | Intermediate |
| agent-memory-systems | 5 | 8mo | No flags | Advanced |
Try saying
Example prompts that trigger this skill in your AI assistant.
More by itsmostafa
View all by itsmostafa →You might also like
langchain
zechenzhangAGI
Framework for building LLM-powered applications with agents, chains, and RAG. Supports multiple providers (OpenAI, Anthropic, Google), 500+ integrations, ReAct agents, tool calling, memory management, and vector store retrieval. Use for building chatbots, question-answering systems, autonomous agents, or RAG applications. Best for rapid prototyping and production deployments.
reasoningbank-with-agentdb
ruvnet
Implement ReasoningBank adaptive learning with AgentDB's 150x faster vector database. Includes trajectory tracking, verdict judgment, memory distillation, and pattern recognition. Use when building self-learning agents, optimizing decision-making, or implementing experience replay systems.
agent-memory-systems
davila7
Memory is the cornerstone of intelligent agents. Without it, every interaction starts from zero. This skill covers the architecture of agent memory: short-term (context window), long-term (vector stores), and the cognitive architectures that organize them. Key insight: Memory isn't just storage - it's retrieval. A million stored facts mean nothing if you can't find the right one. Chunking, embedding, and retrieval strategies determine whether your agent remembers or forgets. The field is fragm
senior-ml-engineer
davila7
World-class ML engineering skill for productionizing ML models, MLOps, and building scalable ML systems. Expertise in PyTorch, TensorFlow, model deployment, feature stores, model monitoring, and ML infrastructure. Includes LLM integration, fine-tuning, RAG systems, and agentic AI. Use when deploying ML models, building ML platforms, implementing MLOps, or integrating LLMs into production systems.
dspy
davila7
Build complex AI systems with declarative programming, optimize prompts automatically, create modular RAG systems and agents with DSPy - Stanford NLP's framework for systematic LM programming
ai-engineer
sickn33
Build production-ready LLM applications, advanced RAG systems, and intelligent agents. Implements vector search, multimodal AI, agent orchestration, and enterprise AI integrations. Use PROACTIVELY for LLM features, chatbots, AI agents, or AI-powered applications.