openai-vision
Extracts information from images or image sequences by detecting objects and describing visual content.
Install
mkdir -p .claude/skills/openai-vision && curl -L -o skill.zip "https://agentskills.codes/api/skills/download/5088" && unzip -o skill.zip -d .claude/skills/openai-vision && rm skill.zipInstalls to .claude/skills/openai-vision
Activation
This is the description your AI agent reads to decide when to run this skill — the better it matches your request, the more reliably it fires.
Analyze images and multi-frame sequences using OpenAI GPT vision modelsKey capabilities
- →Analyze image content and scenes
- →Extract text from images
- →Compare multiple images
- →Process video frames for temporal analysis
- →Generate structured JSON analysis
How it works
It uses OpenAI vision models to process visual inputs, extracting descriptive data, objects, and text based on provided prompts.
Inputs & outputs
When to use openai-vision
- →Describing UI mockups
- →Analyzing image datasets
- →Extracting text from images
About this skill
OpenAI Vision Analysis Skill
Purpose
This skill enables image analysis, scene understanding, text extraction, and multi-frame comparison using OpenAI's vision-capable GPT models (e.g., gpt-4o, gpt-4o-mini). It supports single images, multiple images for comparison, and sequential frames for temporal analysis.
When to Use
- Analyzing image content (objects, scenes, colors, spatial relationships)
- Extracting and reading text from images (OCR via vision models)
- Comparing multiple images to detect differences or changes
- Processing video frames to understand temporal progression
- Generating detailed image descriptions or captions
- Answering questions about visual content
Required Libraries
The following Python libraries are required:
from openai import OpenAI
import base64
import json
import os
from pathlib import Path
Input Requirements
- File formats: JPG, JPEG, PNG, WEBP, non-animated GIF
- Image sources: URL, Base64-encoded data, or local file paths
- Size limits: Up to 20MB per image; total request payload under 50MB
- Maximum images: Up to 500 images per request
- Image quality: Clear, legible content; avoid watermarks or heavy distortions
Output Schema
Analysis results should be returned as valid JSON conforming to this schema:
{
"success": true,
"images_analyzed": 1,
"analysis": {
"description": "A detailed scene description...",
"objects": [
{"name": "car", "color": "red", "position": "foreground center"},
{"name": "tree", "count": 3, "position": "background"}
],
"text_content": "Any text visible in the image...",
"colors": ["blue", "green", "white"],
"scene_type": "outdoor/urban"
},
"comparison": {
"differences": ["Object X appeared", "Color changed from A to B"],
"similarities": ["Background unchanged", "Layout consistent"]
},
"metadata": {
"model_used": "gpt-4o",
"detail_level": "high",
"token_usage": {"prompt": 1500, "completion": 200}
},
"warnings": []
}
Field Descriptions
success: Boolean indicating whether analysis completedimages_analyzed: Number of images processed in the requestanalysis.description: Natural language description of the image contentanalysis.objects: Array of detected objects with attributesanalysis.text_content: Any text extracted from the imageanalysis.colors: Dominant colors identifiedanalysis.scene_type: Classification of the scenecomparison: Present when multiple images are analyzed; describes differences and similaritiesmetadata.model_used: The GPT model used for analysismetadata.detail_level: Resolution level used (low,high, orauto)metadata.token_usage: Token consumption for cost trackingwarnings: Array of any issues or limitations encountered
Code Examples
Basic Image Analysis from URL
from openai import OpenAI
client = OpenAI()
def analyze_image_url(image_url, prompt="Describe this image in detail."):
"""Analyze an image from a URL using GPT-4o vision."""
response = client.chat.completions.create(
model="gpt-4o",
messages=[
{
"role": "user",
"content": [
{"type": "text", "text": prompt},
{
"type": "image_url",
"image_url": {
"url": image_url,
"detail": "high"
}
}
]
}
],
max_tokens=1000
)
return response.choices[0].message.content
Image Analysis from Local File (Base64)
from openai import OpenAI
import base64
client = OpenAI()
def encode_image_to_base64(image_path):
"""Encode a local image file to base64."""
with open(image_path, "rb") as image_file:
return base64.standard_b64encode(image_file.read()).decode("utf-8")
def get_image_media_type(image_path):
"""Determine the media type based on file extension."""
ext = image_path.lower().split('.')[-1]
media_types = {
'jpg': 'image/jpeg',
'jpeg': 'image/jpeg',
'png': 'image/png',
'gif': 'image/gif',
'webp': 'image/webp'
}
return media_types.get(ext, 'image/jpeg')
def analyze_local_image(image_path, prompt="Describe this image in detail."):
"""Analyze a local image file using GPT-4o vision."""
base64_image = encode_image_to_base64(image_path)
media_type = get_image_media_type(image_path)
response = client.chat.completions.create(
model="gpt-4o",
messages=[
{
"role": "user",
"content": [
{"type": "text", "text": prompt},
{
"type": "image_url",
"image_url": {
"url": f"data:{media_type};base64,{base64_image}",
"detail": "high"
}
}
]
}
],
max_tokens=1000
)
return response.choices[0].message.content
Multi-Image Comparison
from openai import OpenAI
import base64
client = OpenAI()
def compare_images(image_paths, comparison_prompt=None):
"""Compare multiple images and identify differences."""
if comparison_prompt is None:
comparison_prompt = (
"Compare these images carefully. "
"List all differences and similarities you observe. "
"Describe any changes in objects, colors, positions, or text."
)
content = [{"type": "text", "text": comparison_prompt}]
for i, image_path in enumerate(image_paths):
base64_image = encode_image_to_base64(image_path)
media_type = get_image_media_type(image_path)
# Add label for each image
content.append({
"type": "text",
"text": f"Image {i + 1}:"
})
content.append({
"type": "image_url",
"image_url": {
"url": f"data:{media_type};base64,{base64_image}",
"detail": "high"
}
})
response = client.chat.completions.create(
model="gpt-4o",
messages=[{"role": "user", "content": content}],
max_tokens=2000
)
return response.choices[0].message.content
Multi-Frame Video Analysis
from openai import OpenAI
import base64
from pathlib import Path
client = OpenAI()
def analyze_video_frames(frame_paths, analysis_prompt=None):
"""Analyze a sequence of video frames for temporal understanding."""
if analysis_prompt is None:
analysis_prompt = (
"These are sequential frames from a video. "
"Describe what is happening over time. "
"Identify any motion, changes, or events that occur across the frames."
)
content = [{"type": "text", "text": analysis_prompt}]
for i, frame_path in enumerate(frame_paths):
base64_image = encode_image_to_base64(frame_path)
media_type = get_image_media_type(frame_path)
content.append({
"type": "text",
"text": f"Frame {i + 1}:"
})
content.append({
"type": "image_url",
"image_url": {
"url": f"data:{media_type};base64,{base64_image}",
"detail": "auto" # Use auto for frames to balance cost
}
})
response = client.chat.completions.create(
model="gpt-4o",
messages=[{"role": "user", "content": content}],
max_tokens=2000
)
return response.choices[0].message.content
Full Analysis with JSON Output
from openai import OpenAI
import base64
import json
import os
client = OpenAI()
def analyze_image_to_json(image_path, extract_text=True):
"""Perform comprehensive image analysis and return structured JSON."""
filename = os.path.basename(image_path)
prompt = """Analyze this image and return a JSON object with the following structure:
{
"description": "detailed scene description",
"objects": [{"name": "object name", "attributes": "color, size, position"}],
"text_content": "any visible text or null if none",
"colors": ["dominant", "colors"],
"scene_type": "indoor/outdoor/abstract/etc",
"people_count": 0,
"notable_features": ["list of notable visual elements"]
}
Return ONLY valid JSON, no other text."""
try:
base64_image = encode_image_to_base64(image_path)
media_type = get_image_media_type(image_path)
response = client.chat.completions.create(
model="gpt-4o",
messages=[
{
"role": "user",
"content": [
{"type": "text", "text": prompt},
{
"type": "image_url",
"image_url": {
"url": f"data:{media_type};base64,{base64_image}",
"detail": "high"
}
}
]
}
],
max_tokens=1500
)
# Parse the response as JSON
analysis_text = response.choices[0].message.content
# Remove markdown code blocks if present
if analysis_text.startswith("```"):
analysis_text = analysis_text.split("```")[1]
if analysis_text.startswith("json"):
analysis_text = analysis_text[4:]
analysis = json.loads(analysis_text.strip())
result = {
"success": True,
"filename": filename,
"analysis": analysis,
"metadata": {
---
*Content truncated.*
When not to use it
- →For medical diagnostic analysis
Limitations
- →Cannot access EXIF or GPS metadata
- →Limited accuracy for small or rotated text
- →Not for medical diagnostics
How it compares
It provides structured, multimodal data extraction and comparison rather than simple image captioning.
Compared to similar skills
openai-vision side by side with the closest alternatives in the catalog.
| Skill | Installs | Updated | Safety | Difficulty |
|---|---|---|---|---|
| openai-vision (this skill) | 1 | 6mo | No flags | Intermediate |
| quant-analyst | 103 | 2mo | No flags | Advanced |
| umap-learn | 6 | 2mo | Review | Intermediate |
| embedding-strategies | 8 | 2mo | No flags | Intermediate |
Try saying
Example prompts that trigger this skill in your AI assistant.
More by benchflow-ai
View all by benchflow-ai →You might also like
quant-analyst
zenobi-us
Expert quantitative analyst specializing in financial modeling, algorithmic trading, and risk analytics. Masters statistical methods, derivatives pricing, and high-frequency trading with focus on mathematical rigor, performance optimization, and profitable strategy development.
umap-learn
K-Dense-AI
UMAP dimensionality reduction. Fast nonlinear manifold learning for 2D/3D visualization, clustering preprocessing (HDBSCAN), supervised/parametric UMAP, for high-dimensional data.
embedding-strategies
wshobson
Select and optimize embedding models for semantic search and RAG applications. Use when choosing embedding models, implementing chunking strategies, or optimizing embedding quality for specific domains.
building-automl-pipelines
jeremylongshore
Build automated machine learning pipelines, including feature engineering, model selection, and performance evaluation.
model-compare
rawwerks
Compare 3D CAD models using boolean operations (IoU, Dice, precision/recall). Use when evaluating generated models against gold references, diffing CAD revisions, or computing similarity metrics for ML training. Triggers on: model diff, compare models, IoU, intersection over union, model similarity, CAD comparison, STEP diff, 3D evaluation, gold reference, generated model, precision recall 3D.
matchms
davila7
Mass spectrometry analysis. Process mzML/MGF/MSP, spectral similarity (cosine, modified cosine), metadata harmonization, compound ID, for metabolomics and MS data processing.