Tags
Best Image Generation Skills for AI Agents
63 Image Generation skills for AI coding assistants — ranked by popularity.
This collection provides specific AI-agent skills designed to integrate image generation and editing capabilities directly into your development workflow. Whether you use Claude Code, Codex, or Cursor, these skills allow you to automate visual tasks like generating scientific diagrams, building complex ComfyUI workflows, or editing images via the Gemini API. Each skill acts as a bridge between your agent and specialized visual processing tools. You will find modules for text-to-image synthesis, video creation, and technical image refinement, such as watermark removal or multimedia processing. These tools are for developers who want to reduce context switching by letting their agents handle the heavy lifting of visual asset creation and manipulation. By adding these files to your agent's configuration, you can execute image generation requests directly within your terminal or editor environment using simple natural language prompts.
Top Image Generation skills
jimeng-mcp-skill
wwwzhouhui
使用jimeng-mcp-server进行AI图像和视频生成。当用户请求从文本生成图像、合成多张图片、从文本描述创建视频或为静态图像添加动画时使用此技能。支持四大核心能力:文生图、图像合成、文生视频、图生视频。需要jimeng-mcp-server在本地运行或通过SSE/HTTP访问。
gemini-logo-remover
bear2u
Remove Gemini logos, watermarks, or AI-generated image markers using OpenCV inpainting. Use this skill when the user asks to remove Gemini logo, AI watermark, or any logo/watermark from images.
ai-multimodal
mrgoonie
Process and generate multimedia content using Google Gemini API. Capabilities include analyze audio files (transcription with timestamps, summarization, speech understanding, music/sound analysis up to 9.5 hours), understand images (captioning, object detection, OCR, visual Q&A, segmentation), process videos (scene detection, Q&A, temporal analysis, YouTube URLs, up to 6 hours), extract from documents (PDF tables, forms, charts, diagrams, multi-page), generate images (text-to-image, editing, composition, refinement). Use when working with audio/video files, analyzing images or screenshots, processing PDF documents, extracting structured data from media, creating images from text prompts, or implementing multimodal AI features. Supports multiple models (Gemini 2.5/2.0) with context windows up to 2M tokens.
ai-image
tyrchen
Generate AI images using OpenAI's gpt-image-1 model with customizable aspect ratios and artistic themes. Use when the user wants to create images, generate artwork, or mentions image generation with specific styles like Ghibli, futuristic, Pixar, oil painting, or Chinese painting.
baoyu-article-illustrator
JimLiu
Analyzes article structure, identifies positions requiring visual aids, generates illustrations with Type × Style two-dimension approach. Use when user asks to "illustrate article", "add images", "generate images for article", or "为文章配图".
social-media
langchain-ai
Use this skill when creating short-form social media content for LinkedIn, Twitter/X, or other platforms
nano-banana-pro-prompts-recommend-skill
YouMind-OpenLab
Recommend suitable prompts from 6000+ Nano Banana Pro image generation prompts based on user needs. Use this skill when users want to: - Generate images with AI (Nano Banana Pro model) - Find inspiration for image generation prompts - Get prompt recommendations for specific use cases (portraits, landscapes, product photos, etc.) - Create illustrations for articles, videos, podcasts, or other content - Translate and understand prompt techniques
gifgrep
openclaw
Search GIF providers with CLI/TUI, download results, and extract stills/sheets.
image-generation
onyx-dot-app
Generate images using nano banana.
baoyu-image-gen
JimLiu
AI image generation with OpenAI, Google and DashScope APIs. Supports text-to-image, reference images, aspect ratios. Sequential by default; parallel generation available on request. Use when user asks to generate, create, or draw images.
baoyu-cover-image
JimLiu
Generates article cover images with 5 dimensions (type, palette, rendering, text, mood) combining 9 color palettes and 6 rendering styles. Supports cinematic (2.35:1), widescreen (16:9), and square (1:1) aspects. Use when user asks to "generate cover image", "create article cover", or "make cover".
nano-image-generator
solidSpoon
Generate images using Nano Banana Pro (Gemini 3 Pro Preview). Use when creating app icons, logos, UI graphics, marketing banners, social media images, illustrations, diagrams, or any visual assets. Triggers include phrases like 'generate an image', 'create a graphic', 'make an icon', 'design a logo', 'create a banner', or any request needing visual content.
baoyu-danger-gemini-web
JimLiu
Generates images and text via reverse-engineered Gemini Web API. Supports text generation, image generation from prompts, reference images for vision input, and multi-turn conversations. Use when other skills need image generation backend, or when user requests "generate image with Gemini", "Gemini text generation", or needs vision-capable AI generation.
stable-diffusion-image-generation
davila7
State-of-the-art text-to-image generation with Stable Diffusion models via HuggingFace Diffusers. Use when generating images from text prompts, performing image-to-image translation, inpainting, or building custom diffusion pipelines.
isometric-asset-sheets
amilich
Generate sprite sheet images for the isometric city game using the GenerateImage tool. Use when creating new game assets, sprite sheets, vehicle sprites, building sprites, or any visual assets for the isometric city builder. Ensures consistent format, sizing, and isometric projections.
infographic-generator
vfarcic
Generate retro arcade style infographic prompts for documentation pages
prompt-extractor
huangserva
自动化提取AI绘画提示词的模块化结构,从海量提示词中提炼可复用的模块组件
intelligent-prompt-generator
huangserva
智能提示词生成器 v2.0 - 支持人像/跨domain/设计三种模式,语义理解、常识推理、一致性检查
gap-focus-storyteller
amao2001
Generates high-quality AI image prompts focusing on "Gap Focus" or "Keyhole Reveal" composition. It emphasizes cinematic storytelling by creating a narrow visual corridor that evokes voyeurism and mystery.
gemini-api
MadAppGang
Google Gemini 3 Pro Image API reference. Covers text-to-image, editing, reference images, aspect ratios, and error handling.
ideogram-core-workflow-a
jeremylongshore
Execute Ideogram primary workflow: Core Workflow A. Use when implementing primary use case, building main features, or core integration tasks. Trigger with phrases like "ideogram main workflow", "primary task with ideogram".
midjourney-card-news-backgrounds
bear2u
Generate Midjourney prompts for 600x600 card news background images based on topic, mood, and style preferences. Use when user requests card news backgrounds or Instagram post backgrounds.
clip
davila7
OpenAI's model connecting vision and language. Enables zero-shot image classification, image-text matching, and cross-modal retrieval. Trained on 400M image-text pairs. Use for image search, content moderation, or vision-language tasks without fine-tuning. Best for general-purpose image understanding.
ideogram-hello-world
jeremylongshore
Create a minimal working Ideogram example. Use when starting a new Ideogram integration, testing your setup, or learning basic Ideogram API patterns. Trigger with phrases like "ideogram hello world", "ideogram example", "ideogram quick start", "simple ideogram code".
How to choose a Image Generation skill
Select a skill based on your specific infrastructure needs. Look at the underlying API or model; for instance, choose Gemini-based skills if you rely on Google's stack, or the ComfyUI builder if you need granular control over node-based workflows. Check the documentation for the specific input requirements, as some skills like Jimeng support video generation while others focus solely on static image editing. Consider the maintenance frequency and whether the skill requires local server components, such as a running MCP server, to function correctly.