groq-deploy-integration
Recipes for deploying Groq-based applications to production, including edge function handlers and container configuration.
Install
mkdir -p .claude/skills/groq-deploy-integration && curl -L -o skill.zip "https://agentskills.codes/api/skills/download/5978" && unzip -o skill.zip -d .claude/skills/groq-deploy-integration && rm skill.zipInstalls to .claude/skills/groq-deploy-integration
Activation
This is the description your AI agent reads to decide when to run this skill — the better it matches your request, the more reliably it fires.
Deploy Groq integrations to Vercel, Cloud Run, and containerized platforms.Key capabilities
- →Deploy Groq applications to Vercel Edge
- →Deploy Groq applications to Cloud Run
- →Deploy Groq applications to Docker containers
- →Configure platform-specific secrets for Groq API keys
- →Add health checks for Groq integrations
How it works
The skill provides recipes for writing API handlers, storing secrets securely on platforms like Vercel or Cloud Run, deploying the application, and adding a health check endpoint.
Inputs & outputs
When to use groq-deploy-integration
- →Deploying Groq-powered apps to Vercel Edge
- →Containerizing AI applications for Cloud Run
- →Configuring production environment variables for Groq
- →Setting up streaming server-side handlers for AI models
About this skill
Groq Deploy Integration
Overview
Deploy applications using Groq's inference API to Vercel Edge, Cloud Run, Docker, and other platforms. Groq's sub-200ms latency makes it ideal for edge deployments and real-time applications.
This SKILL.md is the high-level workflow. Every platform recipe — full source for the Vercel Edge Function, Dockerfile, Cloud Run command, Express health-check server, and Vercel AI SDK handler — lives verbatim in references/implementation.md. End-to-end walkthroughs that chain those recipes are in references/examples.md.
Prerequisites
- Groq API key stored in
GROQ_API_KEY - Application using
groq-sdk(or@ai-sdk/groqfor the Vercel AI SDK path) - Platform CLI installed (
vercel,docker, orgcloud)
Instructions
Pick the deployment target, then follow its recipe in references/implementation.md.
- Write the handler. For Vercel Edge, create
app/api/chat/route.tswithexport const runtime = "edge"and stream Server-Sent Events when the request asks for them; otherwise return a JSON completion. See Step 1 inreferences/implementation.md. - Store the secret. Never bake
GROQ_API_KEYinto an image. Use the platform's secret store — see the Environment Variable Config table below. - Deploy.
vercel --prodfor Vercel (Step 2); build the Dockerfile (Step 3) andgcloud run deploy --source .for Cloud Run (Step 4) — all inreferences/implementation.md. - Add a health check. The Express server (Step 5) exposes
/healththat pings Groq with the cheapest model (llama-3.1-8b-instant,max_tokens: 1) and reports latency, so orchestrators can probe liveness cheaply. - Keep instances warm. On serverless platforms set
min-instances=1to keep cold-start latency off the request path.
The essential Vercel Edge skeleton looks like this — the full streaming body is in the reference:
// app/api/chat/route.ts
import Groq from "groq-sdk";
export const runtime = "edge";
export async function POST(req: Request) {
const groq = new Groq({ apiKey: process.env.GROQ_API_KEY! });
const { messages } = await req.json();
const completion = await groq.chat.completions.create({
model: "llama-3.3-70b-versatile",
messages,
max_tokens: 2048,
});
return Response.json(completion);
}
Environment Variable Config
| Platform | Command |
|---|---|
| Vercel | vercel env add GROQ_API_KEY production |
| Cloud Run | gcloud secrets create groq-api-key --data-file=- |
| Fly.io | fly secrets set GROQ_API_KEY=gsk_... |
| Railway | railway variables set GROQ_API_KEY=gsk_... |
| Docker | -e GROQ_API_KEY=gsk_... or Docker secrets |
Output
Following this skill produces:
- A deployed Groq inference endpoint (
POST /api/chat) on the chosen platform that streamstext/event-streamchunks on demand and returns JSON completions otherwise. - The secret registered in the platform's secret store — never committed to source or an image layer.
- A
/healthliveness endpoint returning{ status: "healthy", groq: { connected: true, latencyMs: N } }(HTTP 200) or{ status: "unhealthy", ... }(HTTP 503) for orchestrator probes. - A warm serverless configuration (
min-instances=1) keeping cold-start latency off the request path.
Error Handling
| Issue | Cause | Solution |
|---|---|---|
| Rate limited (429) | Too many requests | Implement request queuing with backoff |
| Edge timeout | Response > 25s | Use streaming for long completions |
| Model unavailable | Capacity or deprecation | Fall back to llama-3.1-8b-instant |
| Cold start latency | Serverless function init | Set min-instances=1 on Cloud Run |
| API key not found | Secret not configured | Check platform secret config |
Examples
Full worked walkthroughs live in references/examples.md:
- Example A — Vercel Edge streaming chat: drop in the Step 1 handler,
vercel env add+vercel --prod, get a streamingPOST /api/chatURL. - Example B — Cloud Run with a liveness probe: Dockerfile
HEALTHCHECK+ Express/health+gcloud run deploy --min-instances=1, yielding a 200/503 health signal Cloud Run consumes. - Example C — Vercel AI SDK path: swap the raw client for
@ai-sdk/groqstreamText+toDataStreamResponse()for zero manual stream plumbing.
Resources
- Groq API Documentation
- Vercel AI SDK + Groq
- Groq Client Libraries
- Full implementation recipes
- Worked examples
Next Steps
For multi-environment setup (separate dev/staging/prod secrets and pipelines), see the groq-multi-env-setup skill in this pack.
When not to use it
- →When not deploying Groq-powered applications
- →When not using Vercel, Cloud Run, or Docker for deployment
- →When a different AI SDK than Vercel AI SDK is preferred
Prerequisites
Limitations
- →Edge timeout occurs if response is over 25 seconds
- →Model unavailability may require falling back to a different model
- →Cold start latency can occur on serverless functions
How it compares
This skill offers platform-specific deployment instructions for Groq applications, ensuring secure secret handling and performance-tuned health checks.
Compared to similar skills
groq-deploy-integration side by side with the closest alternatives in the catalog.
| Skill | Installs | Updated | Safety | Difficulty |
|---|---|---|---|---|
| groq-deploy-integration (this skill) | 1 | 26d | Review | Intermediate |
| openevidence-deploy-integration | 1 | 26d | Caution | Intermediate |
| apollo-deploy-integration | 1 | 26d | Caution | Intermediate |
| agent-sandbox | 1 | 6mo | No flags | Intermediate |
Try saying
Example prompts that trigger this skill in your AI assistant.
More by jeremylongshore
View all by jeremylongshore →You might also like
openevidence-deploy-integration
jeremylongshore
Deploy OpenEvidence integrations to healthcare production environments. Use when deploying to production, setting up staging environments, or configuring cloud deployments for clinical AI applications. Trigger with phrases like "deploy openevidence", "openevidence staging", "openevidence production deploy", "release openevidence".
apollo-deploy-integration
jeremylongshore
Deploy Apollo.io integrations to production. Use when deploying Apollo integrations, configuring production environments, or setting up deployment pipelines. Trigger with phrases like "deploy apollo", "apollo production deploy", "apollo deployment pipeline", "apollo to production".
agent-sandbox
ruvnet
Agent skill for sandbox - invoke with $agent-sandbox
docker-node-version-compat-modules
Disentinel
|
DevOps Engineer Skills
lanmata
Consolidated skill set for the DevOps Engineer agent — Maven build, Docker, GitHub Actions, CI/CD orchestration, and release management
gcp-cloud-run
mk-knight23
Specialized skill for building production-ready serverless applications on GCP. Covers Cloud Run services (containerized), Cloud Run Functions (event-driven), cold start optimization, and event-driven architecture with Pub/Sub.