rfp-ingest
Automates RFP data collection and normalization into a canonical schema.
Install
mkdir -p .claude/skills/rfp-ingest && curl -L -o skill.zip "https://agentskills.codes/api/skills/download/12616" && unzip -o skill.zip -d .claude/skills/rfp-ingest && rm skill.zipInstalls to .claude/skills/rfp-ingest
Activation
This is the description your AI agent reads to decide when to run this skill — the better it matches your request, the more reliably it fires.
Ingest RFP opportunities from multiple data sources (SAM.gov, eMMA, RFPMart). Use when adding new data sources, modifying ingestion logic, or debugging data fetching issues.Key capabilities
- →Ingest RFP data from SAM.gov, eMMA, RFPMart, and GovTribe
- →Normalize RFP data into a canonical schema
- →Deduplicate ingested RFP opportunities
- →Detect eligibility flags based on text patterns
- →Log ingestion status and record counts
- →Handle API errors with retries or alerts
How it works
The skill fetches RFP data from specified sources, normalizes it to a standard schema, and stores it in a database, handling deduplication and eligibility flagging.
Inputs & outputs
When to use rfp-ingest
- →Ingesting new RFP data
- →Adding a new RFP data source
- →Debugging data fetching logic
About this skill
RFP Ingestion Skill
Overview
This skill helps implement multi-source RFP data ingestion with canonical schema normalization and deduplication.
Supported Data Sources
| Source | Priority | API Type | Rate Limits |
|---|---|---|---|
| SAM.gov | P1 | REST API | 10 req/sec, 10k/day |
| Maryland eMMA | P1 | Web scraping | Respectful crawling |
| RFPMart | Current | REST API | As documented |
| GovTribe | P2 | REST API (paid) | Per subscription |
Canonical Schema
All sources must normalize to this schema:
interface Opportunity {
externalId: string; // Source-specific ID
source: "sam.gov" | "emma" | "rfpmart" | "govtribe";
title: string;
description: string;
summary?: string;
location: string;
category: string;
naicsCode?: string;
setAside?: string; // "Small Business", "8(a)", etc.
postedDate: number; // Unix timestamp
expiryDate: number; // Unix timestamp
url: string;
attachments?: Attachment[];
eligibilityFlags?: string[];
rawData: Record<string, unknown>;
ingestedAt: number;
}
SAM.gov Integration
API Endpoint
https://api.sam.gov/opportunities/v2/search
Required Headers
{
"Accept": "application/json",
"X-Api-Key": process.env.SAM_GOV_API_KEY
}
Example Query
const params = new URLSearchParams({
postedFrom: "2024-01-01",
postedTo: "2024-12-31",
limit: "100",
offset: "0",
ptype: "o", // Opportunities only
});
Field Mapping
| SAM.gov Field | Canonical Field |
|---|---|
noticeId | externalId |
title | title |
description | description |
postedDate | postedDate (parse to timestamp) |
responseDeadLine | expiryDate (parse to timestamp) |
placeOfPerformance.state | location |
naicsCode | naicsCode |
setAsideDescription | setAside |
Convex Implementation
Ingestion Action
// convex/ingestion.ts
import { action, internalMutation } from "./_generated/server";
import { v } from "convex/values";
import { internal } from "./_generated/api";
export const ingestFromSam = action({
args: { daysBack: v.optional(v.number()) },
handler: async (ctx, args) => {
const apiKey = process.env.SAM_GOV_API_KEY;
if (!apiKey) throw new Error("SAM_GOV_API_KEY not configured");
const fromDate = new Date();
fromDate.setDate(fromDate.getDate() - (args.daysBack ?? 7));
const response = await fetch(
`https://api.sam.gov/opportunities/v2/search?` +
`api_key=${apiKey}&postedFrom=${fromDate.toISOString().split("T")[0]}&limit=100`,
{ headers: { Accept: "application/json" } }
);
if (!response.ok) {
throw new Error(`SAM.gov API error: ${response.status}`);
}
const data = await response.json();
let ingested = 0;
let updated = 0;
for (const opp of data.opportunitiesData ?? []) {
const result = await ctx.runMutation(internal.rfps.upsert, {
externalId: opp.noticeId,
source: "sam.gov",
title: opp.title ?? "Untitled",
description: opp.description ?? "",
location: opp.placeOfPerformance?.state ?? "USA",
category: opp.naicsCode ?? "Unknown",
postedDate: new Date(opp.postedDate).getTime(),
expiryDate: new Date(opp.responseDeadLine).getTime(),
url: `https://sam.gov/opp/${opp.noticeId}/view`,
rawData: opp,
});
if (result.action === "inserted") ingested++;
else updated++;
}
// Log ingestion
await ctx.runMutation(internal.ingestion.logIngestion, {
source: "sam.gov",
status: "completed",
recordsProcessed: data.opportunitiesData?.length ?? 0,
recordsInserted: ingested,
recordsUpdated: updated,
});
return { ingested, updated, source: "sam.gov" };
},
});
Upsert Mutation
// convex/rfps.ts (internal mutation)
export const upsert = internalMutation({
args: {
externalId: v.string(),
source: v.string(),
title: v.string(),
description: v.string(),
location: v.string(),
category: v.string(),
postedDate: v.number(),
expiryDate: v.number(),
url: v.string(),
rawData: v.optional(v.any()),
},
handler: async (ctx, args) => {
const existing = await ctx.db
.query("rfps")
.withIndex("by_external_id", (q) =>
q.eq("externalId", args.externalId).eq("source", args.source)
)
.first();
const now = Date.now();
if (existing) {
await ctx.db.patch(existing._id, { ...args, updatedAt: now });
return { id: existing._id, action: "updated" as const };
}
const id = await ctx.db.insert("rfps", {
...args,
ingestedAt: now,
updatedAt: now,
});
return { id, action: "inserted" as const };
},
});
Deduplication Strategy
- Exact match:
externalId+sourcecombination - Title similarity: Fuzzy match titles within same deadline window
- URL canonicalization: Normalize URLs before comparison
Eligibility Pre-Filtering
Detect disqualifiers during ingestion:
const DISQUALIFIER_PATTERNS = [
{ pattern: /u\.?s\.?\s*(citizen|company|organization)\s*only/i, flag: "us-org-only" },
{ pattern: /onshore\s*(only|required)/i, flag: "onshore-required" },
{ pattern: /on-?site\s*(required|mandatory)/i, flag: "onsite-required" },
{ pattern: /security\s*clearance\s*required/i, flag: "clearance-required" },
{ pattern: /small\s*business\s*set[- ]aside/i, flag: "small-business-set-aside" },
];
function detectEligibilityFlags(text: string): string[] {
return DISQUALIFIER_PATTERNS
.filter(({ pattern }) => pattern.test(text))
.map(({ flag }) => flag);
}
Scheduled Ingestion
// convex/crons.ts
import { cronJobs } from "convex/server";
import { internal } from "./_generated/api";
const crons = cronJobs();
crons.interval(
"ingest-sam-gov",
{ hours: 6 },
internal.ingestion.ingestFromSam,
{ daysBack: 3 }
);
export default crons;
Error Handling
| Error Type | Action |
|---|---|
| Rate limit (429) | Exponential backoff, retry after delay |
| Auth error (401/403) | Log error, alert admin |
| Server error (5xx) | Retry up to 3 times |
| Parse error | Log raw data, skip record |
Testing Approach
- Mock API responses for unit tests
- Use sandbox/test endpoints when available
- Validate schema transformation
- Test deduplication logic
- Verify eligibility flag detection
When not to use it
- →When the data source is not one of the supported platforms
- →When a custom schema is required that deviates from the canonical Opportunity interface
- →When real-time ingestion is needed without batch processing
Prerequisites
Limitations
- →Only supports SAM.gov, Maryland eMMA, RFPMart, and GovTribe as data sources
- →Relies on a predefined canonical schema for RFP opportunities
- →Rate limits and API errors are handled according to defined strategies
How it compares
This skill automates data fetching, normalization, and deduplication across multiple RFP sources, unlike manual collection and processing.
Compared to similar skills
rfp-ingest side by side with the closest alternatives in the catalog.
| Skill | Installs | Updated | Safety | Difficulty |
|---|---|---|---|---|
| rfp-ingest (this skill) | 0 | 6mo | Review | Intermediate |
| apify-ultimate-scraper | 0 | 3mo | Review | Intermediate |
| web-scraper | 0 | 1mo | No flags | Intermediate |
| apify-trend-analysis | 0 | 3mo | Review | Intermediate |
Try saying
Example prompts that trigger this skill in your AI assistant.
You might also like
apify-ultimate-scraper
Anhvu1107
ALWAYS use this when the request matches Apify Ultimate Scraper: AI-driven data extraction from 55+ Actors across all major platforms.
web-scraper
shobcoder
Scrape, crawl, and extract data from websites. Use when users ask to scrape web pages, extract content, crawl websites, or collect data from the internet.
apify-trend-analysis
Anhvu1107
ALWAYS use this when the request matches Apify Trend Analysis: Discover and track emerging trends across Google Trends, Instagram, Facebook, YouTube, and TikTok to inform content strategy.
segment-cdp
davila7
Expert patterns for Segment Customer Data Platform including Analytics.js, server-side tracking, tracking plans with Protocols, identity resolution, destinations configuration, and data governance best practices. Use when: segment, analytics.js, customer data platform, cdp, tracking plan.
databuddy
databuddy-analytics
Integrate Databuddy analytics into applications using the SDK or REST API. Use when implementing analytics tracking, feature flags, custom events, Web Vitals, error tracking, LLM observability, or querying analytics data programmatically.
analytics-pipeline
dadbodgeoff
Real-time analytics with Redis counters, periodic PostgreSQL flush, and time-series aggregation. High-performance event tracking without database bottlenecks.