An automated pipeline that discovers AI startups on Product Hunt, enriches their websites, and scores them against a weighted rubric — built to prove the rubric holds up before spending money or taking on ToS risk with paid data sources. Along the way, two findings changed the build: the rubric itself was tuned wrong, and the "official" data source turned out not to be the only honest one.
Product Hunt GraphQL API, or a token-free public-feed fallback. Same output schema either way.
Playwright visits each resolved domain for pricing/docs pages, socials, and a GitHub link.
A four-category weighted rubric, composite 0–100, every sub-score kept.
Top-decile vs. random sample from a historical window, checked against real outcomes.
Sorted CSV/JSON, plus the Scout Deck dashboard shown below.
The rubric scores four categories — Traction, Team, Market/Product, Momentum — and three of its ten terms started as plain flags: was a LinkedIn page found, a pricing page, a docs/demo page. Before shipping any weight, each was swept across a grid and checked against the top 20 by composite score. The result: moving a flag's weight from zero to the very first nonzero value tested reordered more than half the leaderboard. There was no small safe weight to land on — the moment a binary term entered the sum, it dominated it.
The fix wasn't a smaller weight, it was a different shape. Each flag was folded into a fractional signal instead: team_social_presence is the share of {LinkedIn, Twitter/X} actually found, market_site_maturity is the share of {pricing, docs/demo} pages found. Zero standalone binary terms ship in the rubric that runs today.
Product Hunt's GraphQL API returns 401 without a token. But its public Atom feed and product pages return 200 to a plain HTTP request, no login, no automation — and their static HTML already carries the real website domain (past any tracking redirect), a GitHub link, topics, and the full maker list. That's most of what the official API gives you, for free, from pages Product Hunt already serves to anyone.
What it doesn't give you: vote and comment counts only render client-side via JavaScript, and loading the page with an automated browser to read them triggers Product Hunt's Cloudflare bot challenge. That line wasn't crossed — defeating a platform's active bot detection was out of scope no matter which source was involved. So the token-free path ships as a clearly labeled fallback: Traction and vote-based Momentum score zero for every company, visible in the dashboard below as a hatched "not measured" segment rather than a silent zero.
A live run: 50 launches scanned, 29 confirmed AI-related, all 29 enriched with zero failures, scored, and ranked — captured in the token-free mode described above.
Built, unit-tested, and verified against a live run — 29 real companies, 0 enrichment failures, 11 tests passing.
The official API (needs a token, full history, votes/comments) and a public-feed fallback (no token, ~50 recent launches, no vote data).
Harness is built — samples a top-decile and a random group from a 6–12-month-old window and compares good-outcome rates. Needs the API token; the feed source has no date-range window to draw a historical sample from.
Crunchbase, Wellfound, and outreach automation stay explicitly out of scope until the rubric is backtest-proven.