Case study · Personal project · Phase 1

Signal & Noise

An automated pipeline that discovers AI startups on Product Hunt, enriches their websites, and scores them against a weighted rubric — built to prove the rubric holds up before spending money or taking on ToS risk with paid data sources. Along the way, two findings changed the build: the rubric itself was tuned wrong, and the "official" data source turned out not to be the only honest one.

4
rubric categories, all sub-scores stored
2
ingestion sources, one needs zero credentials
11
unit tests passing
$0
spent — free-tier APIs and public pages only
01

Ingest

Product Hunt GraphQL API, or a token-free public-feed fallback. Same output schema either way.

02

Enrich

Playwright visits each resolved domain for pricing/docs pages, socials, and a GitHub link.

03

Score

A four-category weighted rubric, composite 0–100, every sub-score kept.

04

Backtest

Top-decile vs. random sample from a historical window, checked against real outcomes.

05

Rank & view

Sorted CSV/JSON, plus the Scout Deck dashboard shown below.

Finding 01

The rubric that broke itself

Methodology

Three "yes/no" signals were quietly deciding the entire ranking

The rubric scores four categories — Traction, Team, Market/Product, Momentum — and three of its ten terms started as plain flags: was a LinkedIn page found, a pricing page, a docs/demo page. Before shipping any weight, each was swept across a grid and checked against the top 20 by composite score. The result: moving a flag's weight from zero to the very first nonzero value tested reordered more than half the leaderboard. There was no small safe weight to land on — the moment a binary term entered the sum, it dominated it.

The fix wasn't a smaller weight, it was a different shape. Each flag was folded into a fractional signal instead: team_social_presence is the share of {LinkedIn, Twitter/X} actually found, market_site_maturity is the share of {pricing, docs/demo} pages found. Zero standalone binary terms ship in the rubric that runs today.

weight 0 weight 3 100% 0% 55% top-20 dominance share, LinkedIn-presence term
First grid point tested past zero: 55% of the top‑20 already carried the flag. Full sweep and the fix are documented in src/scoring.py.
Finding 02

The data source hiding in plain sight

Sourcing

No developer token doesn't have to mean no data — but it has a real ceiling

Product Hunt's GraphQL API returns 401 without a token. But its public Atom feed and product pages return 200 to a plain HTTP request, no login, no automation — and their static HTML already carries the real website domain (past any tracking redirect), a GitHub link, topics, and the full maker list. That's most of what the official API gives you, for free, from pages Product Hunt already serves to anyone.

What it doesn't give you: vote and comment counts only render client-side via JavaScript, and loading the page with an automated browser to read them triggers Product Hunt's Cloudflare bot challenge. That line wasn't crossed — defeating a platform's active bot detection was out of scope no matter which source was involved. So the token-free path ships as a clearly labeled fallback: Traction and vote-based Momentum score zero for every company, visible in the dashboard below as a hatched "not measured" segment rather than a silent zero.

GET /feed → 200 GraphQL, no token → 401 product page, plain HTTP → 200 product page, Playwright → Cloudflare challenge
Scout Deck

What it looks like

A live run: 50 launches scanned, 29 confirmed AI-related, all 29 enriched with zero failures, scored, and ranked — captured in the token-free mode described above.

Where it actually stands

Shipped vs. still open

Ingestion, enrichment, scoring, ranking, dashboard

Built, unit-tested, and verified against a live run — 29 real companies, 0 enrichment failures, 11 tests passing.

Two ingestion sources

The official API (needs a token, full history, votes/comments) and a public-feed fallback (no token, ~50 recent launches, no vote data).

Historical backtest

Harness is built — samples a top-decile and a random group from a 6–12-month-old window and compares good-outcome rates. Needs the API token; the feed source has no date-range window to draw a historical sample from.

Phase 2 sources

Crunchbase, Wellfound, and outreach automation stay explicitly out of scope until the rubric is backtest-proven.