AI Product Opportunity
Methodology

One formula. Evidence or a capped score.

Every opportunity on the board is scored by the same formula, from the same kinds of evidence, with the same caps. This page explains exactly how that score is computed, what the number on the ring means — and shows the honest statistics behind the leaderboards.

The formula

Four pillars, sixteen dimensions

Each idea is scored on four weighted pillars. Every pillar is the average of four dimensions, each scored 0–10 with cited evidence. A small convergence bonus (capped at +10) rewards ideas that independent sources keep resurfacing, and reader feedback can adjust the total by at most ±3 — evidence always dominates.

Demand

30%

Do people demonstrably have this problem?

  • Pain severity
  • Job frequency
  • Workaround cost
  • Reachable demand

Monetization

30%

Will someone pay, and is the model clear?

  • Willingness to pay
  • Model clarity
  • Comparable revenue
  • Cost to serve

Buildability

20%

Can a small team actually ship this?

  • Feasibility
  • Data availability
  • Time to MVP
  • Regulatory risk

Openness

20%

Is there durable white space left?

  • White space
  • Differentiation
  • Gap durability
  • Counter-positioning

Two dimensions used to sit in Demand and no longer do: market size and demand breadth. We removed them. A TAM figure has no free, checkable source — scoring one meant scoring a number nobody can verify, which is exactly what the rest of this page exists to prevent. In their place Demand asks the question that actually decides whether a product ships: reachable demand — can you find and close these buyers, through a channel that exists, at a price they already pay?

The other three additions close gaps we could name. Cost to serve is in Monetization because AI inference is a real marginal cost: the usual 65–70% software gross margin is an assumption that breaks when every request costs money. Regulatory risk is in Buildability because “can a small team ship this” includes whether they are allowed to. Counter-positioning is in Openness and asks something differentiation does not: not “are we better?” but “could the incumbent even respond, or would answering cost them their existing business?”

Opportunities scored under the earlier version of this formula keep the payload they were scored with — nothing is silently re-scored under a formula it has no judgement for — so an older opportunity’s breakdown may still list the previous set of 12 rather than today’s 16.

Anti-fabrication

No verified evidence, no high score

Any dimension whose cited evidence cannot be verified is capped at 6 out of 10 before pillar averaging. An LLM asserting “huge demand” without a checkable citation cannot push an idea to the top.

Verification is scoped per dimension, not per pillar: a URL only lifts the cap for the specific claims its kind of evidence can actually prove. A Reddit complaint proves how badly something hurts; it never proves you can reach and close those people. An arXiv abstract proves feasibility; it is not a schedule, so it cannot lift time to MVP. Scoping this by pillar was a category error in both directions, and fixing it is what let cost to serve — a claim about model size and inference load, provable from code and papers — be verified at all.

Citations are checked against the evidence the scorer was actually shown. A URL that appears in a rationale but not in the evidence set is dropped rather than counted, so a fabricated link can no longer stand in for a real one — and the “dimensions with evidence” figures below are counts of admissible citations, not of citations claimed.

One citation is not proof of a consensus. A single admissible URL lifts a dimension to 8.5 out of 10; going above that takes at least two admissible citations on different websites. Twenty threads on one site are one origin.

Two of the sixteen dimensions — white space and gap durability — assert that something does not exist, and no document can prove that. They are licensed by search coverage instead: a disconfirming search has to have actually run and surveyed live results before either claim is believed above the cap.

Evidence also expires. A dimension whose every admissible citation is older than roughly two and a half years is treated as unverified — a 2019 complaint is not evidence of demand today. An unknown publication date never trips this: absence of a timestamp is not evidence of age.

Evidence classLifts the cap for these dimensionsWhere it comes from
PainPain severity, Job frequency, Workaround costHacker News, Reddit, GitHub issues, Stack Exchange, app-store reviews, G2
BuyerWillingness to pay, Comparable revenue, Model clarity, Reachable demand, Differentiation, Counter-positioningApp Store and G2 reviewers, paid Atlassian and Shopify listings — people who already paid for something in this category
RepoFeasibility, Data availability, Time to MVP, Cost to serveGitHub repositories
ResearchFeasibility, Data availability, Cost to serve, Regulatory riskarXiv, Hugging Face
MarketReachable demand, Willingness to pay, Model clarity, Comparable revenue, Regulatory risk, Differentiation, Counter-positioningProduct Hunt, marketplace listings (Atlassian, WordPress, Shopify), Y Combinator, SEC EDGAR, verifier-supported web results
Search coverageWhite space, Gap durabilityNot a citation: a live disconfirming search that actually surveyed the field — no document can prove a product does not exist

And one thing that never enters the formula: popularity. GitHub stars, upvotes and points are provenance — a filter on what is worth reading during discovery — but they are not quality, so they carry zero weight in the score.

Reading the number

The displayed score is position and size

The formula produces an internal raw score. Because of the evidence cap, raw scores cluster below 60 — only ideas with verified evidence across their dimensions can rise above it. That clustering is honest, but it makes raw numbers hard to compare at a glance.

The displayed score used to be a pure percentile, and a pure percentile lies about distance: two ideas a tenth of a point apart can display 30 points apart if the board is thin around them, and two ideas genuinely far apart can display as neighbours. So the number you see is now half percentile position, half raw magnitude — the idea’s place on the board averaged with a direct map of its raw score onto 1–100, refreshed after every discovery run. Both halves move the same way with raw score, so this changed what the number says, never who ranks above whom.

Read it as a comparison, not a certainty: a displayed 90 means this idea sits near the top of the board and scored well in absolute terms — not 90 points of proof.

Raw-score history is kept as an append-only series, so trend lines and movers always compare raw against raw — never a displayed score against a formula total.

What the score does not say

Confidence is a second number, on purpose

Every score carries a separate confidence value and band (low, moderate or high), built from how much of the evidence cleared verification, how many distinct evidence classes and independent websites it spans, how fresh it is, whether the scorer agreed with itself, and whether a disconfirming pass actually ran. It comes with a score interval that widens as confidence falls.

It is deliberately not folded into the quality score, and the reasons are worth stating plainly. A single blended number cannot tell you the difference between a good idea with thin evidence and a weak idea with airtight evidence — and those are opposite instructions: one is a research to-do, the other is a pass. Shrinking scores by confidence would also double-count, since the cap already docks a dimension for missing evidence. Worst of all, confidence correlates with how crowded a space already is: mature markets have reviews, listings and revenue comparables; genuine white space has none of that by definition. Multiplying scores by confidence would systematically bury the exact thing this engine exists to find.

Alongside it runs an adversarial pass: a separate step searches for who already does this, names the incumbents it finds, and states the strongest objection to the idea. An incumbent it cannot point at in the evidence is recorded without a link rather than invented one.

Honest statistics

The distribution, unretouched

Live aggregates from the current board — the raw distribution every displayed score is derived from. No cherry-picking: this is the same data the leaderboards rank on.

Raw score distribution (3472 published ideas)

  • 0100
  • 10200
  • 20300
  • 30400
  • 4050158
  • 50601519
  • 60701401
  • 7080382
  • 809012
  • 901000

Displayed scores are derived from this distribution onto 1100 — half percentile position, half raw magnitude, refreshed every run. That is why we chart raw scores here, not displayed ones.

51.7%

of published ideas score a raw 60+ — above the unverified ceiling, on verified evidence

11.3%

reach a raw 70+ — the strongest evidence-verified signals on the board

26.6%

of 55,552 scored dimensions were capped for lacking verified evidence — caps we publish, not hide

Evidence source mix

  • Web search14,265
  • GitHub665
  • Hacker News629
  • GitHub issues620
  • Hugging Face450
  • Reddit182
  • Y Combinator153
  • arXiv140
  • Product Hunt124
  • Shopify App Store100
  • WordPress.org96
  • Atlassian Marketplace92
  • App Store reviews78
  • Stack Exchange78

Board bookkeeping

Published ideas
3,472
Archived (never deleted)
114
Business domains
44
Evidence items stored
17,672

Ideas that fall off a leaderboard are archived, never deleted, and re-compete if their evidence resurfaces.

Now read the board knowing what the numbers mean.

Every opportunity ships with the evidence behind its score — open the citations and check us.

← Why this exists