The PodKnowledge|← All Episodes
Based on Lenny's Podcast data
Lenny's Knowledge Sketch

Expert Evals Are Creating
the Fastest-Growing AI Cos

Brendan Foody
CEO & co-founder of Mercor
SEP 18 2025
The Opportunity

Expert Knowledge +
AI Evals = Moat

EXPERT EVAL MARKET
"If the model is the product, then the eval is the product requirement document."
  • Domain experts writing evals = defensible moat that can't be replicated with scale alone
  • Medical AI needs doctors. Legal AI needs lawyers. Financial AI needs investment bankers.
  • The expert eval flywheel: better evals → better models → more expert demand → better evals
  • Brendan's thesis: expert eval companies are the fastest-growing category in AI infrastructure
Framework

The Expert Eval Business Model

DOMAIN EXPERTSEVAL WRITINGMODEL IMPROVEMENT
$95/hr
median expert pay on Mercor
1,600%+
Mercor net retention
16 mo
$1M to $400M revenue run rate
  • The network is the moat: expert annotators don't leave once they're integrated
  • Domain specificity: one-size-fits-all doesn't work for expert evals
  • Trust and verification: expert credentials must be verified, not just claimed
  • The pricing power: median $95/hr on Mercor vs ~$30/hr crowdsourcing average, up to $500/hr for deep expertise
Brendan's insightThe bottleneck for AI improvement is not compute or architecture, it's the availability of domain experts who can evaluate model outputs.
Expert Eval Categories

Where Expert Evals Create Most Value

  • Doctors and lawyers evaluating professional judgment
  • Investment bankers on financial workflows
  • Software engineers reviewing code outputs
  • Emmy-winning screenwriters and the Harvard Lampoon on creative capabilities
  • Goldman bankers, McKinsey analysts, Fang software engineers for higher-skill work
A concrete example

Say you want a model to write a redline for a contract the way a lawyer would. A lawyer creates a rubric, similar to how a professor grades a deliverable, of what "excellent" looks like.

Professional vs. creative

Labs are leaning into economically valuable professional domains, but the creative side still matters — funnier models, comedy writers, screenwriters.

Playbook

Build Expert Eval Capacity

  • Identify your 3 highest-stakes AI outputs, those need expert evals, not crowdsourced
  • Build relationships with domain experts before you need them at scale
  • Create a certification process: how do you verify expert qualification for your domain?
  • Offer performance-based compensation: experts who write better evals should earn more
The competitive moatAn expert eval network is hard to replicate at quality. It's the AI moat that money can't buy quickly.
Contrarian

Eval Business Myths

Crowdsourcing scales expert qualityINSTEAD →Crowdsourcing scales volume. Expert quality doesn't scale, it's earned through credentials and judgment.
AI can evaluate its own outputsINSTEAD →AI can evaluate syntactic correctness. Domain experts evaluate semantic correctness and safety.
Models will soon not need evalsINSTEAD →Better models need harder evals. The bar rises with capability.
One eval framework fits all domainsINSTEAD →Domains require custom eval frameworks. Medical, legal, and code evals share almost no methodology.
Based on Brendan Foody's episode on Lenny's Podcast. All ideas on this page are from the episode.Watch on YouTubeFollow @BrendanFoody on X
Every sketch here began as a transcript.

Send us yours and we’ll draw it. $1 a transcript.

Sketch mine →
𝕏︎ X / Twitterin LinkedIn📸 Instagram🔗 Copy link
0:00