VP blog / native X article — humanized final draft

There's a pattern in how people are adapting Karpathy's autoresearch.

On February 27, Thomas Wolf asked why the NanoGPT speedrun challenge isn't fully AI-automated yet. Karpathy responded the next day. Ten days later he dropped autoresearch. 8 million views. 20K+ stars in 48 hours.

People started porting the loop everywhere. One team pointed it at markets: 25 AI agents debating macro, scoring picks against outcomes, worst performer gets its prompt rewritten by the system. 378 iterations. 54 prompt modifications. 16 survived. +22% in 173 days.

Hypothesis, test, score, keep or discard, repeat.

But every adaptation I've seen has a clean metric at the end. Training loss. Sharpe ratios. Click-through rates. Things with a number.

I wanted to know what happens when you point the loop at decisions that don't have a number.

So I tried it with investment research. Not the code, the methodology. And the first useful output wasn't what the search returned. It was what it didn't.

I ran a structured search across Capital Allocators: 491 episodes over 8 years, the most important podcast in the LP/allocator world. Searched every transcript and all 706 YouTube uploads for MENA sovereign wealth. ADIA, Mubadala, PIF, QIA, KIA, ADQ.

Three sovereign wealth episodes. Norway twice, Australia once. Zero MENA.

Same gap across 20VC, Invest Like the Best, Venture Unlocked, Alt Goes Mainstream. Thousands of episodes across five shows. Near-zero Gulf coverage.

Here's the thing: MENA sovereign wealth funds deployed $56.3 billion in the first nine months of 2025, with the US as top destination. That's 54% of all sovereign capital deployed globally. Mubadala alone put $12.9 billion into AI. PIF announced a $100 billion AI fund. A third of all sovereign capital on earth, $4.9 trillion, doesn't show up in the content the US venture industry uses to understand its own LP base.

Any search tool would tell you "491 episodes analyzed, comprehensive survey." A tool that tracks gaps would tell you "zero coverage on the capital source writing the biggest checks in your market."

Karpathy's system works because evaluation is locked down. Same experiment twice, same score. Investment decisions aren't like that. The data is always incomplete, outcomes take years, and the variables that matter most are ones you can't Google.

The loop still works. You just have to change what you're measuring.

First, write down what you expect to find before you search. I've talked to enough VCs to know that most start with a thesis and look for confirmation. Writing down the prediction is the only thing that catches you doing it.

Second, track three states instead of two. Keep, discard, or insufficient data. "I searched for independent revenue verification and found nothing" is not the same as "revenue is unverified." It means the data doesn't exist publicly. That's its own signal, and it generates a specific question: who do you call to get this?

Third, log what you couldn't find. After 50 decisions, that log is the more interesting one. It maps where public information stops in your space, which claims are never verifiable from a desk, which gaps only your network can fill.

Every VC has a deal story like this. Deck was polished, metrics looked right, market was big. One phone call to a customer or a former employee surfaced the thing that wasn't in any database.

That phone call is what the tool should produce. Not the answer. The question. And who to ask.

I'm building this and running it against real portfolios. If you want in, DM me.

Tweet flow (article-first format)

Launch tweet (Day 1)

I ran a structured search across 5 major VC podcasts. 3,000+ episodes. Looking for MENA sovereign wealth coverage.

Three episodes. Norway twice, Australia once. Zero Gulf.

These funds deployed $56.3B last year. The US was their #1 destination.

Wrote about what this means for research tools:

[article link]

Follow-up thread (3 tweets)

1/ Everyone adapting Karpathy's autoresearch is pointing it at clean metrics. Sharpe ratios, click-through rates, training loss. Things with a number at the end.

What happens when you point the loop at decisions that don't have a number?

2/ The most useful output from my investment research loop wasn't what the search found. It was what it didn't.

$4.9 trillion in MENA sovereign capital. 54% of all sovereign deployment globally. Near-zero coverage across the top 5 LP/allocator podcasts.

3/ I'm adapting the autoresearch methodology for VC due diligence. Three changes: pre-register your priors, track "insufficient data" as its own state, and log the gaps.

After 50 decisions, what you couldn't find is more valuable than what you did.

Day 3 repost (different hook)

The absence of data is data.

I searched 3,000+ episodes of the top VC podcasts for MENA sovereign wealth coverage. The gap I found says more about the US venture ecosystem than any episode could.

[article link]

Day 7 quote-tweet

Still thinking about this. The venture industry has a $4.9 trillion blind spot and nobody's building tools to surface it.

Sources (verified)

◆   ◆   ◆