LLMs Outscore Investor Benchmarks in Startup-Picking Test

Oxford and Vela Research introduced VCBench, a benchmark testing whether language models can identify startups that later achieve major milestones. Several models outperformed investor baselines on the anonymized dataset, though retrospective scores do not prove an advantage in live deal sourcing. The open benchmark gives researchers and investors a way to test reproducibility, bias and practical value before relying on AI for investment decisions.

INVESTINGUSAGEFUTURETOOLS

The AI Maker

12/21/20262 min read

Glowing data streams converge on a taller planter with a bright sprout, among smaller planters and a rising bar chart.
Glowing data streams converge on a taller planter with a bright sprout, among smaller planters and a rising bar chart.

A new benchmark from University of Oxford (https://www.ox.ac.uk/) and Vela Research (https://vela.partners/research/vcbench) suggests that several large language models can identify successful startups more accurately than common venture-capital benchmarks. In tests using anonymized founder profiles and company data, GPT-4o (https://openai.com/index/hello-gpt-4o/) recorded the strongest F0.5 score, while DeepSeek-V3 (https://www.deepseek.com/news/deepseek-v3/) delivered more than six times the precision of the market baseline.

The study, “VCBench: Benchmarking LLMs in Venture Capital” (https://arxiv.org/abs/2509.14448), evaluates whether AI systems can distinguish founders whose companies later reached major milestones, such as an exit or initial public offering. Its dataset contains 9,000 anonymized founder profiles, around 810 of which were labeled successful. That small share reflects a familiar challenge in startup investing: major outcomes are rare, making reliable prediction difficult.

The researchers removed names and direct identifiers to limit the chance that models could rely on recognition of prominent founders. They also conducted adversarial tests intended to detect re-identification from public information. The paper reports that these measures reduced re-identification risk by 92 percent while retaining features useful for prediction.

VCBench compares model results with several investor benchmarks. The market baseline recorded 1.9 percent precision, or about one successful company for every 50 selections. Y Combinator (https://www.ycombinator.com) reached 3.2 percent, while tier-one venture firms reached about 5.6 percent. DeepSeek-V3 exceeded six times the market’s precision, and GPT-4o led on F0.5, a measure that combines precision and recall while placing greater weight on precision. Claude 3.5 Sonnet (https://www.anthropic.com/news/claude-3-5-sonnet) and Gemini 1.5 Pro (https://blog.google/technology/ai/google-gemini-next-generation-model-february-2024/) also performed above the market baseline.

Those results suggest that leading models can rank potential investments effectively in this test. They do not establish that an AI system would outperform experienced investors across live deal flow. A retrospective benchmark depends on how success is defined, which founders appear in the data, and whether the information available to a model matches what investors could have known at the time. Those questions matter particularly in a field where historical outcomes are unevenly distributed and incomplete.

The researchers have made VCBench available at vcbench.com and invited others to evaluate models against it. An open benchmark could help researchers test whether results hold across systems and methods, rather than relying on a single paper’s comparison. It may also expose weaknesses that aggregate scores conceal, including whether models favor particular founder backgrounds or company profiles.

For venture firms, the near-term use case is likely screening and prioritization, not handing investment decisions to a model. AI-assisted tools could help surface less obvious candidates or organize large volumes of founder information, but human diligence would remain necessary to assess markets, teams and execution. The benchmark’s next test is whether its results can be reproduced on new data and translated into better decisions outside a controlled evaluation.

Cited: https://decrypt.co/340418/ai-now-way-better-predicting-startup-success-vcs

Your Data, Your Insights

Unlock the power of your data effortlessly. Update it continuously. Automatically.

Answers

Sign up NOW

info at aimaker.com

© 2024. All rights reserved. Terms and Conditions | Privacy Policy