Sitemap

On what actually predicts YC startup success

8 min readApr 30, 2026

Since Rebel Fund invests exclusively in seed-stage Y Combinator startups, the dozens of blog posts I’ve published over the years are largely focused on what separates the future YC unicorns from the rest of the pack. As my long-time readers know, Rebel maintains the most comprehensive database of YC startups, founders, and outcomes that exists outside of YC itself, and we use the millions of data points we’ve collected to train our proprietary Rebel Theorem ML algorithm to predict which YC startups are most likely to deliver outsized investor returns.

Last year I published On Rebel Theorem 4.0, unveiling the latest release of that algorithm after 5+ years of development. But our work doesn’t stop there — we also now compliment our core ML algorithm with a series of AI-derived scores on each startup to create a “composite” ML/AI score that is much more predictive of YC startup success than Rebel Theorem scores alone.

This post takes a step back from Rebel Theorem to focus on some of the AI-derived scores that we now assign to every YC startup before Demo Day, and which best predict a startup’s likelihood to raise a Series A round or beyond.

The setup

To answer the question rigorously, we worked from a clean dataset of 1,275 startups across five recent YC batches (S22 through S24). Of those companies, roughly 100 — about 8% — have already raised a Series A round or beyond¹. Before I talk about which startup characteristics produce the highest statistical lift over that baseline, I’d like to illustrate just how powerful our composite ML/AI scores are.

The chart below shows how a company’s Series A+ likelihood climbs by our composite ML/AI score band. Companies in the highest score band raise a Series A or beyond at ~4x the baseline average.

Press enter or click to view image in full size

The gains curve chart below illustrates that by focusing on the top ~10% of YC companies per our composite score, we capture ~4x as many Series A+ winners than we’d expect due to chance. While we’d capture even more winners in absolute terms if we invested beyond the top ~10% of each batch, we’d get diminishing returns in terms of Series A+ hit rate.

Press enter or click to view image in full size

The signals that matter most

1. Progress

The single strongest predictor of future success we found is what our internal scoring system calls “Progress” — essentially, a structured AI-derived measure of how much real progress a company has made on its product and business by the time we evaluate it. Progress encompasses things like whether they’ve shipped a real product, whether they’re actively iterating with customers, whether they’re moving fast, and whether their progress feels proportionate to how long they’ve been at it.

Companies in the top 10% on Progress are several times more likely to reach Series A than a randomly chosen YC startup. It’s the single strongest signal in the entire dataset.

The intuition is straightforward: Series A investors aren’t betting on ideas, they’re betting on momentum. A company that has already shipped, learned, iterated, and shipped again is dramatically more likely to keep doing that, which is exactly what gets you to a Series A round.

2. Product

Closely behind Progress sits what we call Product — an AI-derived measure of how thoughtfully the actual product itself is built. Is the founder’s core insight a real one? Is the product elegant rather than bloated? Is there evidence of taste in how decisions were made?

Top-10% Product scores also deliver a lift several times above baseline.

It’s worth pausing on this for a moment. The two strongest qualitative signals we found are both about the thing the founders have built. It’s not about the market size, the cap table, or deal heat. It’s about the work itself and the founders’ rate of progress on it.

3. Revenue

After Progress and Product, the next strongest signal is Revenue. This isn’t surprising since Revenue growth is downstream of Product and Progress, but the magnitude of the effect at even quite modest revenue tiers is.

Even a company with as little as $100K in ARR is more than twice as likely to reach Series A as a random YC startup. By $1M+ in ARR, you’re looking at a lift comparable to the strongest qualitative signals.

The takeaway for investors: don’t dismiss small-revenue YC startups as “too early” to evaluate. Even a few hundred thousand in ARR at the seed stage is a meaningfully positive signal. And for founders: getting to even modest paid traction before Demo Day puts you in a substantially different statistical bucket than companies that haven’t yet.

4. Founder fundraising track record

Founders who’ve previously raised significant capital for prior ventures show up as a real signal — though more modest than the ones above. Specifically, we found that the more capital a founding team has previously raised, the more likely their current YC startup is to reach Series A.

A subtle but important point: this is “have you previously raised capital,” not “have you previously exited.” The two are correlated but distinct, and prior fundraising turns out to be the stronger of the two signals at the YC Demo Day stage.

5. Rebel Theorem scores

Unsurprisingly, our Rebel Theorem scores also produce a significant lift for companies scoring well. It’s a useful signal, especially given that it operates on every company objectively based on hundreds of discrete data points, and these scores are additive to our newer AI-derived scores.

But the more interesting finding is that the Rebel Theorem score is asymmetric. The lift at the top end is moderate, but the downside signal at the bottom end is sharper: companies with very low scores reach Series A at rates well below the baseline. In other words, Rebel Theorem is a stronger filter for ruling weak companies out than for pulling strong companies in.

That’s an important insight for any data-driven investor building this kind of model. A score doesn’t have to be a perfect oracle of upside to be valuable — being a reliable filter on the downside is, on its own, a hugely useful property at the seed stage, where avoiding losers is half the battle.

6. Team quality

Now to the result that surprised us most. Of the dozen or so AI scoring dimensions we assess, the one called Team — covering things like founder pedigree, technical skill, prior experience, and complementary co-founder skill sets — produced only a middling lift.

When we dove in deeper, we found the situation is actually quite nuanced. It’s not that Team scores don’t matter, it’s that they only matter at the extreme. In other words, really good teams perform extraordinarily well, but nearly ok or good teams don’t do much better than average.

Bear in mind that YC’s own admissions process is, among other things, an exceptionally rigorous filter on team quality. YC’s admissions rate is <1%, so by the time a company shows up in a YC batch, the variance in team quality across the cohort is already compressed dramatically relative to the broader founder population. The Team signal still discriminates within that compressed range — just less powerfully than dimensions like Progress or Product, where there’s more variance to work with.

This squares with something we’ve believed for a long time at Rebel: at the YC Demo Day stage, the question isn’t really “is this a good team?” since most of them are. The question is “is this an exceptional team and are they executing quickly on the right thing?”

The signals that don’t matter (as much)

Just as interesting as the signals that work are the ones that don’t. A few worth calling out:

Launch timing

How early a company launches publicly before Demo Day has little predictive value, with a notable exception when companies launch very early in the batch. This probably reflects the fact that companies who get organized early are more likely to be high-Progress companies — which we’re already capturing directly. Once you control for the other signals, launch timing doesn’t tell you much, but the statistical uptick in the performance of companies that launch early is an interesting phenomenon.

Demo Day “heat”

I covered this in detail in On investing in “hot” Y Combinator startups, so I’ll keep it brief here. Demo Day investor enthusiasm has a real but quite weak correlation with eventual outcomes. Given the higher entry valuations that hot deals command, the effect is roughly washed out by the time you account for price. In other words: deal heat is at best a weak positive signal that gets fully priced in by the market. It’s not a useful basis for an allocation decision.

The compounding effect

One last observation, which is arguably the most important methodological finding of all: the signals compound. Companies where two or more of the strong signals fire together are dramatically more likely to succeed than companies where any one fires alone.

We now screen companies using a combinatorial ML/AI scoring system with hundreds of variables precisely because of this. The marginal value of any single signal is real but limited — the value of precisely weighted combinations of signals is much greater.

This is one of the reasons that, despite all the recent hype around AI in venture, only ~1% of funds today employ truly rigorous and backtested data-driven investment processes. It’s hard work, and it’s only valuable if you have the right data and team to train a sophisticated multi-variate model, which took us years and millions of R&D dollars.

Knowing which signals drive YC startup performance is not sufficient on its own. You also need the technical infrastructure to collect vast amounts of data on every YC startup and founder, know how much to weight each variable, and then score and rank them in an objective and scalable way.

What this means

For investors at the YC Demo Day stage, the implications are pretty clear. If you’re not weighting product progress and revenue traction at least as heavily as team pedigree, you’re probably mispricing the cohort. The instinct to “bet on the team” is right at the population level, but at the YC stage, where team variance is already compressed by selection, it’s the product itself and rate of progress that does the real discriminating.

For founders, the message is one I’ve made before but worth restating: the most valuable thing you can do in the weeks before Demo Day is ship visible product progress and get to real paying customers. Both move the statistical needle far more than the things founders typically obsess over — pitch polish, deck design, networking, who’s “hot” this batch, etc.

For us at Rebel, the underlying lesson is that even an extremely sophisticated, data-driven approach to YC startup investing must continually evolve. Every six months or so, we retest our priors against the latest cohort of outcomes data. Some signals get stronger, some weaker, and some may fall out entirely. The whole point of building the world’s most comprehensive YC dataset and ML/AI scoring system trained on it, is that we can keep doing exactly that, indefinitely.

¹ “Series A or later” is the success threshold I use throughout this post, both because it’s a clean and observable milestone that we can assess in relatively recent batches. What ultimately drives returns for a seed-stage investors like us are really the unicorns, but it’s still a useful intermediate filter. While only a minority of the Series A startups ultimately become unicorns, every unicorn raised a Series A at some point.

--

--

Jared Heyman
Jared Heyman

Written by Jared Heyman

Tech guy and investor. Founder of Rebel Fund and previously Pioneer Fund, CrowdMed (YC W13), Infosurv & Intengo (acq. LON: NFC). Ex-Bain consultant. Data nerd.