How Accelerators Actually Track Cohort Progress, and Why Spreadsheets Break at Ten Companies
What the evidence says about tracking startup progress through an accelerator cohort, the metrics real programs collect, and why spreadsheet based tracking degrades as cohorts grow.
Every accelerator program manager has lived the same week. Demo day is fourteen days out, a funder wants a cohort performance summary, and the only source of truth is a spreadsheet where four of the twelve companies last updated their numbers in March.
This is not a discipline problem. It is a structural one, and it becomes visible at a predictable point: roughly the moment a cohort passes ten companies.
What the evidence actually says about accelerator impact
Before discussing how to track progress, it is worth being precise about what accelerators demonstrably do, because that determines what is worth measuring.
The strongest available evidence is a 2025 meta analysis published in the Journal of Technology Transfer, which pooled 21 primary studies and 68 effect sizes. It found a statistically significant positive effect of acceleration on venture performance, with substantial heterogeneity between programs. It also found evidence of publication bias, meaning the literature over represents positive results.
Two of its moderator findings should shape how any program thinks about cohort design. Longer programs produce better outcomes. And larger cohorts reduce effectiveness, which the authors attribute to diluted attention.
Canada has unusually good data here. Innovation, Science and Economic Development Canada published a study in August 2024 combining its Business Accelerator and Incubator Performance Measurement Framework survey data from 2017 to 2020 with Statistics Canada tax filing records from 2014 to 2020. It covered 8,062 firms, of which 7,152 were matched to tax data, an 88.7 percent match rate.
The findings are more sobering than most accelerator marketing suggests:
- Employment was 14 percent higher during the participation year and 13.1 percent the following year, with the advantage diminishing over time.
- Revenue was 13 percent higher during the participation year. The following year, the advantage disappeared and became statistically insignificant.
- 6.1 percent of participating firms qualified as high growth firms under the OECD definition, against 0.4 percent of the comparison group.
The implication for tracking is direct. If the measurable revenue advantage of your program concentrates in the participation year and fades afterward, then annual retrospective reporting will systematically miss the effect you are trying to demonstrate. You need in program measurement, at intervals short enough to capture change while it is happening.
The concentration problem nobody designs for
The Global Accelerator Learning Initiative, a partnership between Emory University and the Aspen Network of Development Entrepreneurs, has assembled data on roughly 23,000 entrepreneurs applying to more than 360 programs across 150 countries.
Its single most useful finding for program operations is this: only 10 percent of ventures account for over 95 percent of total equity investment.
Outcomes in a cohort are not normally distributed. They follow a severe power law. This breaks two things that most programs do by default.
It breaks average based reporting. A cohort mean is dominated by one or two outliers and tells a funder almost nothing about the typical experience. GALI's companion report, A Rocket or a Runway, examined 2,599 ventures across 212 programs and found that growth concentrates among top performers while median ventures show only modest gains. Report medians and distributions, not means.
It also breaks uniform attention allocation. If nearly all measurable value sits in a small subset, but you cannot reliably identify which subset in advance, you need tracking that surfaces divergence early rather than a monthly check in that treats all companies identically. GALI is blunt on the prediction problem, noting that accelerators select fewer than 13 percent of applicants and are "not always successful at predicting which entrepreneurs will succeed."
What a rigorous program actually collects
The best documented tracking model in the published literature belongs to Creative Destruction Lab, described in detail in an August 2025 National Bureau of Economic Research working paper covering 15,791 applications and 3,813 admitted ventures from 2012 to 2024.
CDL runs a full day session every eight weeks. The structure is worth understanding because the measurement is inseparable from the program design.
At each session, ventures set measurable objectives for the following two month interval. Small group meetings of four to six mentors per founder are followed by a moderated large room discussion where objectives are finalised. Ventures that fail to secure formal mentor commitments are dropped from subsequent sessions, meaning selection is continuous rather than a single intake decision.
The data captured every session includes:
- Proposed and finalised objectives, classified by business function
- Chief executive progress reports covering what is working and what is blocked
- Financial updates: cash burn, revenue, runway and employee count
- Attendance records
- Verbatim transcripts of mentor and founder discussions
That financial set of four metrics, collected on a fixed eight week cadence alongside structured objectives, is the most defensible tracking framework published by any real program. It is small enough that founders will actually maintain it and rich enough to show trajectory.
Note what is absent. There is no sprawling metrics dashboard, no thirty field monthly survey. Four financial numbers, a set of objectives, and a qualitative progress report.
Why spreadsheets fail, with actual data
Programs almost universally start in spreadsheets, and the failure is usually attributed to founder discipline. The research suggests the tool itself carries a substantial error rate.
Ray Panko's 2015 review of the spreadsheet error literature found that across 14 laboratory studies involving 967 participants, the average cell error rate was 3.9 percent, consistent with the one to five percent range predicted by general human error research. Across 85 intensively inspected operational spreadsheets, errors were found in 94 percent of them. Field audit cell error rates ranged from 1.2 to 2.5 percent.
Panko adds a caveat worth repeating: reported error rates likely understate the problem, because inspectors report only the errors they found, not those they missed. It should also be noted that several of the 85 audited spreadsheets were inspected by commercial firms with an interest in the finding.
Apply a two percent cell error rate to a cohort tracker. Twelve companies, eight metrics each, updated monthly across a four month program gives roughly 384 data points. At two percent you expect around eight wrong numbers, and you will not know which eight. That is the number you are putting in front of a funder.
The failure is not that spreadsheets are bad software. It is that manual transcription across many contributors accumulates error at a rate that scales with cohort size, and there is no validation layer to catch it.
The reporting burden is documented, and it is real
Program managers who feel that data collection consumes a disproportionate share of their time have institutional evidence on their side.
The National Research Council's evaluation of the Canada Accelerator and Incubator Program, a $92,990,612 program that funded 16 Canadian accelerators and incubators between 2014 and 2019, found that evaluators encountered "reluctance by some A/Is to fully participate in the data collection, in part due to a perceived administrative burden." The evaluation also cited a "complex and time consuming claim review process" and "difficulty in collecting performance data from CAIP recipients."
Its recommendation was that reporting requirements should be clearly specified before a contribution agreement is signed, noting there had been insufficient time to understand the program and develop efficient processes.
This is the clearest independent documentation that accelerator data collection imposes genuine operational cost and degrades data quality when it is bolted on after the fact.
It is worth being honest about a gap here. There is no published research quantifying the burden on the founder side, on survey fatigue or repeated reporting requests. Vendor claims about time wasted on data cleaning are marketing without methodology. What is documented is the burden on programs, not on founders.
Three failure modes worth naming
Stale founder self reporting. Every metric in most trackers is self reported by the founder, on a cadence set by the program, into a form disconnected from anything the founder uses for their own decisions. Reporting is therefore pure overhead for them and it degrades accordingly. CDL's design avoids this by making the objectives the founder sets the same objectives the program tracks. The reporting is the program, not an addition to it.
No cross cohort comparability. Programs change their metrics between cohorts, usually for good reasons. The result is that year three cannot be compared with year one, precisely when a funder asks whether the program is improving. ISED's Performance Measurement Framework exists specifically to address this, with the stated aim of "standardising and improving the quality of data reported."
The pre demo day scramble. Because tracking is retrospective, the reporting deadline creates a compressed period of chasing founders for numbers they should have supplied gradually. The cost is not just the manager's time. It is that the data arrives too late to act on.
What to measure, practically
Based on what real programs collect and what funders in Canada actually assess:
Core financial set, every four to eight weeks. Cash burn, revenue, runway, headcount. Four numbers.
Structured objectives per interval. What the company committed to and whether it was achieved. This carries more signal than any metric, because it measures execution rather than outcome, and outcomes lag beyond program duration.
The three outcome measures that appear in nearly every framework. Revenue, employees, and capital raised. GALI collects these across 360 programs, which makes them the de facto standard and gives you external comparability.
Funder aligned indicators. In Canada, ISED's framework focuses on job creation, startup growth, and entrepreneurship among under represented groups. If you receive public funding, collecting these from day one avoids reconstructing them later.
What to skip. Vanity metrics that flatter the program without informing anyone: cumulative mentor hours, session attendance as a headline, social media reach, aggregate portfolio valuation. These correlate with nothing in the outcomes literature.
The distinction that matters most is between metrics that prove programme return on investment to your funders and metrics that genuinely help your founders. Most programs conflate the two and end up collecting a large volume of data that serves neither.
Frequently asked questions
How often should an accelerator collect metrics from its cohort? Creative Destruction Lab operates on an eight week cycle, collecting financials and objectives at each session. For a typical three to four month program, a four to eight week cadence captures change while it is happening. The ISED study found revenue advantages that were significant during participation and insignificant the following year, which argues strongly against relying on annual retrospective collection.
What metrics do accelerators report to funders? In Canada, ISED's Business Accelerator and Incubator Performance Measurement Framework focuses on job creation, startup growth, and entrepreneurship among under represented groups. Internationally, revenue, employee count, and capital raised are the near universal core, used by GALI across more than 360 programs.
At what cohort size does spreadsheet tracking stop working? There is no research establishing a precise threshold. What the spreadsheet error literature shows is that error accumulates with the number of manually entered data points, which scales with cohort size multiplied by metrics multiplied by reporting periods. The practical breaking point most programs report is around ten to twelve companies, at which the manual chasing and reconciliation exceeds the time available.
Do accelerators actually improve startup outcomes? The 2025 meta analysis in the Journal of Technology Transfer found a statistically significant positive effect with substantial variation between programs, alongside evidence of publication bias. The regression discontinuity study of Start Up Chile, published in the Review of Financial Studies in 2018, found that basic services alone, meaning cash and coworking space, had no measurable effect, while entrepreneurship schooling combined with those services significantly increased performance. Program design matters more than program participation.
Sources
- Seitz, Buratti, Lehmann and Kurrle, A meta analysis towards the effectiveness of startup accelerators, Journal of Technology Transfer, 2025
- ISED Canada, The effect of Business Accelerators and Incubators on business performance, 2024
- ISED Canada, Business Accelerator and Incubator Performance Measurement Framework
- GALI, Does Acceleration Work? Five Years of Evidence
- GALI, A Rocket or a Runway?
- Sariri et al., NBER Working Paper 34127, Creative Destruction Lab
- Panko, What We Don't Know About Spreadsheet Errors Today
- NRC, Evaluation of the Canada Accelerator and Incubator Program
- Gonzalez-Uribe and Leatherbee, Review of Financial Studies, 2018