We track Tellstone's predictions against outcomes the studio reports back over time, whether the predicted risks materialized, whether recommended team changes worked out, whether the team shipped on the predicted timeline. In our most recent manual validation pass across nine teams whose outcomes we tracked qualitatively with the studios involved, the composite's primary risk prediction matched what the team subsequently reported in 8 of 9 cases.
This figure should be read with the appropriate caveats:
- Small sample. Nine teams is a starting point, not a peer-reviewed result. We're explicit about this.
- Why nine and not fourteen. Our dataset includes 14 AI team analyses, but only these nine have had their real-world outcomes confirmed back to us by the studios so far, and the other five are awaiting outcome data. We score the validation only on the teams we can actually check.
- Outcome tracking is being formalized. The earlier validation was manual, confirming outcomes with studio operators in conversation rather than via structured in-product reporting. We've since built outcome-tracking into the product, and the dataset will compound from here as more teams complete the assessment → simulation → outcome cycle.
- Operationalization matters. "Match" means the team confirmed that the flagged primary risk either manifested or didn't, as predicted. It's a specific operationalization of predictive validity, not a universal "accuracy" claim.
- Acting on a prediction changes the outcome. If we flag a risk and the team mitigates it, the risk may never materialize, which is the system working, not failing (the self-defeating-prophecy and Hawthorne effects). So we validate two ways: whether the people who know the team recognize the flagged dynamic as true when we name it, and whether the mitigations we recommend get adopted and work, not only whether the raw outcome occurred.
We'd rather report our sample size honestly than claim a larger validation set than we have. As our data grows with more outcomes, this section gets richer. Several of those tracked cases are below.
What the tracked cases actually look like
The figure above is only as good as the cases underneath it. Here are several of the tracked predictions, confirmed by three different kinds of witness: an investor who wrote back in his own words, an operator who knew the team firsthand, and a founder confirming the call about himself.
Predicted: Before a cofounder pairing at a Paragraph portfolio company, Tellstone flagged a specific dynamic: the commercially adaptable cofounder would suppress disagreements with his independent technical counterpart to keep the peace, letting resentment build instead of surfacing healthy conflict. The report set out ground rules to force early, structured debate.
What happened: eleven days later, Paragraph's managing partner wrote back, unprompted: “The analysis was very helpful and feels directionally accurate. We're working through some of those issues this week.” The predicted dynamic was already live. Paragraph later became an investor in Tellstone.
Predicted: Analyzing a Purdue Innovates accelerator team, Tellstone, given the roles as entered, flagged a “Shadow CEO Standoff”: the listed CEO and COO profiles fit each other's seats, a role overlap that had to resolve before the team could sell.
What happened: the program director confirmed on the spot that the roles were in fact reversed in real life, the model had detected the swap from the founders' profiles alone. “That's actually really cool.”
Predicted: In early testing, a funded tech startup offered to pilot the assessment with its founding team. Tellstone typed them as “Methodical Missionaries” and flagged a Delivery Bottleneck: ambition outrunning capacity, work slipping past deadlines.
What happened: on the debrief call a cofounder validated it immediately: “That's honestly the thing I've been working on for eight years now. I bite off more than I can chew, and I end up deadline slipping. The fact that it pulls that out is good.” The team then asked to add their third partner and re-run the analysis.
Run August 2025 on our earlier instrument (pre-2026 rebuild).
Whether or not a flagged risk ever fires, the mitigation plan is what the studio acts on, and for many teams it is the most useful part of the report. Take the cofounder-tension case above. Alongside the risk, the report gave three concrete, do-it-on-day-one moves:
- a sales-commitment protocol, a pre-approved menu of what the commercial cofounder can promise a customer without sign-off, plus the triggers that require the technical cofounder's approval, so flexibility doesn't quietly break the roadmap;
- monthly founder-friction reviews, a standing session built to surface role and equity tension early rather than waiting for it to surface on its own;
- an equity-and-decision-rights agreement written before the company forms, naming who owns which calls, product scope versus deal terms.
None of that depends on us being right about the risk. A studio can put it in place immediately, and it holds its value even if the tension never materializes. That is the point. You are not buying a prophecy, you are buying a sharper read and a plan you can act on.
We publish these because personality-based prediction invites doubt, not in spite of it. The point isn't applause. It's that the dynamic we flagged is the one the people who knew the team confirmed, sometimes while it was still unfolding. And an independent measurement scientist reviewed the instrument itself.
Predictions → outcomes → calibrated model
The peer-reviewed psychometric and founder-failure literature gives Tellstone a strong starting point. But the most important thing we're building isn't the assessment or the AI layer. It's the compounding outcome dataset sitting underneath them.
Every team that gets assessed and analyzed produces a prediction: a team archetype, a set of risk theses, and a mitigation plan. Every team that operates over the months that follow produces an outcome: the predicted risks either manifested or didn't, the mitigations either worked or didn't, the team either hit its milestone or stalled where we said it might. When those outcomes flow back into the system, through our in-product outcome tracking and direct confirmation with studio operators, they don't just measure accuracy. They calibrate the next prediction.
Over time, this produces a proprietary dataset that no competitor can replicate: real founding teams, real assessment profiles, real predicted risks, real documented outcomes. Every academic source we cite is a population-level average derived from studies of broad worker samples. The dataset we're building is narrower and far more relevant, early-stage founding teams specifically, scored with our own instrument, with outcomes that actually map to what venture studios and accelerators care about.
The practical consequence for the studios and accelerators using Tellstone today: every team you analyze, and every outcome you track, makes the next prediction sharper, not just for the next person, but for your studio specifically. The model increasingly draws on what actually happened at your studio with your portfolio companies, not just generic patterns from the literature.