A multi-cue composite, not a single magic score
No single signal predicts team success well. Personality assessments alone correlate with job performance at roughly r = .20-.30 (Barrick & Mount, 1991); even cognitive ability tops out at r ≈ .51 (Schmidt & Hunter, 1998). Any platform claiming to predict team outcomes from one input is overselling.
What does work is combining independent signals. Schmidt & Hunter's foundational meta-analysis showed that adding a structured interview to general mental ability lifts composite validity to r = .63; adding an integrity test lifts it to r = .78. The math is straightforward. When signals are partially independent, their errors partially cancel when combined, and the composite predicts performance better than any single component could, capturing more of the true signal and less method-specific noise (Campbell & Fiske, 1959).
Tellstone is built on this principle. We combine:
- A forced-choice psychometric assessment, 10 personality dimensions (5 dichotomies) plus 10 soft-skill dimensions
- Work-history analysis, functional background, leadership tier, industry exposure, stage experience
- Deterministic structural risk flags, drawn from Wasserman, CB Insights, and Startup Genome failure-pattern research
- An AI synthesis layer, pulling the signals together into a team archetype, risk thesis, and mitigation recommendations
- A compounding outcome dataset, predicted risks are tracked against what actually happened, so every outcome calibrates the next prediction
Each component is modest on its own. The composite is what matters, and the dataset that powers the composite gets sharper over time.
85+ years of selection research
Tellstone's prediction model combines six categories of input: personality, soft skills, functional background, leadership tier, industry footprint, and structural risk flags. Each carries only modest predictive validity on its own (in the research literature, single predictors of this kind typically correlate r = .25-.45 with team outcomes). Combined, they reach the composite-validity ranges established by 85+ years of selection research.
"When predictors are combined optimally, composite validity routinely exceeds any single predictor. General mental ability alone has validity r ≈ .51 for job performance; combined with a structured interview, the composite reaches mean validity r ≈ .76; combined with an integrity test, r ≈ .78."
This isn't a guess about how Tellstone works. It's the empirical pattern across a century of personnel-selection research (Schmidt, Oh & Shaffer, 2016; Schmidt & Hunter, 1998). A more recent recalibration by Sackett and colleagues, which corrected for systematic over-correction of range restriction in prior meta-analyses, reduced most validity estimates by .10-.20, though the rank order of predictors largely held (Sackett et al., 2022). When signals carry independent variance, their errors partially cancel, and the composite captures variance no single signal could (Campbell & Fiske, 1959, multitrait-multimethod matrix).
Team-level evidence, modest but real
The most-cited team-level Big Five meta-analysis is Peeters et al. (2006), which reported that team-mean Conscientiousness (ρ = .20) and Agreeableness (ρ = .24) predicted team performance, with variability in these traits negatively related to performance (Peeters, Van Tuijl, Rutte & Reymen, 2006). A 2024 update by Han, Krieger, Kim, Nixon and Greiff incorporates newer studies and finds weaker effects (r = -.05 to .13) than the 2006 figures, though several moderators (team type, task interdependence) matter (Han et al., 2024). Bell (2007) found that deep-level composition variables predict team performance more strongly in field than in lab settings (Bell, 2007). The reasonable read: team-personality effects are real but modest on their own, which is exactly why composites that add work history and structural risk flags outperform team-personality alone.
Layered on top of the composite, Tellstone applies deterministic structural risk flags, rule-based checks that fire when specific patterns appear (e.g., no vesting agreement; solo non-technical founder building a product). Paul Meehl's foundational Clinical vs. Statistical Prediction (1954) and the Grove et al. 2000 meta-analysis of 136 studies showed that mechanical rules-based prediction matches or beats expert clinical judgment in roughly 90% of comparisons, and averages ~10% more accurate. These rules catch dealbreakers that continuous models miss.
Descriptive accuracy vs. predictive validity
A subtle but important distinction in psychometric methodology: describing a person accurately and predicting an outcome accurately are different criteria, and a tool can succeed at one while being imperfect at the other (Campbell & Fiske, 1959; Sackett & Lievens, 2008).
A founder reading their Tellstone profile may say, "the structured/flexible axis isn't quite right. I'm more flexible than this suggests." That's a descriptive critique. It's worth listening to, and it's why we show the assessment back to the team for review. But the question that matters for the studio decision is different: does the composite predict how this team is likely to perform together? Those are separately validated criteria, and the second is the one our 8-of-9 validation figure addresses.