HomeScienceFAQ
Science FAQ

How Tellstone works, in plain answers

The questions prospects, investors, and skeptics actually ask about the methodology.

How does Tellstone predict team outcomes?

Tellstone combines several partially-independent signals into a single composite: a forced-choice psychometric assessment (personality plus soft skills), work-history analysis, deterministic structural risk flags drawn from startup-failure research, and an AI synthesis layer. No single signal predicts team success well, but more than 85 years of personnel-selection research shows that combining partially-independent signals produces validity that exceeds any one of them on its own.

Why does Tellstone use forced-choice questions instead of rating scales?

Rating (Likert) scales let everyone rate themselves high on everything, which inflates self-reports through social-desirability bias. Forced-choice paired-trait questions make respondents trade off between options, so the format is more resistant to faking. Meta-analytic research (Salgado and Tauriz, 2014; Cao and Drasgow, 2019) finds forced-choice formats yield higher operational validity and more faking resistance than Likert scales.

What does Tellstone actually measure?

Five personality dichotomies (introversion versus extraversion, analytical versus intuitive, structured versus flexible, collaborative versus independent, and creative versus methodical), plus ten soft skills: leadership, teamwork, adaptability, critical thinking, empathy, self-awareness, time management, verbal communication, written communication, and conflict resolution.

Is Tellstone based on the Big Five?

Four of the five personality dimensions overlap substantially with the Big Five, also called OCEAN or Norman's Big Five, the academic gold-standard model. Introversion versus extraversion maps to extraversion, structured versus flexible to conscientiousness, collaborative versus independent to agreeableness, and creative versus methodical to openness. We deliberately replace the fifth domain, emotional stability (neuroticism), with a decision-style axis, analytical versus intuitive, because how a founder thinks and decides is more useful for composing a team, and because emotional stability is the weakest and least consistent Big Five predictor of job performance. The Myers-Briggs made the same trade decades ago. The ten soft skills are grounded in their own published competency rubrics rather than in the Big Five.

Has Tellstone's assessment been independently reviewed?

Yes. An independent biostatistics professor, a specialist in measurement and validity, took the full assessment and reviewed it. He concluded that it has face validity: it overlaps substantially with established personality science while also going its own way in directions that make practical sense.

How accurate are Tellstone's predictions?

We deliberately don't lead with a single accuracy number. In our most recent manual validation across nine teams whose outcomes we tracked with the studios, the composite's primary risk prediction matched what the team later reported in 8 of 9 cases. A "match" is a specific operationalization (the flagged risk either showed up or didn't, as predicted), not a universal accuracy claim, and it's a small, honestly-reported sample, not a peer-reviewed result. Acting on a prediction also changes the outcome, because a risk you mitigate may never materialize, which is the tool working. So the value doesn't hinge on being right every time. It comes from a sharper, structured read on the team and a concrete mitigation plan, with outcome tracking now built into the product so the dataset compounds.

What are the limits of a self-report assessment like Tellstone?

Every self-report instrument measures how someone describes themselves, not how others observe them. Self-other agreement on personality averages roughly r = .38 to .46 (Connelly and Ones, 2010). This is a structural feature of all self-report tools, including DISC, MBTI, Hogan, the Big Five, and the Enneagram. Tellstone treats each self-report score as one input into a multi-cue composite, never as a standalone verdict about an individual.

Who is Tellstone for, and who is it not for?

Tellstone is most useful when you do not yet have rich lived-experience data on a team: venture studios placing founders, accelerators evaluating cohorts, investors doing seed-stage diligence, and founders building a team from scratch. It is less useful for long-established teams that already have ground-truth observations of each other; for them, 360-degree feedback or facilitated team-dynamics work is more appropriate.

How is Tellstone different from DISC, MBTI, or the Predictive Index?

Most assessments describe an individual. DISC, the Predictive Index, and Gallup CliftonStrengths profile a person, and Myers-Briggs sorts people into types that are unstable on retest, which is why its own publisher says not to use it for hiring. Tellstone differs on the level of analysis. It's built to support a team decision, pairing cofounders, composing a team, or evaluating a founding scenario, by combining psychometrics with work history, structural risk flags, and team-level analysis. It uses a forced-choice format for faking resistance, frames items around real situations, and tracks its predictions against real outcomes so the model improves over time. We're not trying to out-measure Hogan at the individual level, which is its strength. We operate at the team level, which the others don't natively address.

Can't people just game the assessment?

It's harder than it looks. The format is forced-choice, so you can't rate yourself high on everything, since every question makes you trade one trait or skill against another. The on-screen option order is randomized, so position gives no clue to the best answer, and options are kept similar in length so length is no tell. We also write the lower options as dignified, genuine self-descriptions, and several of the top options openly admit a flaw, so the right answer isn't obvious to pick. Peer-reviewed work (Salgado and Tauriz, 2014; Cao and Drasgow, 2019) finds forced-choice formats are meaningfully more resistant to faking than rating scales for exactly these reasons.

How do you know the questions are right before you have data?

Two ways. First, every item was audited against a written definition of what its construct measures and excludes, including adversarial checks: we ask whether someone high on the opposite trait, or strong on only one of a paired skill, could pick an option for the wrong reason, and rewrite it if so. Second, we have pre-registered what we expect the data to show, for example score means drifting toward the scale midpoint, before collecting it, so the results will either confirm those predictions or tell us exactly what to fix. Design gets us a defensible starting point; the data is what proves it.

Want the full detail behind any of these? Start with the methodology overview or browse the references.

Bring measurement to the team decisions you're already making

See how Tellstone fits into your studio's diligence workflow. Free tier, no credit card.

Start Free