HomeScienceHonest Limits
Honest Limits

What self-report can and can't do

The companies you should be skeptical of are the ones that won't tell you where their method stops. Here's where ours does.

Every self-administered personality assessment, including Tellstone's, has the same structural limitation: it measures how the test-taker describes themselves, not how others observe them. The peer-reviewed literature is clear on this.

Connelly & Ones' 2010 meta-analysis of 263 studies (N = 44,178) found that self-other agreement on the Big Five averages r ≈ .38 for casual acquaintances and r ≈ .46 for well-acquainted dyads. Even teammates who have worked together for years only agree with self-reports at moderate correlations. This is a structural feature of every self-report instrument (DISC, MBTI, Hogan, Big Five, Enneagram, all of them), not a defect of any specific tool.

Simine Vazire's Self-Other Knowledge Asymmetry (SOKA) model explains why: people are the best judges of their internal traits (anxiety, neuroticism), but observers are better judges of evaluative, visible traits (communication style, conscientiousness). Practically, this means a self-report assessment can describe how someone sees themselves accurately, while their teammates describe them differently, and both are true.

This dynamic is more pronounced for the soft-skill measures than for the personality dichotomies. Communication, leadership, teamwork, and self-awareness sit precisely in the externally-evaluated category SOKA predicts will diverge most between self and observer. Our soft-skill distributions sit close to the practical midpoint with healthy variance, which suggests the instrument itself is well-calibrated, but a colleague reading the report may still see a teammate differently than that teammate sees themselves. That divergence is the predicted phenomenon, not an instrument failure. Tellstone treats every soft-skill score as one input into a multi-cue composite, never as a standalone behavioral verdict about an individual.

What this means for our product

Tellstone's predictive validity comes from the composite (work history, structural risk flags, and personality signals together), not from any single self-report being perfect. The composite is where that predictive signal lives; the individual personality typing is one signal among many.

Where we choose our words carefully: Tellstone's reports describe how a founder reports themselves, not how they will be perceived by others. These are related but distinct constructs.

Who This Is For

Built for team diligence, not executive coaching

Tellstone is most useful when you don't yet have rich lived-experience data on a team. Venture studios placing founders, accelerators evaluating cohorts, investors doing seed-stage diligence, founders building a team from scratch. These are the contexts where a measurement layer adds real value, because there's no decade of working-together history to draw on.

Tellstone is less useful for established teams that have been together for many years. At that point, the team already has its own ground-truth observations of each other, which any self-report assessment will lose to in a head-to-head comparison. For those teams, 360-degree feedback, executive coaching, or facilitated team-dynamics work is the right intervention, not a self-administered assessment. We're explicit about this because being honest about who a product is for is part of doing the work well.

Boundaries

What Tellstone is not

  • Not a substitute for human judgment. A studio operator with 20 years of pattern-matching brings something our model can't replicate. Tellstone supports that judgment with structured measurement; it doesn't replace it.
  • Not a personality "type" labeling system. We measure continuous trait dimensions. Archetype labels are interpretive overlays for narrative clarity, not categorical destinies.
  • Not a hiring tool. We're built for team-fit prediction, not candidate selection. Different methodology, different validation, different ethical constraints.
  • Not 360-degree feedback. Different methodology designed for a different use case (established teams with observation data on each other).
  • Not predictive at the individual level for performance. Our predictions are about team-level outcomes (risk manifestation, goal achievement), not whether a specific founder will succeed.
When We Get Doubted

A deal we lost, and the limit it marks

A venture studio asked us to analyze a two-founder team in the middle of a raise. From the psychometrics and work histories alone, before we'd sat in on a single one of their meetings, the report named the pair “synchronized”: two highly extraverted, collaborative founders who present a single, unified front to investors. That alignment was their headline strength. The named risk was that the same alignment could become an echo chamber, scoring low on the analytical/critical-thinking end and running on tight consensus, the pair might leave key assumptions unstress-tested, so that under hard investor scrutiny their agreement could prove “surface-level consensus rather than rigorous validation.”

The company's CEO didn't buy it. He'd taken Myers-Briggs, the Enneagram, “a lot of these,” and this read like more of the same: it's “like fortune telling… or horoscopes,” he said. He disputed the sharpest inference, that the team's alignment would cost them financial discipline: “I very well stress tested the financials.” He had every reason to trust his own read of his team over a report built from a questionnaire, and on the financials he may simply be right. We were missing context on the resources around them, and we said so on the call. No model gets every read right, and this may be one it didn't.

Why we keep this one

The team wanted more proven accuracy than a young instrument can promise, and the friction the conversation created was a real cost. It marks the limit we most want a reader to see: a self-report instrument this young can inform a decision, but it can't overrule how a capable founder reads his own team, and it hasn't yet earned the track record that would justify asking it to. That is why a score is never the verdict here. It is one input among many, held lightly, and worth only as much as the judgment of the person reading it.

Bring measurement to the team decisions you're already making

See how Tellstone fits into your studio's diligence workflow. Free tier, no credit card.

Start Free