Most personality assessments use Likert scales, the "rate yourself 1-7 on this trait" format. It has a known structural problem. Everyone can rate themselves high on everything. Social-desirability bias and self-deceptive enhancement (Paulhus's BIDR framework) inflate self-reports, especially on desirable traits like "leadership" and "collaboration."
Tellstone uses forced-choice paired-trait questions instead. Each question presents four options across two opposing trait poles (e.g., "introversion" vs. "extraversion"). The respondent can't pick "I'm high on both." They have to trade off.
This isn't just a stylistic choice. The peer-reviewed evidence on forced-choice formats is strong: a comprehensive meta-analysis by Salgado and Táuriz found that quasi-ipsative forced-choice personality inventories produced higher operational validity for predicting job performance than single-stimulus Likert formats across nine occupational groups (Salgado & Táuriz, 2014, European Journal of Work & Organizational Psychology). Subsequent meta-analytic work confirms forced-choice formats are also more resistant to deliberate faking (Cao & Drasgow, 2019). This is the same methodological family that Hogan Forced-Choice and several Big Five forced-choice variants use. In contexts where social-desirability pressure is high, and founder assessments shown to investors and studios are exactly that, it's the methodologically sounder choice.
We measure five personality dichotomies and ten soft skills. The five dichotomies are:
- Introversion vs. Extraversion, outward energy preference
- Structured vs. Flexible, process orientation
- Collaborative vs. Independent, how someone works with others
- Creative vs. Methodical, whether someone reaches for a novel approach or a proven one
- Analytical vs. Intuitive, whether someone decides by working through the logic or by reading the pattern
Alongside them we measure ten soft skills: leadership, teamwork, adaptability, critical thinking, empathy, self-awareness, time management, verbal communication, written communication, and conflict resolution.
Four of these overlap substantially with the Big Five (the OCEAN model, sometimes called Norman's Big Five), the academic gold standard in personality assessment, with internal-consistency alphas of .73 to .80 (recent reliability meta-analysis, 2025), without inheriting the Likert format's bias. Introversion vs. Extraversion maps to Extraversion, Structured vs. Flexible to Conscientiousness, Collaborative vs. Independent to Agreeableness, and Creative vs. Methodical to Openness, the novel-versus-proven axis at the heart of that trait. Mapping to the Big Five also lets these four dimensions inherit its best-documented property, rank-order stability over time (Roberts & DelVecchio, 2000), the consistency that sets the Big Five apart from type-based tools like the Myers-Briggs.
The overlap is deliberate but not complete, and that is a design choice for a workplace tool. We measure four of the five Big Five domains and replace the fifth, emotional stability (neuroticism), with a decision-style axis, Analytical vs. Intuitive, drawn from the dual-process tradition of System 1 and System 2 thinking (Kahneman, 2011). The best-known workplace instrument in history made the same trade. The Myers-Briggs maps to four Big Five domains and substitutes information and decision style for neuroticism (McCrae & Costa, 1989). We make that trade for two reasons. First, how a founder thinks and decides is more useful for composing a team than how anxious they are. Second, the evidence. Across the meta-analytic record, emotional stability is the weakest and least consistent Big Five predictor of job performance, roughly .09 to .13 against about .19 to .23 for conscientiousness (Zell & Lesick, 2022; Sackett et al., 2022). Its stronger associations are with outcomes we do not claim to predict, like turnover and strain, and it is also the trait most prone to feeling stigmatizing in a team-diligence context.
Measures anchored in published rubrics
The ten soft skills are measured differently from the personality dichotomies, and this is the part people most often misread, so to be clear, it is not a Likert self-rating. Nobody is asked to score themselves from 1 to 5 on "leadership." Instead each skill is assessed through short, scenario-based items, realistic situations where the respondent picks how they would actually handle it. The options for every item form a proficiency ladder, a progression from a lower-skill to a higher-skill response, and the options are scrambled so the higher-skill answer is not guessable by position. The ordering is anchored to the published competency literature for that specific skill, not to our own opinion of what "good" looks like, so a respondent cannot simply pick the "5."
Those ladders draw on the academic frameworks that define each construct:
- Communication (verbal & written), the AAC&U Oral and Written Communication VALUE rubrics, and reader- versus writer-based prose (Flower & Hayes)
- Critical thinking, the Facione/Delphi consensus framework and Toulmin's model of argument
- Leadership, House's path-goal theory (direction-setting under ambiguity) and the ownership/accountability (internal locus of control) literature
- Teamwork, Salas, Sims & Burke's "Big Five" of teamwork
- Adaptability, Pulakos et al.'s adaptive-performance taxonomy (reactive to proactive to anticipatory)
- Empathy, Davis's Interpersonal Reactivity Index (perspective-taking, empathic concern) and the cognitive/affective empathy distinction
- Self-awareness, Eurich's internal/external distinction and Vazire's self-other knowledge asymmetry
- Time management, the Eisenhower/Covey importance-over-urgency distinction
- Conflict resolution, the Thomas-Kilmann modes and Jehn's task/relationship conflict research
Two principles carry over from the forced-choice philosophy above. No option is written to be the obvious "right" answer. Each is a genuine, defensible way a real person at that level would describe themselves, so the good-to-better ordering lives in the underlying behavior, not in wording a test-taker can pattern-match. And each item is balanced across the two skills it touches, so a respondent can't score high on one by accident of the other. Both choices exist for the same reason as the forced-choice format, to measure how someone actually works, not how they'd like to be seen.
Audited for construct validity, and built to learn from real data
An instrument is only as trustworthy as its weakest item, so we audit every one of them against a written definition of what its construct measures and, just as importantly, what it must not reward. In a 2026 rebuild, each of the ten soft skills was re-derived against its published proficiency ladder, and each of the five personality dichotomies against a defined pole anchored to the Big Five. Items were stress-tested for the failure modes that quietly undermine self-report scales: options that are effectively duplicates, a lower-skill behavior accidentally scored as higher (an "inversion"), wording that flatters or repels regardless of the respondent's true level, and paired-skill items where one skill quietly rides the other's rung. Anything that didn't clear the bar was rewritten and re-checked against the literature.
We then attacked the items the way a skeptic, or a test-taker trying to game them, would. For each personality item we built the cross-profile (someone high on the opposite trait) and asked whether they'd pick an option for the wrong reason; where they could, we reworded it. For the paired soft-skill items we ran the same attack from both lopsided profiles (strong on one skill, weak on its partner) so no rung can be claimed on one skill alone. We balanced the on-screen option order so the "best" answer is never learnable by position, kept option lengths even so length is no tell, and engineered desirability deliberately. Low options are written as dignified, genuine self-descriptions, and several top options openly admit a flaw, so neither the "right" answer nor a faked one is obvious.
Finally, we log each response at the item level (de-identified), which lets us watch how every individual question behaves across the population, whether an option is never chosen, whether scores pile at the ceiling, whether each option attracts the people who genuinely score that way. Instrument improvement becomes surgical: when a question underperforms, we can see it and fix that one item without re-testing anyone. We'll publish the rebuilt instrument's distributions as responses accumulate.