How this test is built
Every number on this page comes from public data you can download and recompute yourself. If you find an error in it, we want to hear about it.
The model
The test measures the Big Five (also called the Five-Factor Model): five broad trait dimensions that were not invented by us or by anyone else as a theory. They were found by taking the vocabulary people use to describe personality and looking at which descriptions move together. The same five keep reappearing across languages, across cultures, and in ratings made by other people rather than by yourself.
The five are scales, not types. You are not one of sixteen boxes; you are somewhere on each of five continuous dimensions, and about half of everyone sits near the middle of any given one. That is the main reason we report percentiles instead of letters, a person at the 51st percentile and a person at the 49th are the same person, but a letter-based test would give them opposite labels.
The questions
All 120 statements come from the International Personality Item Pool (IPIP), a public-domain item set built by academic researchers for research use. Public domain means anyone may copy, translate or use the items for any purpose, including commercial products, without permission or fees. We did not write the items and we did not reword them: rewording would break the link to the published psychometrics below.
Specifically we use the IPIP-NEO-120 as described by Johnson (2014), which measures the five domains through 30 narrower facets. The free test you take is a 60-item short form of it, two items per facet.
Does the short form still work?
We chose the 60items on one random half of Johnson's sample and then measured how well they perform on the other half - 202,473 people the selection never saw. That split matters: measuring quality on the same data you used to pick the items would flatter the result.
| Scale | Agreement with the full 120-item inventory | Internal consistency (α) |
|---|---|---|
| Negative Emotionality | r = 0.971 | 0.817 |
| Extraversion | r = 0.966 | 0.783 |
| Openness | r = 0.940 | 0.695 |
| Agreeableness | r = 0.951 | 0.731 |
| Conscientiousness | r = 0.969 | 0.834 |
| Average | r = 0.959 | 0.772 |
Openness is the weakest of the five (α = 0.695), which is a known property of that domain: it covers more varied ground than the others. Read your Openness score with a little more slack than the rest.
Does our scoring actually match the published instrument?
A fair question, since anyone can claim to use a research inventory. So we ran our own scoring code over Johnson's raw data - 619,150 respondents, and recomputed the 30 facet reliabilities he published.
In other words we are not loosely inspired by the instrument, we reproduce its psychometrics. The scripts that do this live in the project repository and run in one command.
Where your percentile comes from
Your raw score means nothing on its own, so it is compared against real people: the 404,041respondents from Johnson's dataset who answered every item, reported their sex, and were between 14 and 80 (raw data on OSF). Personality norms differ by sex and shift with age, so the comparison is made within your own group, not against everybody at once.
| Comparison group | People in it |
|---|---|
| Male, 14-17 | 24,390 |
| Male, 18-21 | 56,542 |
| Male, 22-29 | 45,881 |
| Male, 30-39 | 22,464 |
| Male, 40-49 | 10,579 |
| Male, 50-80 | 5,956 |
| Female, 14-17 | 39,863 |
| Female, 18-21 | 83,865 |
| Female, 22-29 | 59,066 |
| Female, 30-39 | 30,293 |
| Female, 40-49 | 16,856 |
| Female, 50-80 | 8,286 |
If you skip the sex or age question, you are compared against the whole sample instead, and your reading says so. A percentile of 74 means roughly 74 out of 100 people in your comparison group described themselves as lower on that trait than you did. It is not a score out of 100 and it is not a grade.
The archetype name
The measurement is the five percentiles. The archetype. The Architect, The Seeker, and so on, is our shorthand for which trait sits furthest from the middle in your profile, and in which direction. There are ten such names because there are five traits and two directions.
It is a label we invented to make the numbers easy to talk about. It is not a category found in nature, it is not a type, and two people with the same archetype can differ a great deal. When we say "11.27% of people" or "1.1% share your combination", those percentages are counted in the reference sample above, not invented for effect.
What the work section is based on
Traits are not abilities, and this is not an aptitude or placement test. Two of the five traits have a well-measured relationship with occupational interests, from meta-analysis (Larson, Rottinghaus & Borgen, 2002): Openness correlates .48 with artistic and .28 with investigative interests; Extraversion correlates .41 with enterprising and .31 with social interests.
For the other three the link to fields is weak, so your reading talks about working conditions instead, deadlines, autonomy, conflict load, how feedback is delivered. The one strong finding there: Barrick & Mount (1991) found Conscientiousness predicted job performance across every occupational group they studied, at about .20–.22. A correlation that size is a tendency across thousands of people. It cannot tell any individual what to do for a living, and we do not pretend otherwise.
What the growth section is based on
Traits can change. Roberts et al. (2017) reviewed 207 intervention studies and found an average shift of d = .37 over about 24 weeks, with emotional stability moving most. And Hudson & Fraley (2015) found that people who set out to change a trait did shift both their self-reports and their daily behaviour over 16 weeks, but the version of the intervention that worked taught them to write if-then plans; simply resolving to be different did not.
That is why every recommendation in your reading has the form "if X happens, then I will do Y", and why we tell you to expect a few percentile points over months rather than a new personality. Behaviour moves first. The trait follows slowly, if at all.
Limitations, stated plainly
- •The reference sample is 404,041 internet volunteers, not a census. About 72% are from the United States and the median age is 22. Read your percentile as "relative to this sample", not "relative to humanity".
- •The data were collected mostly in the 2000s. Population norms drift slowly, but they do drift.
- •This is a self-report questionnaire. It measures how you described yourself in one sitting, not how you behave, and not what other people see.
- •A bad week can move a score by a few points. If a number surprises you, retake it in a month.
- •The free form has two items per facet, which is enough for the five broad scales but not for facet-level detail, so we do not report facets from it.
- •We have not measured how stable our own scores are over time (test-retest) and do not claim a figure. We plan to publish one once enough people have retaken the test.
- •Nothing here is a clinical or diagnostic instrument. It cannot detect a disorder, and it must not be used to make decisions about another person, including hiring.
What we do not claim
We do not claim the test predicts your future, your income, your relationships or your success. We do not claim it was validated or endorsed by any university or professional body. We do not claim an accuracy percentage, anyone who gives you one from a 60-question questionnaire is selling something. And we are not the Myers-Briggs instrument, nor a version of it: this measures traits on continuous scales, which is a different thing.
Sources
- Items: International Personality Item Pool, public domain.
- Instrument and norms: Johnson, J. A. (2014). Measuring thirty facets of the Five Factor Model with a 120-item public domain inventory: Development of the IPIP-NEO-120. Journal of Research in Personality, 51, 78–89. Raw data: osf.io/wxvth.
- Interests: Larson, Rottinghaus & Borgen (2002). Journal of Vocational Behavior, 61, 217–239.
- Job performance: Barrick & Mount (1991). Personnel Psychology, 44, 1–26.
- Trait change: Roberts et al. (2017), Psychological Bulletin; Hudson & Fraley (2015), Journal of Personality and Social Psychology.
That is the whole method. If something here does not hold up, the test is wrong, not you.
Take the test →