HomeThe science

A thousand simulations before anything is built.

AI has changed how an assessment can be designed. It has not changed how one has to be proven. This page shows what we built with it, a behavioral framework, a personality instrument, a judgment studio and a reasoning test, and the evidence behind each one.

I trained as an aerospace engineer in the nineties, and I arrived just as something changed. Computers got powerful enough to simulate the physical world, and design stopped being build-it-and-see. You could try a thousand wings before you cut any metal.

The same shift is now happening in psychology, in the part of the work that used to set the pace: writing the items. We generate thousands of candidate items. Each one is reviewed and either edited or removed: some are off target, some are pitched at too high a reading level, some are worded in a way that could disadvantage a group of candidates. Then Pseudo Factor Analysis runs across the whole pool. It is a structural read of the items themselves, done before a single person has answered one, and it tells us which items are likely to behave. Only those go forward to trial.

That work used to take months. It now takes a day.

What we do not skip is people. Every instrument is trialed on real candidates, without exception. The simulation finds the design; the trial is what proves it, and that order never changes.

The result is not a cheaper test. It is a test built for the job.

— Tariq Shaban, Co-Founder & Product & Innovation Director

The method

One method, four times over.

The same four steps built the behavioral framework, the personality instrument, the judgment studio and the reasoning test. The first three take a day. The fourth takes as long as it needs.

  1. Step 1 · one day

    Generate

    Thousands of candidate items, written against tight design briefs and anchored to specific behaviors.

  2. Step 2 · one day

    Review and edit

    Every item is checked for target, reading level and wording that could disadvantage a group. Items come back edited, or they come out.

  3. Step 3 · one day

    Pseudo Factor Analysis

    A structural read of the items themselves, showing which ones are likely to behave before anyone has answered one.

  4. Step 4 · as long as it takes

    Trial on people

    Every instrument, without exception. The simulation finds the design; the trial proves it.

Every step is checked against the real thing.

Nothing in the pipeline is taken on trust. Subject-matter experts' judgments are checked for agreement before they become a scoring key, and every stage a model touches is checked against human data on the same content. The chart shows the example we publish: we ran Pseudo Factor Analysis on a set of items, ran a conventional factor analysis on real responses to the same items, and compared what each one recovered. Pseudo Factor Analysis gets most of the way there on its own, which is all it needs to do. Its job is to decide which items go to trial, not to replace the trial.

Pseudo Factor Analysis compared against conventional factor analysis on the same content
Pseudo Factor Analysis compared with a conventional factor analysis of the same items.

01 · The behavioral framework

Before you can measure behavior, someone has to name it.

Every assessment predicts something, and that something is usually vague. In most organizations, “performance” ends up meaning one manager's overall rating, which says little about what the person actually did.

Competency models were meant to solve that, but most are written in a two-day workshop. Ours was built from published evidence. We collected 898 behavior statements from 46 published sources, covering four decades of research taxonomies and the competency dictionaries in professional use, and let the structure emerge from that material rather than from a room of people agreeing with each other.

We then applied the same method we use for items: generate, review, Pseudo Factor Analysis, then expert adjudication. The result is six clusters, eighteen competencies and seventy-seven observable behaviors, each one traceable to the source it came from.

898behavior statements
46published sources
6·18·77the structure that resulted
Behavior networkDrag · zoom · click a node

All seventy-seven behaviors, placed by how closely they relate to one another. The six clusters were not assigned afterwards; they are where the behaviors grouped themselves. Drag, zoom, or click any node to see the behavior and its source.

An organization that cannot name the behaviors it values cannot select for them, develop them, or defend its decisions about them.

02 · Personality

Built in weeks. Documented like a decade.

Six factors, twelve traits, twenty-four facets and forty-five scales, in about twenty minutes. The instrument was calibrated on 662 working adults and has since been taken by more than fifteen thousand applicants. Every scale meets the reliability standard, and the structure we designed is the structure that shows up in the data. The technical manual runs to two hundred pages and is published in full, which very few publishers do.

The illustrated item format
Two formats, one test. The instrument runs as text or as illustrated scenes, so reading load never decides who can take it. None of the 96 items behaves differently between formats, and the largest score difference on any factor is negligible. One norm table, one score, and the candidate chooses the format.
What happens when a language model takes the assessment
What happens when an AI takes it. We pointed a language model at our own instrument and told it to look like the ideal candidate. It raised itself on the traits the role rewards and paid for them on other traits, because the forced-choice format makes it choose. Knowing what that trade-off looks like is how you spot it.
Second-language English candidates compared
Second-language English. Candidates whose first language is not English score no differently once you account for who they are and where they are. All 96 items fall inside the negligible band.

How the scales relate to behavior

Alongside the inventory, people reported on their own conduct at work: the helpful behaviors that go beyond the job description, and the counterproductive ones that damage the organization.

Helping colleagues tracks Agreeableness. Going beyond the role for the organization tracks Conscientiousness, and counterproductive behavior toward the organization tracks it in reverse. For every behavior, the strongest relationship is with the factor the research literature predicts, and the correlations sit at or above the top of the published range.

These behaviors are self-reported and were measured in the same sitting as the inventory. A study against operational outcomes is under way with a global staffing partner.

03 · Judgment

Situational judgment tests come with a manual here.

Most SJT documentation stops at the job analysis and the critical incidents. We treat that as the starting line. Every scenario bank we build is documented from end to end: where each scenario came from, who approved it, how the scoring key was derived, and what the fairness testing found.

Nine gates, each one signed

From your material to a delivered assessment in nine stages, and no stage opens until a named person has approved the one before it. That record is what makes the finished instrument defensible. It exists because the test was built properly, not because a file was assembled afterwards.

We tested the method on itself

Almost everyone building SJTs with AI has landed on the same design: a panel of AI raters, each given a different expert persona, averaged into a scoring key. We built one, then measured what its scores were actually responding to. The personas contributed nothing. Almost all of the variation came from which model family the rater was built on. We rebuilt our panel around that finding, and it now agrees with the full panel at a fraction of the cost.

Fairness, before anyone sits it

We ran 512 tests for items that behave differently across demographic groups. None were flagged. Every scenario also carries its own audit trail, so a question about any single item has a documented answer rather than an opinion.

Each character is written down, approved, and then locked. That text drives every frame the character appears in, in every language, so the cast does not drift between scenes.

04 · Reasoning

A test you have to hold.

The industry's instinct has been to look for a kind of question a language model cannot answer, and build a product on it. That position keeps collapsing: every benchmark designed to be hard is solved within months.

So we built a test where nothing depends on a model being weak. The question is only visible while the candidate is physically holding it on the screen. Let go, to reach for a phone, open another window or take a screenshot, and the question disappears while the clock keeps counting. A screenshot taken without holding it captures an empty screen.

Underneath that mechanic is an ordinary adaptive test with a calibrated item bank, and the design work is documented the same way as everything else on this page. Calibration is under way, and we will publish what it finds.

Held

Item on screen. Clock counting.

Released

Screen empty. Clock still counting.

The clock never speeds up or pauses. Letting go costs you the question, not the time.

Three real items from the bank, at the easiest live level, in the test's own interface. The orb drops to the floor and waits to be picked up; the clock does not start until it is. Part way through, one release hides the item and drops the orb. The panel beside it fills in with the movement data as the interaction produces it.

Answering means holding on, not clicking.

Back to← Home PreviousReasoning NextPlatform →

Open Research Program

Five studies, registered before the data existed.

Five studies: validation, synthetic respondents, faking resistance, cross-cultural invariance, and the criterion validity of the adaptive and maladaptive scales. The analysis rules were fixed in advance, with a written commitment to publish the results either way.

Everything on this page is the minimum we think a buyer should expect, not a ceiling. It is what anyone selling an assessment should be able to show you, and we would rather be measured against that standard than argue about it.