Book a demo
Home›The science
AI has changed how an assessment can be designed. It has not changed how one has to be proven. This page shows what we built with it, a behavioral framework, a personality instrument, a judgment studio and a reasoning test, and the evidence behind each one.
I trained as an aerospace engineer in the nineties, and I arrived just as something changed. Computers got powerful enough to simulate the physical world, and design stopped being build-it-and-see. You could try a thousand wings before you cut any metal.
The same shift is now happening in psychology, in the part of the work that used to set the pace: writing the items. We generate thousands of candidate items. Each one is reviewed and either edited or removed: some are off target, some are pitched at too high a reading level, some are worded in a way that could disadvantage a group of candidates. Then Pseudo Factor Analysis runs across the whole pool. It is a structural read of the items themselves, done before a single person has answered one, and it tells us which items are likely to behave. Only those go forward to trial.
That work used to take months. It now takes a day.
What we do not skip is people. Every instrument is trialed on real candidates, without exception. The simulation finds the design; the trial is what proves it, and that order never changes.
The result is not a cheaper test. It is a test built for the job.
— Tariq Shaban, Co-Founder & Product & Innovation Director
◆ The method
The same four steps built the behavioral framework, the personality instrument, the judgment studio and the reasoning test. The first three take a day. The fourth takes as long as it needs.
Thousands of candidate items, written against tight design briefs and anchored to specific behaviors.
Every item is checked for target, reading level and wording that could disadvantage a group. Items come back edited, or they come out.
A structural read of the items themselves, showing which ones are likely to behave before anyone has answered one.
Every instrument, without exception. The simulation finds the design; the trial proves it.
Nothing in the pipeline is taken on trust. Subject-matter experts' judgments are checked for agreement before they become a scoring key, and every stage a model touches is checked against human data on the same content. The chart shows the example we publish: we ran Pseudo Factor Analysis on a set of items, ran a conventional factor analysis on real responses to the same items, and compared what each one recovered. Pseudo Factor Analysis gets most of the way there on its own, which is all it needs to do. Its job is to decide which items go to trial, not to replace the trial.
◆ 01 · The behavioral framework
Every assessment predicts something, and that something is usually vague. In most organizations, “performance” ends up meaning one manager's overall rating, which says little about what the person actually did.
Competency models were meant to solve that, but most are written in a two-day workshop. Ours was built from published evidence. We collected 898 behavior statements from 46 published sources, covering four decades of research taxonomies and the competency dictionaries in professional use, and let the structure emerge from that material rather than from a room of people agreeing with each other.
We then applied the same method we use for items: generate, review, Pseudo Factor Analysis, then expert adjudication. The result is six clusters, eighteen competencies and seventy-seven observable behaviors, each one traceable to the source it came from.
An organization that cannot name the behaviors it values cannot select for them, develop them, or defend its decisions about them.
◆ 02 · Personality
Six factors, twelve traits, twenty-four facets and forty-five scales, in about twenty minutes. The instrument was calibrated on 662 working adults and has since been taken by more than fifteen thousand applicants. Every scale meets the reliability standard, and the structure we designed is the structure that shows up in the data. The technical manual runs to two hundred pages and is published in full, which very few publishers do.
Alongside the inventory, people reported on their own conduct at work: the helpful behaviors that go beyond the job description, and the counterproductive ones that damage the organization.
Helping colleagues tracks Agreeableness. Going beyond the role for the organization tracks Conscientiousness, and counterproductive behavior toward the organization tracks it in reverse. For every behavior, the strongest relationship is with the factor the research literature predicts, and the correlations sit at or above the top of the published range.
These behaviors are self-reported and were measured in the same sitting as the inventory. A study against operational outcomes is under way with a global staffing partner.
◆ 03 · Judgment
Most SJT documentation stops at the job analysis and the critical incidents. We treat that as the starting line. Every scenario bank we build is documented from end to end: where each scenario came from, who approved it, how the scoring key was derived, and what the fairness testing found.
From your material to a delivered assessment in nine stages, and no stage opens until a named person has approved the one before it. That record is what makes the finished instrument defensible. It exists because the test was built properly, not because a file was assembled afterwards.
Almost everyone building SJTs with AI has landed on the same design: a panel of AI raters, each given a different expert persona, averaged into a scoring key. We built one, then measured what its scores were actually responding to. The personas contributed nothing. Almost all of the variation came from which model family the rater was built on. We rebuilt our panel around that finding, and it now agrees with the full panel at a fraction of the cost.
We ran 512 tests for items that behave differently across demographic groups. None were flagged. Every scenario also carries its own audit trail, so a question about any single item has a documented answer rather than an opinion.
Each character is written down, approved, and then locked. That text drives every frame the character appears in, in every language, so the cast does not drift between scenes.
◆ 04 · Reasoning
The industry's instinct has been to look for a kind of question a language model cannot answer, and build a product on it. That position keeps collapsing: every benchmark designed to be hard is solved within months.
So we built a test where nothing depends on a model being weak. The question is only visible while the candidate is physically holding it on the screen. Let go, to reach for a phone, open another window or take a screenshot, and the question disappears while the clock keeps counting. A screenshot taken without holding it captures an empty screen.
Underneath that mechanic is an ordinary adaptive test with a calibrated item bank, and the design work is documented the same way as everything else on this page. Calibration is under way, and we will publish what it finds.
Item on screen. Clock counting.
Screen empty. Clock still counting.
The clock never speeds up or pauses. Letting go costs you the question, not the time.
Answering means holding on, not clicking.
◆ Open Research Program
Five studies: validation, synthetic respondents, faking resistance, cross-cultural invariance, and the criterion validity of the adaptive and maladaptive scales. The analysis rules were fixed in advance, with a written commitment to publish the results either way.
Everything on this page is the minimum we think a buyer should expect, not a ceiling. It is what anyone selling an assessment should be able to show you, and we would rather be measured against that standard than argue about it.