Every recommendation carries the grade of evidence behind it — and what science is still arguing about.
Every brain is biased. This one knows it.
Hypotheses about situations, never labels about people. No individual scores, no surveillance.
A manager observes a behaviour, forms an explanation within seconds and acts on it. That explanation is almost never tested. When it is wrong, the intervention is wrong with it — and the cost shows up months later.
of managers received any formal training in managing people.
The rest learned by watching.
Gallup ¹
or more of transformations fail to deliver — and the reason is behaviour, not technology.
The new tool arrives; the old workflow carries on.
change management literature ²
feedback interventions worsen the performance of the person who received them.
Doing something is not better than doing nothing.
Kluger & DeNisi, 1996 ³
¹ Gallup, survey of managers. ² Meta-analyses and surveys of organisational transformation programmes. ³ Kluger, A. N., & DeNisi, A. (1996). The effects of feedback interventions on performance. Psychological Bulletin. — Sources verified before publication; figures rounded for readability.
Observable behaviour, not interpretation. “The spreadsheets still go out by email”, not “the team is resistant”. If interpretation comes in, the tool hands the question back.
Each question separates competing hypotheses. The theory being tested never appears — asking “is this territoriality?” contaminates the answer.
Support shown in three notches — low, moderate, high. Never a percentage: there is no measurement behind one, and pretending otherwise is where the lying starts.
You record what you expected to observe, and a date. At closing, an intervention that did not work becomes evidence against the hypothesis that motivated it.
Not all information withholding is a problem: it can be commercial confidentiality, need-to-know, or protection against exposure at the wrong stage. When that is the case, the diagnosis ends in legitimate withholding — a valid outcome, not a failure of the tool.
We show even what science still disputes — because trust is built on honesty. A D grade does not mean “useless”: it means you know what ground you are standing on.
The grade measures the state of the scientific literature on that model — not how useful it is to you. It answers “how firm is the ground I am standing on?”.
Several independent studies reached the same result, and a meta-analysis pooled them. The direction of the effect is reliable; its size still varies with context.
The base is solid, but without the replication density of an A. Use it and observe what happens in your team.
The effect exists, but its size and conditions are disputed among researchers. Treat it as a hypothesis to test, never as the sole basis for a decision.
Popular, but with a thin or contested empirical base. Useful for naming the phenomenon in a conversation — not for justifying a decision.
The grade is not the recommendation. A model graded A may have nothing to do with your case, and a C may be shouting in the signals the tool found. The grade speaks about the science; support speaks about your session. The two rulers appear together and are never summed.
When–where–how written down before acting closes the gap between intention and behaviour.
A specific, difficult goal beats “do your best” — provided there is feedback on progress.
Without the belief that exposing error is safe, information does not circulate, however many channels exist.
Every behaviour requires capability, opportunity and motivation. It is the skeleton of the diagnosis.
Feedback aimed at the self, rather than the task, tends to worsen performance.
Losing weighs more than gaining the equivalent — what change takes away is felt first.
The first number said organises every number after it, project deadlines included.
Autonomy, competence and relatedness sustain motivation that does not depend on chasing.
What the group actually does weighs more than what policy says should be done.
The artefact becomes “mine”: changing its authorship is felt as losing territory.
The current state is preferred purely for being current, even with no verifiable advantage.
Effort is calibrated by comparison: perceived unfairness adjusts effort downwards.
Cue, routine, reward. The old workflow survives because the cue is still there.
Overload is not unwillingness: it is competition for a finite resource of attention.
Defaults matter — but effect size varies a great deal with context and audience.
Adoption spreads through social layers; the classic typology is more descriptive than predictive.
Investment already made traps the decision — the effect is real, its magnitude is debated.
Temporal landmarks open windows for change; partial replication and context-sensitive.
Avoiding threatening information: plausible in the field, empirical base still thin.
Under heavy methodological review — we use it only as a warning, never as a conclusion.
Grades are reviewed with every update to the base and the history stays public: if an effect loses support in the literature, it is downgraded in plain sight.
“The first principle is that you must not fool yourself — and you are the easiest person to fool.”
Richard Feynman · Caltech, 1974
This is why the tool investigates before concluding — including against your first hypothesis, which is precisely the most comfortable one.
“What you see is all there is.”
Daniel Kahneman · Thinking, Fast and Slow, 2011
What you observed is little, and the mind fills in the rest on its own. The questions exist to bring into view what was left outside it.
“Nothing is as practical as a good theory.”
Kurt Lewin · 1943
Theory here is not decoration: it is what lets you predict what happens if you change the incentive instead of changing the person.
The Culture team sees what repeats across teams: which mechanism dominates the quarter, which interventions worked, how often the diagnosis concluded it was not behaviour. None of it goes through reading anyone's session.
Any slice with fewer than five sessions is suppressed — below that, the aggregate would be re-identifiable.
A label that looks like a person's name is refused at entry. No individual profile exists anywhere.
The aggregate panel has no path back to the individual session. By design, not by an unchecked permission.
Every support read produces a record visible to the session's owner. The log cannot be erased.
Jun–Aug · 46 sessions · 12 teams
absolute session counts — not percentages
When you close a follow-up, the tool records the outcome — including the negative one. Over time, the recommendation starts arriving with its baggage: “in similar cases, this intervention worked in 4 out of 7”.
An intervention applied with no effect is evidence against the mechanism that motivated it — and lowers its support in the next session about the same team.
The base does not rewrite itself from usage. Any change to the taxonomy goes through review by whoever answers for the method.
Every change to a grade, mechanism or intervention is recorded with date and reason. If something was downgraded, you can see when and why.
It reads the description, runs the questions and writes the report. The diagnosis is method — the decision to intervene remains yours.