The BI Method — Behavioural Intelligence for Teams

The future of leadership in an AI-enabled workplace.

When the first draft of the thinking is free, the scarce behaviour becomes knowing which draft to trust.

When the constraint moves

The behaviours that made someone effective when analysis was expensive are not the behaviours that make them effective when analysis is cheap. Producing a considered option used to be most of the job: gathering the numbers, drafting the memo, running the comparison. Producing a considered option is no longer the constraint. Judging one is.

Herbert Simon's account of bounded rationality described decision-makers as satisficing: settling for an option that is good enough, because gathering and comparing every alternative is too costly to be worth it (Simon 1955). The account scaled straightforwardly to how organisations decide. That constraint is precisely the one that has loosened. When producing a comparison is cheap, the old incentive to settle early weakens, and stopping at good enough becomes a choice rather than a necessity forced by scarce time and expensive analysis. Not everyone will make that choice well.

The shift is easiest to see in how a written recommendation now gets produced. A memo that once took a day to draft forced its author to notice gaps in their own reasoning simply through the effort of writing it down. That check happened for free, as a side effect of the slowness, not as a step anyone had to remember to take deliberately. When the draft arrives in seconds, the check has to be reinstated on purpose, or it does not happen at all.

That shift matters because judging and producing draw on different capacities. An account of two systems of thought describes a fast, associative mode that recognises patterns, and a slower, effortful mode that checks them, and the fast mode runs most decisions by default because it is cheap and usually good enough (Kahneman 2011). A fluent, well-formatted answer is exactly the kind of input the fast system accepts without much resistance. That was true before any of this technology existed. It is a larger problem now that the fluent answer arrives instantly and in volume, on demand, for any question at all.

The two capabilities that get expensive

Two capabilities in the framework should become disproportionately valuable. Evidence seeking, because the cost of accepting a fluent but wrong answer has risen. The shortcuts people use under uncertainty were characterised in conditions where information was scarce and effortful to obtain (Tversky & Kahneman 1974). An oversupply of fluent, plausible-sounding answers is a different environment, and there is no reason to assume the same shortcuts behave well in it. That is our expectation rather than a finding. And uncertainty calibration, because systems that express confidence uniformly make it harder to tell a firm conclusion from a plausible one. The habit this rewards is checking a claim against its source rather than against how convincingly it is phrased, which is precisely the discipline confirmation bias erodes, since people search harder for evidence that agrees with a conclusion they already like the look of (Nickerson 1998).

People trade thoroughness for efficiency continuously, in every domain, and the trade is usually invisible until it fails (Hollnagel 2009). A fluent answer that arrives in seconds makes that trade cheaper to make, and easier to make without noticing it has been made at all. The behaviours worth watching for are the ones that preserve thoroughness on purpose: pausing to check a source before acting on it is a small, ordinary act of friction that a genuinely useful tool makes very easy to skip.

Our expectation, which we intend to test rather than assert: given the same system and the same task, the difference in output quality will track how much verification a person does by habit more than how skilled they are at framing a request. That habit is observable, and it is the kind of thing a behavioural read is built to notice, well before it shows up in the quality of what anyone ships.

Producing a considered option is no longer the constraint. Judging one is.

This is not an argument for scepticism as a reflex. Simple heuristics frequently outperform elaborate models (Gigerenzer & Goldstein 1996), and expert intuition is genuinely reliable in domains with fast, valid feedback (Klein 1998). The judgement being asked for is knowing which kind of domain a given decision sits in — one where a fast read can be trusted, or one where it needs to be checked against something outside itself — and that judgement is itself a behaviour that can be observed, and is not evenly distributed across a team.

Where ownership goes when nobody drafted the answer

There is a third, less obvious shift. Ownership becomes harder to locate. When a recommendation has no author, the question of who is answerable for it needs deciding deliberately rather than by default. A memo written by a named person carries an implicit answer to who stands behind it. A synthesis produced by a system does not, and the same underestimation of what the other side does not know that undermines a human handover (Camerer, Loewenstein & Weber 1989) applies at least as strongly here, and plausibly more so, because there is no colleague on the other end to notice the gap and ask. That is our reading of the mechanism, not a measured result. Organisations that do not decide this in advance tend to discover it only once something has already gone wrong.

It cuts the other way too. Escalating-commitment research found that responsibility for an original decision is what makes people keep funding the version of it that is failing (Staw 1976). Strip out the author, and that particular failure mode weakens: nobody is defending a recommendation out of personal investment in having made it. But the same move removes a check that was doing real work — the discomfort of being visibly wrong, which is one of the more reliable forces that eventually stops a bad decision being carried further than it should be.

Human error is more usefully read as a symptom of trouble deeper in the system than as a verdict on the individual who made it (Dekker 2006), and the same posture is worth taking towards a wrong answer produced with the tool's assistance. The question is not only whether the person should have caught it. It is what about the surrounding process made catching it unlikely, and whether that same gap is sitting under every other recommendation this month, waiting to be found the same way.

None of this is a case for slowing adoption down. It is a case for measuring a narrower set of behaviours more deliberately: whether a recommendation was checked against a source before it was acted on, and whether confidence was stated in a way that matched the evidence behind it. Those are observable behaviours, and they are the ones a faster decision environment now depends on.

The research library lists the published work this draws on.

More insights