The BI Method — Behavioural Intelligence for Teams

The difference between personality and observable behaviour.

One asks who someone is. The other asks what happened. Only the second can be improved on purpose.

What a trait label actually buys you

Personality instruments describe disposition, and describe it durably. That is their strength and the reason they are poorly suited to organisational improvement: traits do change, but slowly and rarely on a timescale a team can work with, and a type label tends to become an explanation for behaviour rather than a description of it.

Told that a colleague is disorganised, a team stops asking what specifically went missing and starts explaining every future lapse by the label — a version of the attribution error that leads people to explain what others do by what they are like, and to underweight the situation they were in (Ross 1977). Once the label is in place, evidence that would complicate it tends not to be sought, because the label has already done the explaining. The label does the work the evidence should have done.

The same substitution shows up in how feedback gets given. Told to be more proactive, a person has no next action available to them, because proactive is not something one simply decides to be more of by Tuesday. Told that a client update went out two days late three times this quarter, the same person has an instance to examine, a pattern to notice for themselves, and a specific point at which the next attempt could reasonably go differently.

Traits are not immovable. They drift over a working life, but the timescale is measured in years, not sprints. Even ordinary habits, which are far more tractable than a trait, take considerably longer to become automatic than most change programmes assume, and the time varies enormously between people and between behaviours (Lally et al. 2010). A team cannot wait out a trait, and a quarterly review cannot meaningfully act on one. It needs something that moves on the timescale the team is actually operating on.

A narrower question, asked on purpose

Behaviour Intelligence asks something narrower on purpose. Not what a person is inclined towards, but how often they do things a colleague would be able to see: whether context was shared, whether a concern was raised, whether a commitment was met. Ajzen's work makes the requirement precise: a behaviour has to be specified by action, target, context and time before anything useful can be said about it (Ajzen 1991). “Shares context” is not a behaviour. “Passed on the open questions at the last three handovers” is.

Told to be more proactive, a person has no next action available to them.

There is a reason this narrower question also tends to be more actionable. Hackman and Oldham treated autonomy and feedback as properties of how work is designed rather than of the person doing it (1976) — though their own test measured those characteristics as perceptions rather than observing them, and found the effect depended partly on how much growth the individual wanted from the job. Our working expectation is that behaviour tracks structure more readily than it tracks disposition: change what a role allows and asks for, and the behaviour it produces tends to follow. Change what someone is like and, per the point above, the clock runs in years.

This also explains why an assigned behaviour sometimes sticks and sometimes stops the moment nobody is looking. Behaviour that satisfies a person's own sense of autonomy, competence and connection to others tends to persist without monitoring; behaviour extracted through pressure alone tends to evaporate as soon as the pressure lifts (Deci & Ryan 1985). A behavioural read that only counts instances, without asking why they happened, cannot tell these two apart on its own. That is a real limit, and one worth stating plainly rather than glossing over.

It is a self-report, and it carries the limits self-reports carry. People under- and over-report in patterned, not random, ways: the behaviour that would look bad gets minimised and the behaviour that would look good gets inflated, partly as deliberate self-presentation and partly without noticing at all (Nederhof 1985). Any single self-report instrument has to be designed against its own tendency to inflate — reverse-keyed items, behavioural anchors, corroboration against what colleagues report, rather than trust placed in the honesty of any one respondent (Podsakoff et al. 2003).

Rating what people do, not what they are

This is not a new idea in measurement, even if it remains a less common one in how organisations actually run reviews. The case for behaviourally anchored rating scales argued for exactly this substitution decades ago: anchor a rating to a specific, observable action rather than to a trait judgement, because two raters can agree on whether a deadline was met far more reliably than they can agree on whether someone is proactive (Latham & Wexley 1981). The psychometric gains those scales actually delivered turned out to be more modest than the argument promised; the argument itself survived, and has been available for over forty years. What has been missing is not the theory but the discipline of collecting the instances at scale, consistently enough for a pattern to emerge from them.

Memory research points the same way. Retrieval depends on how well the cue matches the way the event was encoded in the first place (Tulving & Thomson 1973). A colleague asked whether a named behaviour occurred in a specific handover is given a cue that matches an actual episode. The same colleague asked how proactive a person generally is has no such cue, and is reconstructing an impression rather than retrieving one. The behavioural version is simply the easier one to answer honestly.

Naming an event precisely enough to rate it is also the first step in changing it. Any intervention aimed at behaviour eventually has to move one of three things: whether someone is capable of the behaviour, whether the opportunity to do it exists, or whether they are motivated to (Michie, van Stralen & West 2011). A trait label answers none of those questions — it describes an outcome without pointing anywhere. A specific, rated instance at least points towards which of the three is missing, which is the difference between a diagnosis and a description.

What a report can actually say

The same separability argument has proved itself elsewhere in the literature. Burnout, once treated as a single mood, was decomposed into exhaustion, depersonalisation (later relabelled cynicism) and reduced personal accomplishment as three measurable dimensions rather than one (Maslach & Jackson 1981), and the decomposition is what allowed the construct to be acted on at all: a manager can address cynicism and exhaustion very differently, but only once they know which one they are actually looking at. A trait label offers no equivalent decomposition. A behavioural read, tracked across specific domains, does.

The practical difference shows in what a report can say. A trait profile can tell someone they are low in agreeableness, which is difficult to dispute and not obviously actionable. A behavioural read can tell a team that difficult information reaches it late, and name the change that would alter that. One is a description of a person. The other is a description of an event, repeated often enough to be a pattern — and a pattern, unlike a trait, is something a team can decide to change.

The research library lists the published work this draws on.

More insights