The BI Method — Behavioural Intelligence for Teams

Behaviour as a leading indicator of organisational health.

Engagement scores tell you how people feel about last quarter. Behaviour tells you what next quarter will look like.

Why lagging measures arrive late

Much organisational measurement is lagging by construction. Attrition and delivery variance describe a state that has already arrived, and by the time either moves, whatever produced the movement has been operating for a while. Engagement sits differently, and its evidence base is unusually well studied: across a meta-analysis of business units, engagement correlates with a composite of business-unit outcomes at around .22 as observed, rising to .38 once corrected for measurement error and range restriction (Harter, Schmidt and Hayes 2002). Those are one relationship reported at two levels of correction, not a range across outcomes. The study is Gallup-authored and uses Gallup's own engagement instrument, which is worth knowing when reading the size of the effect, and the direction of causation is genuinely contested. The limit is not that these measures are weak. It is that by the time they move, the behaviour that moved them is already established.

Part of the difficulty is that engagement is a harder construct than the single score implies. A careful review of the literature found the term used to describe overlapping but distinct things — a psychological state, a set of dispositions, a set of behaviours — bundled into one number that then gets treated as a single lever to pull (Macey & Schneider 2008). A score that conflates state and behaviour cannot cleanly tell a leader which one to act on, or whether acting on it will move the outcome at all. It can tell them something moved. It cannot reliably tell them what to do next.

Attrition is a clean illustration of the same lag. By the time a valued employee has resigned, the decision was very likely made weeks or months earlier, and everything observable in that window — the missed one-to-one, the quieter contribution in a planning meeting, the concern raised once and then not raised again — happened before the resignation letter did. A measure that only counts the letter is counting the right thing far too late to act on it.

What sits earlier in the chain

Observable behaviour sits earlier in the chain. Whether concerns get raised while they are still small, whether knowledge leaves the person who holds it, whether disagreement is treated as information: these move first, and the outcomes they shape follow. Edmondson's study of 51 work teams found that psychological safety predicted a team's learning behaviour — asking questions, seeking feedback, discussing error openly — and that this behaviour in turn predicted performance (1999). Surfacing a problem is itself the behaviour being measured, and it happens long before the outcome it protects. That is the property a leading indicator needs. It is visible while there is still time to intervene, rather than after the outcome has already been decided.

Safety changes whether a mistake gets surfaced, not whether it happens.

Whether that surfacing happens is itself behavioural, and measurable well before it shows up anywhere else. Managerial openness — whether a manager is actually experienced as receptive — predicted employee voice better than transformational leadership did (Detert & Burris 2007), which is to say that what a policy states is not what people are responding to, and staying silent is not the absence of a behaviour but an active one with its own causes, worth tracking in its own right (Morrison 2011). Both are observable long before an annual engagement survey would register that anything had changed.

The same logic holds outside safety specifically. Group performance itself was predicted less by the average intelligence of members than by two things: how evenly conversational turn-taking was distributed — a behaviour, observable in a single meeting — and members' scores on a test of reading emotional state from photographs of faces, which is not (Woolley et al. 2010, in ad-hoc laboratory groups). Only the first is something a behavioural read can pick up, and it is the one we track. Neither shows up in an engagement score, and both predate whatever outcome the group eventually produces by as long as the group has been meeting at all.

This is also why the categories a behavioural read tracks are not arbitrary. Team processes can be named and separated well enough to be measured on their own terms — how a team plans and coordinates, how it acts once the plan is under way, how members manage one another as people — rather than folded into a single mood score (Marks, Mathieu & Zaccaro 2001). Separating them is what lets a report say which process is slipping, rather than only that something, somewhere, has changed.

The same distinction catches a pattern a lagging measure would miss entirely. A team can hit its numbers by leaning on the heroics of one or two people while quietly running down the collective capability that would let it hit next quarter's numbers without the same strain (Hackman 2002). From outside, the output looks identical either way. The behaviour producing it does not, and only one of the two versions is repeatable.

What the limits of the construct force us to concede

None of this is a licence to treat behavioural measurement as immune to the problems that affect any other construct. A quarter-century of research on psychological safety has also traced where the idea has been over-claimed and applied past the conditions under which it was originally studied (Edmondson & Lei 2014). The honest position is the same one that applies to engagement: a behavioural read is evidence about the conditions producing a result, not a verdict on the result itself, and it should be treated with the same scepticism any single measure deserves.

Building it well is not trivial either. Turning a theory of team behaviour into an instrument that survives psychometric scrutiny is its own discipline (Wageman, Hackman & Lehman 2005), and the difficulty is a large part of why so much organisational measurement defaults to the easier, lagging alternative. Attrition is simple to count. A leading behavioural signal has to be validated before anyone can trust what it is showing them, and that work is slower and less visible than running another survey.

Lagging measures should not be discarded either. They remain the ultimate check on whether a leading signal was reading the right thing, because a leading indicator that never correlates with anything downstream is not actually leading. It is simply early, and beside the point. The two are complementary rather than competing, and a credible programme reports both instead of quietly replacing one with the other and hoping nobody asks why.

A different measurement cadence

The practical consequence is a different measurement cadence. A behavioural read is not an annual verdict, and it is not a substitute for the outcomes it precedes. It is a periodic look at the conditions producing next year's results, taken while there is still time to change them. Where an engagement survey asks how people felt about the last twelve months, a behavioural read asks what is happening right now that the next twelve months will be built from.

The answer is not to measure everything, all the time, which recreates the same fatigue that makes annual surveys go stale in the first place. It argues for choosing a small number of behaviours known to sit early in the chain, and returning to them on a rhythm short enough to still catch the drift while correcting it is still cheap.

The research library lists the published work this draws on.

More insights