26th Aug, 2026 Read time 10 minutes

After the ISO 45003 Assessment: What Should You Actually Measure?

By Ian Collins, Founder and Managing Director, Wellbeing Daily

 

Six months after the assessment, an operations director asked me a fair question. How do we know whether any of this worked?

His team had done the job properly. A survey across three installations, follow-up interviews, a risk register mapped clause by clause. Fatigue, workload and manager quality all came back as significant hazards. The company changed the handover process, added a second person to night shift on two platforms, and trained supervisors in how to hold a conversation about workload. Real money, real disruption, genuine commitment.

Then the board asked for evidence. The only number anyone could reach for was the injury rate.

(The scenario above is a composite drawn from several offshore and mining engagements. Identifying details have been changed.)

The stakes are not in dispute. HSE recorded 776,000 workers in Great Britain with work-related stress, depression or anxiety in 2023-24, accounting for 16.4 million lost working days, and mental ill health has been the single biggest cause of long-term absence in the UK for a decade. What remains unresolved is what a business measures once it decides to do something about that.

This is the point where most psychosocial risk work quietly stalls. Not at the assessment, which organisations are getting better at, but at the question of what to track once the report is written.

 

The number that cannot answer the question

The injury rate gets reached for because it is already on the board pack. It is familiar, it is benchmarked against peers, and everybody in the room knows what it means. It is also the wrong instrument for this job, for two separate reasons.

The first is statistical. Hallowell and colleagues examined this directly in Professional Safety in 2021, and their finding was uncomfortable. For most organisations, at the recording volumes they actually work with, year-to-year movement in the total recordable incident rate is indistinguishable from random variation. The number moves. It just does not reliably mean anything. Yet a two-point drop gets read as a trend, and decisions follow.

The second reason matters more. Even a perfectly reliable injury rate counts events that have already happened. Psychosocial hazards sit upstream of those events, sometimes years upstream. Chronic under-manning does not produce an injury on the Tuesday it starts. It produces a slow erosion of margin, until an ordinary day with an ordinary complication finds a crew with nothing left to absorb it. By the time your lagging indicator moves, the condition you were trying to manage has been present for a long time.

The operations director was in a bind that had nothing to do with his programme. He had made sensible changes to a real problem, and had no instrument capable of detecting whether they helped.

 

What the standard asks for, and what gets skipped

ISO 45003 is guidance rather than a certifiable requirement, which is worth stating plainly because it changes how people use it. In practice, a great many organisations treat it as a survey instrument. They lift the hazard categories, build a questionnaire, run it once, and file the output.

Read further into the standard and the shape of the intended work becomes clear. Identifying psychosocial hazards is the beginning. What follows is control, then monitoring and evaluation of whether those controls are doing anything. The standard follows the management-system logic of ISO 45001, and that logic is a cycle rather than a snapshot.

Stopping at identification is exactly what leaves you with the problem I have described. You have documented your hazards, which is useful and auditable. You have not built any means of knowing whether your response to them is working. The measurement question is not an add-on to the assessment. It is the part of the standard most implementations leave on the table.

 

Measuring capacity instead of counting failures

There is a serious alternative, and it is not new. Sidney Dekker and Michael Tooma set it out in the International Labour Review in 2022, arguing for a capacity index to replace incident-based metrics. Rather than counting how often an organisation has failed, the approach measures its capacity to absorb variability and recover when something goes wrong.

The underlying question changes character. Instead of asking how many people got hurt last quarter, you ask whether the organisation knows what could hurt people, understands whether its controls actually work in practice, resources those controls properly, monitors them, complies with what it has committed to, and verifies any of this independently. Each of those is assessable. None of them requires anyone to be injured first.

Dekker and Tooma built the construct around operational and technical capacity, and it holds up well there. It has a gap, and the gap is people. A crew can sit inside a system where every control is documented, resourced and verified, and still have no capacity left, because they are eleven weeks into a rotation, three short, and reporting to a supervisor who treats a question as a challenge. Nothing in the technical layer sees that. The controls are all present. The margin is gone.

 

Building the human layer

This is where the psychosocial work you have already done becomes useful rather than decorative. The hazards ISO 45003 asks you to identify are the same conditions that determine whether a crew can absorb a bad day, which means your assessment already contains the raw material for a leading measure. It needs structuring into domains you can score and track over time.

Six work reasonably well across shift-based, safety-critical operations. Leadership and relationships is the largest single lever, and offshore it is concentrated in one or two people nobody can walk away from. Work design and demands covers rotation length, watchkeeping pattern, manning levels and the gap between what the job requires and what the crew has to meet it. Flexibility and balance means something different at sea than ashore, where the realistic question is control over rest rather than choice of hours. Recognition and growth matters for retention long before it matters for morale. Purpose and meaning holds up better in these industries than in most, and is worth measuring precisely because it is a genuine asset. Culture and safety covers whether people report early, and whether reporting early has ever cost anyone anything.

The weight I have put on the first of those is not a personal preference. Richard Smith and Sara Silvonen, reviewing five years of UK workplace wellbeing data for the Johns Hopkins Human Capital Development Lab, found that scores rise with every rung of the management ladder, and that frontline managers sit closer to the people below them than to their own management peers. Nearly two thirds of those frontline managers report being excessively stressed by job demands often or almost always. The relationship between confidence in management and wellbeing is close to linear and has held steady for five years. Put a master, an OIM or a shift supervisor in that finding and most readers will recognise it immediately. The tier we rely on to deliver the controls is frequently the most exposed tier in the system, and the one with least authority to change what is exposing it.

Smith and Silvonen also land where this article lands, which is worth noting because they come at it from human resources rather than safety. Their recommendation is job design, stress monitoring and fair people management practices, over awareness programmes, taglines and perks.

I would be honest about the limits. Their data is self-reported perception, gathered from organisations that chose to participate, and the authors say plainly that objective health sits outside the scope of a survey. Perception tells you how a crew feels about the work. It does not tell you the roster is 84 days or that the second engineer is covering two roles. Composite indices of this kind, mine included, are useful management instruments that are not yet independently validated as predictive. Treat them as an early warning system that beats what you have, not as settled science.

 

What you can do on Monday

Three things, none of which require a consultant.

Go back to your assessment and pull out the three hazards that scored worst. Write down, in one sentence each, what would have to be true in twelve months for you to believe the problem had improved. If you cannot write that sentence, you do not yet have a measure.

Then look at what you are already collecting. Rotation length, overtime, turnover by rank, time to fill vacancies, near-miss reporting rates and their distribution across sites. Most organisations hold more leading data than they realise, sitting in systems that never speak to the safety function.

Finally, stop asking the injury rate a question it cannot answer. Keep reporting it, because you are required to and because sharp movements still warrant a look. Just stop treating it as the verdict on work it was never built to assess.

The operations director in that composite eventually got his answer. Not from the injury rate, which did what injury rates do and drifted without meaning. From the fact that on the two installations where night shift was reinforced, crews began reporting fatigue earlier, and their supervisors started raising manning concerns before rather than after a rotation. Neither of those is a safety metric in the traditional sense. Both told him more than the number the board asked for.

 


About the Author: Ian Collins

Ian Collins is Founder and Managing Director of Wellbeing Daily, a psychosocial risk consultancy working with maritime, mining, oil and gas, and offshore operators. His book, The Wellbeing Imperative, publishes with Penguin Random House on 24 November 2026.

 

References

Dekker, S.W.A. and Tooma, M. (2022). A capacity index to replace flawed incident-based metrics for worker safety. International Labour Review, 161, 375-393. https://doi.org/10.1111/ilr.12210

Hallowell, M.R., Quashne, M., Salas, R., Jones, M., MacLean, B. and Quinn, E. (2021). The Statistical Invalidity of TRIR as a Measure of Safety Performance. Professional Safety, 66(4).

Health and Safety Executive (2024). Health and safety at work: Summary statistics for Great Britain 2024.

ISO 45003:2021. Occupational health and safety management — Psychological health and safety at work — Guidelines for managing psychosocial risks.

Smith, R.R. and Silvonen, S. (2025). Fostering Wellbeing at Work in the UK: A Five-Year Review. Johns Hopkins University Human Capital Development Lab and Great Place To Work UK.

 

Brands who we work with

Sign up to our newsletter
Keep up to date with all HSE news and thought leadership interviews