05th Oct, 2026 Read time 9 minutes

What Happens to Psychosocial Data After the First Round

By Ian Collins, Founder and Managing Director, Wellbeing Daily

 

If your second psychosocial diagnostic comes back cleaner than the first, be careful before you take it to the board.

Improved scores can mean the controls worked. They can equally mean a workforce has learned what happens when you answer honestly, and adjusted accordingly. Both look the same on a summary page. Both show an upward arrow.

In an earlier piece for HSE Network, I argued that the injury rate cannot tell you whether psychosocial controls are working, and that organisations need leading measures built around the conditions ISO 45003 asks them to assess. This is the uncomfortable follow-up. Those measures come from people, and people adjust what they say based on what happened the last time they said something.

The first round is the honest one

The first time an organisation runs a psychosocial diagnostic, people usually answer candidly. They have no history with the instrument. They do not yet know what the results will be used for, so they tend to take the stated purpose at face value. Often, there is a backlog of frustration, and the diagnostic is the first official place anyone has offered to put it.

That candour is fragile. Everything the organisation does with the round one results becomes a lesson for round two, and workers watch what happens closely, because the stakes for them are real. Very little of this involves a decision to game anything. It is ordinary human caution.

Donald Campbell described the general pattern back in 1979. The more heavily a quantitative indicator is used to make decisions, the more pressure it comes under, and the more it distorts the processes it was meant to monitor. Whilst Campbell was writing about social policy, there have been many examples of this since. Psychosocial data is unusually exposed, because the people generating the numbers are the same people affected by the decisions made with them.

 

How the honesty gets designed out

Anonymity that turns out not to be. 

Most diagnostics promise nobody will be identified, and most set a minimum group size before results are shown. Then someone asks for the results split by vessel, by rank, by nationality. On a ship with a crew of twenty-two, the engine department cut by rank can be three people. 

Workers do not need to see the dashboard to work this out. They only need to hear that a superintendent asked who wrote a particular comment. One story like that travels a fleet faster than any communication plan.

Results that land on individuals. 

A team reports poor leadership, and head office responds by sending someone to have a conversation with the supervisor. Sometimes that is exactly the right response. 

The crew, though, reads it as a consequence of what they said, and a crew that likes its supervisor, or simply has to live with them for the next ten weeks, will protect that person next time. Scores for that manager go up. Nothing about the conditions that produced the original result has changed.

Nothing happens at all. 

This is the more common failure and it does more damage than the first two together. If round one surfaced a workload problem and a year later nothing visible has changed, people draw a rational conclusion about what their answers are worth.

A plant I worked with had short-notice weekend call-ins as the clearest finding in round one. A couple of departments had both raised it, unprompted, in almost identical terms: it is very difficult to plan anything with your family, because Friday afternoon might turn into Saturday morning. The company’s response was an employee assistance line and a wellbeing week. 

When I went back the following year, one of the team leaders told me the call-ins had not changed, and that most of her shift had not bothered filling the survey in the second time around.

What follows is measurable. Response rates fall, and they do not fall evenly. The people who stop answering are disproportionately the exhausted and the disengaged, the very people whose answers mattered most, so the average improves while the underlying picture gets worse.

Scores become targets.

Once driver scores appear in a site manager’s scorecard, that manager has a reason to influence them, and most do nothing crude. A toolbox talk the week the diagnostic opens, a reminder of how much has been done this year, a mention that head office is watching. Each of those is defensible on its own.

 

Telling a real trend from a learned one

Consider a fleet where round two came back with fatigue scores up eleven points, workload out of the red, and manager support transformed from the weakest result to one of the strongest. Rotations were the same length as the year before. Manning was the same. Two of the four masters flagged in round one were still in post, and the crew under one of them had the highest turnover in the fleet.

(That example is a composite drawn from several maritime and offshore engagements. Identifying details have been changed.)

On the summary page, that fleet is indistinguishable from one where the controls actually worked. The difference shows up in the data around the score.

1. Start with the response rate and read it by group rather than in total. 

Rising scores alongside falling participation deserve suspicion. Rising scores in the groups where participation fell hardest deserve a great deal of suspicion.

2. Look at the spread as well as the average. 

Honest responses are messy, people disagree with each other, and a team of twenty rarely gives uniform answers about workload or manager support. When variation inside a team narrows sharply between rounds, and answers bunch at the favourable end, that pattern often reflects caution rather than consensus. 

The open text tells you something similar. Volume and specificity of written comments are usually the first casualties when trust erodes, so a first round full of detailed problems followed by a second round of short, general praise is worth a second look, whatever the scores say.

3. Check the diagnostic against the data nobody had to fill in. 

Turnover by rank and site. Sick days. Overtime and rotation extensions. Near-miss reporting volume, and where it has gone quiet. None of these can be curated by the person answering the questions. If the diagnostic says fatigue improved while overtime rose and rotations got longer, believe the overtime.

 

I would not push any of this further than it goes. Response rates fall for dull reasons, including a stale distribution list or a survey that opened during a crew change. Variation genuinely does narrow when a real problem gets fixed, and a team that has stopped arguing about workload may simply have less to argue about. 

None of these checks proves anything on its own. What they do is tell you when a clean result deserves a harder look before it becomes the headline number in a board pack.

 

What to settle before round one

Most of this is avoidable, and only if it is decided before the first round rather than after the first results arrive.

   ✔ Decide who sees what, write it down, and publish it to the workforce. Set a minimum reporting group and refuse to break it, including when a senior leader asks. The first exception you make is the one people will hear about.

✔ Keep psychosocial scores out of individual performance measures and bonus calculations. Report them at the level of the system, the site or the department, and treat a poor result as information about working conditions rather than a verdict on a person.

  ✔ Close the loop in public and on a timetable. Tell people what the results said, what will change, and what will not change. That last part carries more weight than most leaders expect, because workers deal perfectly well with being told the rotation cannot change this year and being given the reason. What they do not deal well with is silence.

  ✔ Treat the second round as a test of the first. Where round one was handled well, round two often looks messier in places, because people who trust the process raise new problems once the old ones are being dealt with. A clean, uniformly improved second round is not automatically good news. Sometimes it is the sound of a workforce deciding it is safer to say less.

 


About the Author: Ian Collins

Ian Collins is Founder and Managing Director of Wellbeing Daily, a workplace wellbeing and psychosocial risk consultancy working with safety-critical industries. His book, The Wellbeing Imperative, publishes with Penguin Random House on 24 November 2026.

 

References

Campbell, D.T. (1979). Assessing the impact of planned social change. Evaluation and Program Planning, 2(1), 67-90.

ISO 45003:2021. Occupational health and safety management. Psychological health and safety at work. Guidelines for managing psychosocial risks.

Brands who we work with

Sign up to our newsletter
Keep up to date with all HSE news and thought leadership interviews