Clinician in the AI Loop: a faulty solution to a thorny problem

The vast majority of AI medical devices rely on having a clinician verify their outputs. On the surface, this seems good for safety, but in reality it’s a short-sighted solution to a much deeper problem.

Whenever a new clinical technology emerges, it’s normal to verify its outputs compared to the existing standard of care. For AI tools that perform “cognitive” clinical tasks, that means verifying against the judgement of an appropriately qualified healthcare professional - otherwise known as having a “clinician in the loop”. But over the last decade, having a clinician in the loop has evolved from just being a safety checkpoint to becoming a crutch for manufacturers to defer risk out of their products and into the hands of users. This solution is simply not sustainable as AI devices become more established and prevalent in clinical workflows. There are four main reasons for this: bias, liability, de-skilling and clinical burden.

Bias

Although “clinician in the loop” is meant to serve as a clinical safeguard, there’s no doubt that an AI output suggesting a potential clinical issue will bias the reviewing clinician’s judgement. Additionally, time-constrained clinicians are more likely to be subject to this bias in an effort to save time. The risk of bias is well known, and also something that is being explored in the MHRA Airlock programme

Liability

When manufacturers adopt a clinician-in-the-loop for their intended purpose, often the purpose of the clinician themselves is left quite vague: are the clinicians expected to rely on the AI outputs at all, or are they expected to perform an independent evaluation of the facts beforehand? This ambiguity opens the door for clinicians to be liable for acting on AI-related errors, as the Medical Protection Society recently reported. This liability is a natural consequence of the bias issue we’ve already discussed.

De-skilling

Medicine changes over time, and there are several historical procedures, tests, and devices which are no longer part of a clinician’s toolset. However, in the interim period between transitioning across technologies, there is a real risk of premature de-skilling of the clinical workforce, especially in fundamental clinical skills. We’ve already seen this with the use of stethoscopes, but concerns are mounting even faster for AI-induced de-skilling (see here, here and here). Premature de-skilling, particularly when manufacturers rely on clinicians in the loop as a risk-control measure, poses a real threat to clinical safety, as well as the training of new clinicians. 

Clinical burden

The supposed promise of most AI tools in medicine is relatively simple: they sell themselves as being faster, and potentially cheaper. But if clinicians are responsible for checking AI outputs on top of their existing clinical duties, this could paradoxically increase cognitive burden and workload. Emerging research already shows that this may be the case, where clinicians are forced to act as “janitors” of LLM-generated outputs. This issue is compounded by the fact that several point-solutions may now sit in a clinical workflow, meaning clinicians have to maintain oversight on an increasing number of outputs, each with a different set of biases and limitations to account for.

Three ways forward

There are three main routes forward that manufacturers should consider - and healthcare systems should expect - in order to stop this problem from getting worse.

First, for many products having a clinician in the loop is not going to go away. In these cases, however, AI tools must be designed in such a way that they are concurrent partners to clinicians in the workflow, and not just adding an extra step for clinicians to verify in the process. This means manufacturers should spend more time considering the day-to-day workings of the clinical pathways they are hoping to integrate into. This may increase upfront work and research, but in the long term it has two clear benefits: more efficient clinical workflows, and a higher chance of product-market fit.

Second, developers may want to consider how the use of AI in medical devices may in fact work in the background, rather than in the forefront of clinician-facing user interfaces. A fantastic example of this is Hardian client Oxipit (recently acquired by Sectra) - who developed a Class IIb certified tool which autonomously reports normal chest X-rays and removes them from radiologist worklists. There remains a human “on-the-loop” regularly auditing a sample of outputs, but not every CXR has to be reviewed by a human.

Third, and related to the second point, manufacturers should now consider being bolder in the claims they want to make; instead of requiring a clinician to “review” AI outputs, with sufficient evidence of safety and effectiveness, they should show some ambition and claim that clinicians can “rely” on them directly. Skin Analytics is a company that has already shown this is possible. The evidence burden is certainly higher, but the value generated is correspondingly high too.

This is all to say that having a clinician in the loop is important in many cases, but it must be done properly - and not just as a means to defer risk and maintain a low medical device classification. 

Learn how Hardian can help with mapping relevant clinical pathways to strengthen your value proposition and product strategy

Dr Ankeet Tanna

Dr Ankeet Tanna, Consultant

Next
Next

Is post-market surveillance of AI devices working?