728 x 90

When doctors and AI work together, who is responsible? How to establish medical liability

When doctors and AI work together, who is responsible? How to establish medical liability

Artificial intelligence is rapidly being integrated into hospitals and healthcare clinics. However, the benefits of AI tools come with new risks to patient safety and questions about who is responsible when things go wrong. Conventionally, health care responsibilities are clearly delineated. Doctors must provide an established standard of treatment. Institutions must organize safe treatment processes.

Artificial intelligence is rapidly being integrated into hospitals and healthcare clinics. However, the benefits of AI tools come with new risks to patient safety and questions about who is responsible when things go wrong.

Conventionally, health care responsibilities are clearly delineated. Doctors must provide an established standard of treatment. Institutions must organize safe treatment processes. Manufacturers must deliver non-defective medical devices. Regulators and bodies involved in licensing and credentialing can intervene if any of these parties do not meet expected standards. And courts can assess which party is responsible if patients are harmed.

An incoming wave of medical AI tools is set to blur these lines.

Current AI tools are primarily used as assistants, supported by human controls, as with other medical devices. For example, when AI models trigger an alert that a patient is at risk of sepsis, a physician must review and confirm the physiological rationale before any treatment.

But in the coming years, AI-powered clinical tools are expected to move from synthesizing data to acting on it, with less and less human involvement. They will make diagnoses, devise treatment plans, and make patient management decisions in which doctors play little or no role. While the results of many existing tools are understandable (sepsis models, for example, work by combining physiological variables defined according to a transparent formula).1 — more advanced tools may involve black-box deep learning processes. Doctors will no longer be able to fully follow the tools’ reasoning, even for algorithms that provide some explanations for their judgments.2.

These tools are situated among the accountability standards. Their black box reasoning makes it difficult to determine when they are defective. And when decisions are made jointly by humans and an opaque AI model, it is difficult to assess liability when people are harmed while receiving medical care. Existing medical liability frameworks do not address this crossover point3.

This legal uncertainty means that injured parties could fall into liability loopholes where no one has clearly broken a rule.3. Hospitals are likely to avoid AI for potentially useful functions due to concerns about legal risks, concerns that doctors repeatedly cite as a barrier to safe AI adoption.46 and which some consider to be reasons not to use the black box tools that are already available. And AI vendors could shirk their obligation to monitor safety once a tool hits the market, trusting that it will be difficult to assign blame when things go wrong.

Clear accountability frameworks are needed.

Here we describe how to achieve this, defining seven levels of AI capabilities based on three factors: autonomy, automation and operational scope. In the case of airplanes and autonomous vehicles, regulators already use graduated levels to specify exactly what tasks systems should perform, with what independence, and when humans are expected to take over.79. A similar set of tiers for medical AI systems would help regulators, policymakers, governments, courts, and healthcare professional bodies plan for incoming tools.

Three properties, seven levels

Three properties dictate how AI is used in the clinic.

First, autonomy: how independently the system reasons, without the need for human intervention. Existing sepsis risk models are low-autonomy artificial intelligence tools1. Deep learning systems, by contrast, could arrive at diagnoses from patterns across thousands of variables, making inferences that doctors can’t track.10.

Second, automation: what tasks a medical AI system can perform on its own. Such a system could range from a low-automation robot controlled by a surgeon to a system that sends chemotherapy prescriptions directly to a pharmacy.

Third, the operational scope: the limits within which the system is allowed to operate. A retinal imaging system implemented by a specialist ophthalmology service to verify clinical decisions is less risky than the same tool used in a primary care clinic to make an autonomous decision about whether to refer a patient to a specialist.

A woman undergoes an AI-assisted breast cancer scan while a medical professional monitors the test on a computer.

A scanner being tested in Krakow, Poland, uses AI to help detect breast cancer. Credit: Klaudia Radecka/NurPhoto via Getty

The following seven levels could form a basis for defining risks around medical AI tools. A given tool can be placed at different levels depending on where and how it is used and by whom.

Level 0: Informative. An AI system without autonomy, variable automation and limited operational scope. These systems store, move and display data according to rules set by humans. Examples include algorithms for managing electronic medical records or transmitting data from wearable devices that track vital signs. Algorithmic behavior is transparent and any results are independently verifiable.

This is well-trodden ground. Regulators can assess whether such tools are reliable and offer sufficient cybersecurity and data protection. Courts may apply conventional standards to assess liability.

Level 1: Assistance. AI systems with low autonomy, low to moderate automation, and limited operational scope. These systems can perform a single, clearly defined task supervised by a doctor, such as detecting suspicious rhythms on a heart monitor.11 or highlight possible lung nodules on a scan12. The tools’ algorithmic reasoning may be opaque, but the output can be verified through a clinical review of patient data or through “explainability techniques” that check for correlations between the model’s input and output data.

These tools are regulated as diagnostic aids, similar to existing computer-aided screening systems for mammography and colonoscopy. Hospitals are responsible for ensuring that staff understand their limitations. Users are expected to accurately judge whether alerts are significant.

Level 2: Decision support. AI tools with low autonomy, low to moderate automation, and moderate operational scope. These systems combine multiple streams of data, perhaps including symptoms, medical history, imaging data, and current guidelines. Existing tools at this level can produce a diagnosis and treatment plan, write detailed discharge summaries, and generate risk scores for triage decisions. Like level 1 tools, they are designed to offer advice: their reasoning may be opaque, but to some extent it is verifiable and doctors are expected to make the final decision.

Regulators should require strong evidence that these systems improve care. Hospitals must train doctors to question results rather than accept them by default, and courts must decide whether a doctor’s reliance on a particular recommendation was reasonable under the circumstances. There is a risk that doctors, under time pressure, will rely too much on counseling tools. There will be no easy way to determine the extent to which a doctor was influenced by a machine recommendation.

Level 3: Supervised automation. AI systems with moderate levels of autonomy and automation, and moderate operational scope. The system, rather than the doctor, is the one who makes the decisions. Currently, these tools are scarce (see ‘Classification of current medical AI’). In a strictly defined domain, they can dynamically respond to treatment needs; For example, insulin delivery devices automatically adjust doses based on continuous glucose readings, using built-in algorithms and predefined limits.13. The doctor selects which patients are placed in the system and only intervenes when something seems wrong.

Classification of current medical AI. A stacked horizontal bar chart ranks the use of artificial intelligence (AI) medical devices in the five countries with the highest number of device approvals. The levels range from level 0 (assisted use of AI) to level 3 (supervised automation).

Source: Analysis by K. Lam et al.

Regulators should focus on the quality of instructions for the use and maintenance of these systems and information on the risks inherent in the provision of care. Healthcare organizations should provide guidance for informed consent protocols that respond to the new types of risks inherent in AI-driven treatment, and providers should discuss these risks with patients.

With these systems, physicians can be criticized for enrolling the wrong patient, not responding to device alarms, or not understanding the system’s limitations. Organizations could be judged by how they selected, validated and monitored technology. Manufacturers could be liable if faulty designs or updates contribute to the damage. But conventional liability assessments may not be fit for purpose, if doctors and organizations acted reasonably and the harm is caused by an algorithmically determined decision that is not obviously wrong.

Level 4: AI-enabled human supervision. AI tools with moderate to high levels of autonomy and automation, and broad operational scope. These systems manage entire episodes of therapy in areas such as intensive care. They continually monitor patients and adjust treatments, for example, balancing ventilator settings and sedation with several physiological goals at once. They activate backup protocols, such as summoning a rapid response team when something appears to be going wrong. Doctors are a safety net and act only when the system sounds the alarm. Level 4 deployments have not yet achieved medical breakthroughs, although the necessary technical components are emerging.

For these systems, regulators should require strong evidence that including a human in the loop would reduce safety. Since doctors would have no realistic opportunity to review the actions of these systems, courts could have difficulty determining whether an adverse outcome reflects negligence, a hidden defect, or simply the inherent risks of the medicine.

For more tech updates, stay tuned to our blog.

Posts Carousel

Latest Posts

Top Authors

Most Commented

Featured Videos