Skip to content
TheBrief.Health

Clinical Briefs

Clinical Practices Use AI Even When Clinicians Don’t Know

While many clinicians are unaware that a predictive model is involved in much of the AI that is

Medical practitioner in scrub suit using a laptop for remote consultation and documentation.
Medical practitioner in scrub suit using a laptop for remote consultation and documentation.

While many clinicians are unaware that a predictive model is involved in much of the AI that is being implemented into their clinical workflows, an understanding of how AI is contributing to the many applications of clinical data will enable them to make informed decisions about both their patients’ care and how to effectively utilize the many workflow applications into which so much patient data are being incorporated. The majority of current applications of AI in clinical practice are not interactive, multi-functional, end-to-end clinical applications, but rather a collection of “modules” that support certain clinical applications such as scheduling and billing/coding as well as image interpretation (radiology) and laboratory result interpretation and reporting. Much patient data and clinical information are available through patient portals and are incorporated into the clinical applications that clinicians use on a daily basis through the clinical applications of the electronic health records (EHRs). An understanding of the many contributions of AI in these many clinical applications will help ensure that the best possible care is delivered to each patient. The doctor needs to be able to identify errors, detect bias and apportion blame for the errors, whether these have been caused by clinical error or system error.

Why It Matters

AI can shape care without anyone noticing

Artificial Intelligence (AI) is changing what clinicians focus on in their practice. As AI highlights potentially significant lab values, designates “high-risk” patients, makes a diagnosis, or highlights automatically prioritized messages, clinicians become accustomed to the influence of new actors on the factors that affect their interpretations and decisions with patient data. While clinicians remain ultimately accountable for proper clinical judgment and decision-making, their interactions with AI-generated messages over several days can fundamentally change the way that clinicians manage patient care and make referrals for additional tests and imaging studies, as well as prescriptions for antibiotics and follow-up appointments.

How could a simple tool affect your organization? Well, it could affect clinical decisions made by your staff as well as administrative decisions and clinical outcomes. Perhaps it would decide whether an appointment should be scheduled earlier in the day or later in the day or rescheduled altogether. Maybe it would determine which patients or call center clients would receive optimal service and which might receive less service and fewer payment and collection calls. The tool could also identify which patients are most likely to “no-show” for an appointment and trigger timely communications to confirm the scheduled appointment. Finally, the tool could determine which billing edits and claims processing rules should be reviewed by a coding specialist and/or require additional documentation for correct payment and by when these edits and claims would be addressed. What’s startling is that patients don’t even realize that an algorithm — artificial intelligence — has sorted them out, and is working against them.

Patient safety risks increase when AI is opaque

Providers and care teams want to know when a model is in use and have the ability to check the model’s trustworthiness at the time of use. Providers want to know the specific data sources that a model has been trained on, how recently the model was updated for known weaknesses or bias, and understand the limitations of use of the model. For example, clinicians want to know what data the model is using to generate a score (current vitals, past admissions, claims history, etc.). They want to know how the model performs on pregnant women, pediatric patients, rare diagnoses, complex comorbidities, etc. They want to know if the model has been validated on their specific patient population. We simply need you to provide a few more details so we can review the model to ensure correct and safe usage.

Selection bias, confirmation bias and a less well-recognised bias, the automation bias or over-trust of clinicians of the results provided by an AI system and thus failing to consider relevant evidence, are three biases that need to be addressed when developing AI systems in healthcare. These models are not better than human clinicians, and must be treated as such. The presentation of the AI’s results to the clinician also can have a role, for example presenting the results of the model as definitive labels rather than probabilities that a patient has a condition. Automated clinical decision support can automate a clinical decision. Even though the intent is to use the system to support the clinician’s decision-making, the system can end up driving the decision, by putting orders in the default order sets.

Bias and inequity can be amplified without safeguards

As Hidden AI code grows and multiplies, it can increase inequities embedded deep within the training data. Health and disease data, for example, may contain a variety of health care disparities, including unequal access to health care; unequal use of diagnostic and treatment protocols; unequal intensity or quantity of treatment; and unequal documentation of data. The very best models and applications can therefore disadvantage patients and populations experiencing health and health care disparities. While developers of models and applications may not intend to include sensitive attributes in the training and validation data, they can unwittingly include a number of different proxies related to socioeconomic status, geography, race/ethnicity, language, disability, and/or high/low health care utilization, including insurance-related barriers to services.

The potential harm to underrepresented patients is not so much that they will receive an inaccurate diagnosis, but that they will receive less than optimal treatment. The potential harm will be greatest for patients with rare conditions, atypical presentations, those with medically complex multisystem disease, and for those with limited history. Patients with variable or fragmented medical care, recent immigrants, patients for whom the digital patient portal is not accessible, and prior limited documentation will likely be incorrectly classified as low risk by clinicians using such models for triage decision-making. The clinicians may not even realize the limitations of the model.

Accountability and liability are still catching up

Accountability in the diagnostic process is a significant issue when clinicians incorporate AI into their work. The clinician is ultimately accountable for the care provided to the patient. However, during the diagnostic process, the clinician will rely on AI recommendations for which the clinician will be held accountable if the recommendations later are found to be inaccurate. However, other parties also may have some degree of accountability, including the health system that purchased and implemented the software as well as the vendor that developed and sold the software to the clinician’s health system. As with many other aspects of clinical diagnosis by AI, it will be difficult to tell an AI recommendation (for which the clinician is accountable) from a recommendation influenced by selected AI settings (for which all parties are accountable) from a recommendation influenced by default settings (for which the vendor may be most accountable but which can be difficult to change).

Documentation and audit trails for the AI and/or automation used to generate the alert(s) should also be considered. There may be an instance where an AI generated alert results in delayed care to a patient. In these situations, it is very important to have a record of the tool’s actions, the timing of those actions, and the underlying data inputs to the tool for these decisions. This information is used for investigation of adverse events, trending of adverse events, and verification of the tool’s decision making processes.

Invisible AI can change the economics of care

Efficiency, cost savings, and increased productivity are drivers of the adoption of AI and machine learning in healthcare. However, while there are some positive gains in terms of efficiency and time savings for clinicians in tasks such as documentation, there are unintended and often serious negative consequences. The savings are gained by automating tasks clinicians do not want to do, but they must then supervise the resulting automation that they do not want to use, do not understand, or cannot fix.

The potential impact of AI on burnout has recently been gaining traction in the model operations community. A large part of this discussion revolves around the idea of “hidden work” that systems extract from their human users. While systems may be able to spend hours or even minutes per day on automation or streamlining, this can also be the result of many “clicks” and “alerts” each sucking a handful of cents of cognitive load. Models can consume a tremendous amount of time to monitor when systems are not well integrated into overall workflow, or where there is no clear owner of responsibilities for model maintenance and tuning.

Who It Affects

Patients

As AI enters the healthcare workflow, users of AI will become the patients. They will feel the delay in getting the care and services they need, be directed to triage based on information they key into a computer instead of speaking to a live physician, and have to wait for an appointment. Triage results will be visible to patients and families, such as whether or not a patient needs to come in to see a doctor for an appointment, whether a physician needs to come take a look at the symptoms right away, or if the patient needs a follow-up. Health risks scores used for care management will identify patients who need to have chronic conditions more closely monitored and patients who need earlier cancer screening. Users of AI will become the patients feeling out the clarity or friendliness of automated text messages used to deliver accurate and timely discharge instructions, after-visit summaries and other messages of patients who are logging into their patient portal.

Do Patients a Need to Know that AI Was Involved in Decisions Made About Them? Patients and families have a need to know that health professionals have provided thoughtful consideration to the treatment plan and have determined the timeliness, adequacy and completeness of the services provided. Patients and families may feel deceived if they are unaware of the extent of automated systems’ involvement in their care. Patients have a need to know the information about the decisions made on their behalf so that they may ask questions, seek additional opinions as warranted, and make decisions involved in trade-offs.

Clinicians across roles

Healthcare professionals may not realize how often they are relying on hidden AI. For primary care teams, this might be an automated coding prompt or automated preventive service reminder and referral. For nurses, it might be automated prioritization during triage as well as automated alerts for possible sepsis, falls risk, and patient deterioration. For radiologists, it might be image pre-reads, worklist prioritization, and automated structured reporting prompts. For pharmacists, it might be medication safety alerts, automated interaction checks, and patient adherence predictions. Automated claim edits and documentation prompts are typically developed by the administrative and revenue cycle staff and used as information for providers and clinical documentation specialists to verify accuracy.

As more AI is integrated into the healthcare workflow, there are new challenges for staff who are tasked with reviewing the outputs. If the data used to train the models is not properly labeled, and the resulting outputs from the models are not highlighted in training materials, clinicians and staff will not know what to look for in the workflow to find the specific alert, score, etc. for which they are responsible. They need to see the value that the AI brings to their clinical workflow. Is that score updating every day on their workflow dashboard, or is it a rare flashing alert? Without proper labeling of AI outputs in the workflow, clinicians do not understand the value of the AI and, as importantly, do not know which outputs to focus on to provide meaningful feedback to the model in order to improve its performance.

SERA is used for risk identification, risk assessment and risk prioritisation. It can be used for risk evaluation in terms of safety and nutritional value and for development of safety & nutritional strategies and measures. Additionally, it can be used for monitoring and evaluation of the impact of the developed strategies and measures.

Health system IT professionals implementing and managing clinical IT including AI are trying to understand the applications, systems and technologies. More often than not, the clinical applications and technologies are supplied by vendors that integrate across multiple applications. Some create information silos within health systems because different data sources, update cycles, configurations and performance monitoring functions need to be managed. Alan notes one health system that he has consulted to on AI initiatives, Alan Health, has implemented AI in multiple applications including the EHR, radiology image management, lab middleware, patient portals, call center software and even a couple of their payers. “A more holistic approach and some coordination and management of these applications from a single source would be very beneficial.”

Quality, Safety and Compliance teams want to know what is going on to manage risk. The basic questions would be: What AI/Machine Learning tools are we currently using, on which patients, for what types of decisions, monitored by whom, and what thresholds would trigger review, re-tuning or system shut down. As these models continue to produce harm and unfair bias, these are increasingly pertinent questions.

Payers and policymakers

Automating payer utilization management, prior authorization and claims review activities can result in cost savings and process improvements for providers. However, most patients and clinicians are unaware that critical claims, services requests or utilization management decisions are made in whole or in part by computer algorithms. Patients and their families, as well as physicians and their staffs, are wasting considerable time and experiencing increasing frustration as automation demands more documentation to process claims correctly. Then, when a claim is denied, patients, families, physicians and staffs encounter significant difficulties in processing an appeal. The process is confusing and lacking in important feedback regarding the reasons for denial and steps to reverse the denied claim.

Regulators need to ensure that innovative technologies do not harm their users while fostering beneficial innovation. Policymakers and industry leaders must promote transparency, fairness and accountability by embedding these values into the system through proper disclosure rules, ongoing monitoring and independent evaluation.

What Changes

1) Require transparency and labeling

Make AI visible. Make it clear where AI has been applied and show what it has done. Clearly highlight the outputs of the AI within the EHR record and operational dashboards. Use simple text to inform the user as to what they can expect from the output and explain how it can be used. Health systems do not need to reveal out their confidential commercial code to do this.

We have included in the downloadable assets to this site, within the relevant zip file, a document that we’ve called “model facts”. This document contains fast facts to the clinicians about the model and the output of the model, including what the model has been used for, what data it’s been trained on, when the model was last validated, some information about the limitations of the model and finally, some information about the suitability of the model for different types and numbers of patients. It is then for the individual clinicians to use their professional judgement to decide whether the output from the model is suitable for a particular patient.

2) Establish human-in-the-loop standards

Identify the types of decisions that require human review. Decisions with very significant impacts on diagnosis, treatment or patient access to necessary healthcare that should include meaningful human review, such as high-stakes triage decisions, safety reminders for certain medications, denial of medically necessary services, or delayed evaluations that may impact timely treatment. Human-in-the-loop reviewers are not fast approvers. They review, they can override the AI, and they document the organization’s rationale for the decision made by the AI.

Auditable Trails: Display influences (AI based and other) on results during system use and allow user to learn from them. Record in an auditable log details such as: showing results generated by AI; user acceptance/override of those results; actions generated by the system; actions taken by the user. Auditable trails can enable learning, system improvement and fair accountability.

3) Invest in local validation and ongoing monitoring

Validate your models on your local patient population. While a model can be well validated and validated in multiple patient populations, validation in different use cases can be unpredictable and may vary significantly based on documentation variability and the prevalence of different diseases within your training data. Performance equity by subgroup (such as age, sex, primary language, level of comorbidity, etc.) must also be evaluated to ensure performance is equitable for vulnerable populations.

Continued monitoring of system performance or “performance drift” as clinical and coding practices change as well as the patient population. Identification of any unintended harm. Monitoring of intended outcomes and measures of accuracy including false positives and false negatives. We want to measure the total number of alerts clinicians see, how they feel the tool is burdensome and how accurate they perceive the tool to be. We want to know where the tool is performing now and how it can be easily re-tuned to perform better. We want to assure that the tool can do no harm and have mechanisms in place to re-tune, disable or remove the tool should it do harm.

4) Strengthen governance, contracts, and workforce training

It is best to initially start within your organization with a single team managing an AI-governance program organization-wide. The function of this governance program would be to manage the “inventory” of models and data sets organization-wide. It would classify the risk for each type of model and/or data set, then work to establish processes to approve, monitor over time and respond to incidents as models and/or data sets go into production. The governance requirements for high risk tools would be high as well. Much like there are high standards of governance around the acquisition, distribution and use of pharmaceuticals and medical devices, there should be similarly high standards of governance around models and data sets.

Recommendations for Writing Better Procurement Requirements for Explainability in Health and Medicine. OHS 2019. #WritingRecommendations. Recommend in contracts that all outputs for which the system is used be explainable and that all decisions be fully documented. Include in contracts appropriate mechanisms for updating system as well as a change log, provenance of data, cyber security measures based on risk for type of and amount of data as well as patient population of interest. Clearly assign monitoring responsibility for system (e.g. system alerts, performance over time). Include in contracts a provision prohibiting independent evaluation of either system or outputs. Include in contracts a provision that all clinical staff have full access to all portions of system and all data.

Please make sure training on the AI recognition and safe use is fully embedded into the workflow of all clinicians and staff who will be using the system. My experience is that best learning occurs in small doses over time, as a series of brief pulses of learning, rather than in one long session. Clinicians and staff need to know what a risk score signifies, when to override it, how to document their decisions with the software, and how to report errors and issues with the model. I want these tools to be useful, not nuisance inducing, and need some feedback loops from the front line to continue to fine tune the systems.

I’ve recently been experimenting with some of the new AI driven features in 3DS Max and I’ll let you know if they worked or failed for me before you throw any more hours into testing them out.

I’ve been trying to find out how an object in a .vsd file is tagged by AI and it appears to use a list of pre-defined smart tags from an earlier version of the file. Quite underwhelming really. I was hoping to see AI improvements that obviously aren’t present in VSD files.

Setting goals and outcomes for the integration of AI into clinical decision making. Providers and organizations might ask these questions before introducing an AI system into their clinical decision making practice: What clinical decision is the system intending to support? What benefits does the system provide, and to whom do those benefits matter most? What potential harms could occur? How will these factors be monitored especially when the system fails and produces an error, particularly inequity and increased health disparities. Who owns the tool clinically and operationally?

Downtime / Safe failure: the system should allow for good care to be provided safely when the AI fails to produce recommendations, with minimal downtime/chaos. There should not be too much automation without good fallbacks.

How clinicians can protect patients right now

The Source team reviewed recommendations that you are responsible for. Are the diagnosis, potential risks and/or default actions coming from a model, rule and/or clinical guideline? Even if you can’t identify a problem that is workable for a work-around, please report any issues you find to the appropriate quality committees/informatics leadership.

AI should be one of many factors that managers consider when making decisions, not simply a model that provides answers. If the model recommendation does not match your clinical judgment, history of the case, and the values / preferences of the patient/s, you should overwrite the model output and document the reason.

Policy direction that supports safety and trust

AI is currently used in healthcare but little is known about this by most people who use the healthcare system. Those who do are interested to know when AI will be used in their care and what benefits they can expect to receive from it, such as access to new services, screening programmes, diagnostic tests and treatments. This information needs to be incorporated into the consent process, the patient portal summaries and the plain language summaries. All patient disclosure about the use of AI should clearly set out its capabilities and limitations and patients should have ready access to questioning and challenging the use of AI and the subsequent recommendations which flow from it.

The opportunity is to lift the state of evaluation and reporting of health innovation. As Hidden AI is tested, assessed for concerns around equity, and monitored for incidents, we can work to ensure the “invisible” AI is safe and reliable throughout the health system for everyone.

Looking ahead

Will invisible AI create problems for organisations, or bring massive productivity gains? The answer will depend on many factors, including the culture and governance of an organisation. By labelling, validating, monitoring and supporting the training of AI tools organisations can save much time and derive many benefits including consistency. But as AI tools become invisible, inadequately tested and treated as “set and forget”, they can undermine human judgement, create unfairness and expose organisations to significant legal and reputational risk.

Long term, we think that accountability for use of AI is important. But for now, Health Systems need to be able to set goals and measures around accountability of AI as just one of many clinical tools they can use. Those clinicians who are using the AI will also need to be informed supervisors of the data that went into training the algorithm. Patients have a right to an explanation of how AI was used in their care and how AI can be used to support informed decision making.

As AI technology is being increasingly applied in health care, it is important to implement it safely and responsibly. For this reason, the AI must be transparent, interim-checkable, traceable, and human-centered.

References:

https://www.healthit.gov/buzz-blog/artificial-intelligence/ten-ways-ai-shaping-future-healthcare

ShareFacebook
AI-transparencyclinical-adoption

One story a day

The story of the day, in your inbox

One health journey each morning — no advice, no alarm, just company for the road.

Read next