The Ethics of Integration: Why Healthcare AI Must Be Evaluated Within Clinical and Research Workflows

Author

Maame Yaa A. B. Yiadom, MD, MPH

Publish date

The Ethics of Integration: Why Healthcare AI Must Be Evaluated Within Clinical and Research Workflows
Topic(s): Artificial Intelligence Editorial-AJOB Ethics

This editorial appears in the August Issue of the American Journal of Bioethics

The ethical discourse surrounding artificial intelligence (AI) in healthcare has largely focused on algorithmic performance, bias, privacy, transparency, and explainability. These concerns remain critically important. However, as AI applications increasingly move from development environments into clinical and research operations involving patient health and medical outcomes, a complementary ethical challenge is emerging: how AI changes the workflows through which decisions are made. I argue that ethical evaluation must therefore expand beyond algorithmic performance to include what might be termed workflow integration ethics, the study of how AI influences decisions, behaviors, and outcomes within healthcare delivery systems.

The articles in this issue by Rentzepis et al., Char et al., and Hatherley et al. each examine different AI applications: research recruitment, clinical summarization, and federated learning. Yet collectively they reveal a broader concern. The ethical risks associated with AI frequently arise not from the model itself, but from the interaction between the model, the humans using it, and the organizational systems into which it is deployed. Healthcare organizations often evaluate AI as though it were a standalone technology. In practice, AI functions as an embedded participant in complex sociotechnical systems with pre-established operating goals, workflows, and accountability structures. Healthcare AI is not deployed into a vacuum; it is deployed into systems and practices of healthcare delivery. Ethical assessment must therefore extend beyond evaluating algorithms to evaluating their role and performance within care delivery workflows. As a result, the ethical unit of analysis must expand from algorithms to the sociotechnical systems in which they operate.

From Algorithm Ethics to Workflow Ethics

The ethical concerns raised by AI recruitment tools are often framed as questions of fairness, representativeness, and privacy. An AI recruitment model may alter which patients are approached about clinical trial participation. Similarly, concerns about generative AI summarization focus on accuracy, hallucinations, and clinician trust. A summarization tool may shape how clinicians understand a patient’s history. Federated learning raises questions regarding transparency, accountability, and data governance. A federated learning model may influence predictions despite uncertainty regarding the data on which it was trained.

While these concerns appear distinct, they share a common characteristic: they question how information moves through healthcare delivery systems and ultimately influences human decisions and the experience and health outcomes of patients affected. In each case, the ethical concern is not solely whether the model functions correctly. The concern is whether healthcare organizations can understand, monitor, and govern the workflow consequences that follow from AI-generated outputs.

AI as A Workflow Intervention

Healthcare organizations routinely evaluate interventions that alter clinical workflows. New staffing models, order sets, triage protocols, and quality-improvement initiatives are all assessed according to their effects on care delivery processes and patient outcomes. AI should be treated similarly. An AI tool rarely acts independently. Instead, it changes the timing, sequence, prioritization, or content of decisions made by humans making judgements and performing tasks. These workflow effects may ultimately be more consequential than the model’s computational performance characteristics. For example, an AI recruitment tool that identifies eligible patients more efficiently may increase trial enrollment. However, it may also alter who is approached, when they are approached, and how recruitment resources are allocated across populations. That influence on enrollment can alter the outcomes of the research study. Similarly, a clinical summarization tool may produce technically accurate summaries while subtly changing clinician attention, documentation practices, or information-seeking behaviors in ways that change the quality or nature of care received by their patients. The ethical question therefore becomes how AI reshapes human decision-making within operational systems and whether healthcare organizations can responsibly balance the benefits of those changes against their consequences

Accountability Requires Observability

The articles by Char et al. and Hatherley et al. appropriately highlight concerns regarding transparency and performance visibility. This is not merely a technical problem; it is also an operational and governance problem. Healthcare organizations cannot govern what they cannot observe. Historically, clinical governance has depended upon the ability to reconstruct decision pathways, identify failures, and implement corrective actions. AI introduces new layers of complexity into these processes. Recommendations may be generated external to the organization while being shared within it, training data may be inaccessible, and model outputs may influence workflows in ways that are difficult to detect retrospectively and therefore require prospective evaluation. As a result, healthcare organizations require mechanisms that enable ongoing observation of AI behavior in operational settings. Ethical oversight should not end at deployment. It should include prospective monitoring of workflow effects, user interactions, simulated versus real-world performance, overrides, functional shift and drift, unintended consequences, and differential impacts across patient populations. This requirement is particularly important because many harms emerge only after implementation.

The Importance of human-AI Teaming

A recurring assumption within healthcare AI discussions is that ethical concerns can be addressed by maintaining a “human in the loop.” While human oversight remains important, simply inserting a human reviewer may be insufficient. The relevant question is not whether a human remains involved, but whether the human-AI team functions effectively. In many healthcare settings, the more accurate description is not a human-in-the-loop system, but a human workflow with “AI in the loop.”

Research across healthcare operations demonstrates that outcomes depend upon communication structures, role clarity, feedback mechanisms, and organizational culture. Similar principles should guide AI implementation where AI functions as a member of a care delivery team rather than an independent agent. Organizations should therefore evaluate whether: 1) users understand AI outputs, 2) workflows sequence and objectives are understood well enough to permit meaningful review, 3) accountability for action and workflow outcomes remain clear, 4) users are empowered to override, adjust, or reevaluate recommendations, 5) feedback loops exist to improve performance. These factors influence safety and effectiveness as much as algorithmic accuracy. In this regard, the integration of AI into healthcare workflows may have more in common with the implementation of medical devices and pharmaceuticals than is often acknowledged. The technology itself matters, but so do the systems, training, governance structures, and human behaviors that determine its real-world impact.

Ethical Success Requires Implementation Science

The next generation of healthcare AI ethics should incorporate principles from implementation science, health services research, and systems engineering. Healthcare has repeatedly demonstrated that interventions with strong efficacy can fail when implemented poorly. Conversely, interventions with modest technical advantages may achieve substantial impact when integrated effectively into clinical operations. Ethical evaluation should therefore examine not only whether an AI model works, but whether healthcare organizations possess the governance structures necessary to deploy it responsibly. This shift has important implications for regulators, health systems, sponsors, and institutional review boards. Questions regarding performance, bias, and privacy remain necessary. However, they should be accompanied by questions regarding workflow integration, organizational accountability, monitoring plans, and mechanisms for continuous learning. Ethical success will depend not only on what an AI model does, but on how healthcare organizations choose to implement, monitor, and govern it.

Conclusion

The papers by Rentzepis et al., Char et al., and Hatherley et al. collectively illustrate a transition occurring across healthcare AI. The central ethical challenge is no longer simply evaluating algorithms. It is understanding how AI becomes embedded within the workflows through which healthcare and research are conducted. As AI moves from experimentation to operational deployment, ethical oversight must expand accordingly. The future of ethical AI in healthcare will depend not only on building better models, but also on building more accountable systems in which those models operate. Healthcare organizations should therefore evaluate AI in the same way they evaluate any other intervention intended to improve care delivery; not by what it predicts or computes alone, but by how it changes real-world activities, decisions, and outcomes.

Maame Yaa A. B. Yiadom, MD, MPH

We use cookies to improve your website experience. To learn about our use of cookies and how you can manage your cookie settings, please see our Privacy Policy. By closing this message, you are consenting to our use of cookies.