Medical Imaging AI Safety: Building Trustworthy, Clinically Safe, and Governable Radiology AI Systems
From algorithm validation and bias control to human oversight, workflow integration, cybersecurity, and continuous monitoring across the medical imaging AI lifecycle
Author: Dr. SB Lee
Medical imaging artificial intelligence has moved beyond the experimental stage.
Algorithms can now assist with image reconstruction, segmentation, detection, classification, quantitative analysis, triage, and clinical decision support. The technological trajectory is impressive. But clinical deployment introduces a question that cannot be answered by accuracy metrics alone:
What makes a medical imaging AI system safe?
A model can achieve excellent performance in a development dataset and still create clinical risk after deployment.
It may encounter a different patient population. The scanner manufacturer may change. Acquisition protocols may differ. Image quality may deteriorate. Disease prevalence may shift. A hospital may modify its workflow. An algorithm may generate too many false-positive alerts. A subtle false-negative result may be overlooked because clinicians assume the AI has already checked the examination.
Safety therefore cannot be treated as a property of the neural network alone.
It is a property of the entire clinical system surrounding the AI.
Current international guidance increasingly emphasizes lifecycle management, transparency, fairness, robustness, traceability, usability, and explainability rather than isolated algorithmic performance. The 2025 FUTURE-AI consensus guideline, for example, organizes trustworthy healthcare AI around six principles—fairness, universality, traceability, usability, robustness, and explainability—and addresses the lifecycle from design and validation through deployment and monitoring.
For medical imaging, this distinction is particularly important.
A radiology AI system does not operate in a laboratory vacuum. It receives images from acquisition equipment, interacts with PACS and workflow systems, produces outputs for radiologists, and may ultimately influence patient management.
The safety question is therefore much larger than:
“How accurate is the model?”
The more clinically meaningful question is:
“Can this AI system reliably contribute to patient care within the real-world environment in which it will be used?”
1. Why Is Medical Imaging AI Safety Different from Ordinary Software Safety?
Medical imaging AI operates at the intersection of software engineering, medical diagnosis, human decision-making, and patient safety.
A conventional software failure may cause inconvenience, financial loss, or interruption of a business process.
An AI failure in radiology can potentially contribute to a missed diagnosis, unnecessary investigation, delayed treatment, inappropriate prioritization, or misplaced clinical confidence.
The risk becomes particularly complex because AI errors are not always obvious.
Consider three situations.
Scenario 1: False Negative
An AI system fails to identify a pulmonary nodule.
If the radiologist independently detects the nodule, the AI error may have little clinical consequence.
But if the workflow has encouraged excessive reliance on the algorithm, the same error may contribute to a missed finding.
Scenario 2: False Positive
An algorithm flags a benign pulmonary opacity as suspicious.
The radiologist recognizes the false positive, but repeated unnecessary alerts gradually increase cognitive burden.
Over time, clinicians may experience alert fatigue.
Scenario 3: Workflow Failure
The algorithm itself performs correctly, but its result arrives after the radiologist has already finalized the report.
Technically, the model worked.
Clinically, the system failed.
This illustrates one of the central principles of Medical Imaging AI Safety:
Algorithmic correctness and clinical safety are not synonymous.
FIGURE 1. Medical Imaging AI Safety as an End-to-End Clinical System.
Safety considerations extend from image acquisition and data transfer to AI inference, radiologist verification, reporting, and downstream clinical decision-making.
2. What Are the Major Safety Risks in Medical Imaging AI?
Medical imaging AI safety can be organized into several interconnected domains.
| Safety domain | Typical risk | Clinical consequence |
|---|---|---|
| Data quality | Missing, corrupted, or low-quality images | Incorrect AI output |
| Dataset bias | Underrepresentation of patient groups | Unequal performance |
| Distribution shift | Different scanner/protocol/population | Performance degradation |
| Model error | False positive/negative prediction | Diagnostic risk |
| Workflow integration | Poor timing or routing | Missed clinical opportunity |
| Human factors | Automation bias | Excessive reliance on AI |
| Explainability | Unclear basis for output | Difficult verification |
| Cybersecurity | Manipulation or unauthorized access | Patient and system risk |
| Model drift | Performance changes over time | Silent degradation |
| Governance | Poor ownership and oversight | Uncontrolled clinical risk |
These risks are not independent.
A change in scanner protocol can alter the input distribution. That can affect model performance. A decline in performance can increase false positives or false negatives. If the change is not monitored, clinicians may remain unaware of the degradation.
The result is a systemic safety problem.
3. Why Is Validation the Foundation of Medical Imaging AI Safety?
Validation is the process through which developers and healthcare organizations determine whether an AI system performs adequately for its intended clinical purpose.
But validation should not be reduced to a single accuracy number.
A model may report:
- sensitivity,
- specificity,
- accuracy,
- area under the receiver operating characteristic curve,
- precision,
- negative predictive value,
- calibration,
- segmentation metrics,
yet these metrics may not answer the most important clinical question:
Will the model remain useful in the environment where it will actually be used?
This is why validation should occur at multiple levels.
Technical validation
Does the algorithm function as designed?
Internal validation
Does it perform adequately on data related to development?
External validation
Does performance remain acceptable on data from a different institution, population, scanner environment, or acquisition protocol?
Clinical validation
Does the AI meaningfully support the intended clinical task?
Workflow validation
Does introducing the AI into clinical practice improve or disrupt the workflow?
Post-deployment validation
Does performance remain acceptable after implementation?
The emerging literature increasingly treats transparency and methodological reporting as essential components of trustworthy AI evaluation. STARD-AI, published in Nature Medicine in 2025, specifically addresses reporting requirements for AI-centered diagnostic accuracy studies, including dataset practices, algorithm evaluation, bias, fairness, and generalizability.
4. Dataset Bias: Can an Accurate Model Still Be Unsafe?
Yes.
One of the most important concepts in Medical Imaging AI Safety is that high average performance does not guarantee equitable performance.
Suppose a model is trained predominantly using images from one geographic region, one healthcare system, or a limited range of scanner environments.
Its performance may be excellent within that environment.
But clinical imaging is heterogeneous.
Patients differ in:
- age,
- sex,
- body habitus,
- disease prevalence,
- comorbidities,
- anatomy,
- ethnicity and population characteristics,
- imaging protocol,
- scanner manufacturer,
- reconstruction algorithm,
- field strength,
- image quality.
If these factors are poorly represented during development, the model may behave differently in another environment.
This is not merely a statistical issue.
It is a patient-safety issue.
Therefore, AI evaluation should ask not only:
“How well does the model perform?”
but also:
“For whom does it perform well?”
and:
“Under which imaging conditions does performance deteriorate?”
The FDA's current transparency principles for machine-learning-enabled medical devices explicitly emphasize communicating intended users, environments, target populations, performance, limitations, biases, and data-characterization gaps.
5. Distribution Shift: What Happens When the Real World Changes?
Medical imaging environments are not static.
A hospital may replace a CT scanner.
A protocol may change.
A reconstruction algorithm may be upgraded.
A new MRI sequence may be introduced.
Patient demographics may evolve.
The prevalence of a disease may change.
A model that was validated under the original conditions may encounter a different statistical environment after deployment.
This phenomenon is commonly described as distribution shift.
The dangerous characteristic of distribution shift is that the model may continue producing outputs without an obvious technical failure.
There may be no system crash.
No error message.
No broken software.
The model simply becomes less reliable.
This is why safe Medical Imaging AI requires continuous monitoring.
6. Why Model Drift Monitoring Matters
Model drift refers broadly to changes that can reduce the reliability of an AI system over time.
Monitoring should therefore not stop when an AI product is installed.
A mature monitoring strategy can examine:
- input-data characteristics,
- image quality,
- scanner distribution,
- acquisition protocols,
- output distributions,
- confidence patterns,
- false-positive rates,
- false-negative signals,
- subgroup performance,
- disagreement between AI and clinicians,
- system latency,
- workflow completion,
- model version.
The important conceptual shift is:
AI deployment is not the end of validation. It is the beginning of operational surveillance.
The FDA's January 2025 draft guidance on AI-enabled medical devices similarly frames safety and effectiveness across the total product lifecycle and includes recommendations concerning design, development, maintenance, documentation, transparency, bias, and ongoing monitoring.
FIGURE 2. Continuous Medical Imaging AI Lifecycle Management.
Deployment is followed by monitoring, drift detection, revalidation, controlled updating, and eventual retirement when safety or clinical utility can no longer be assured.
7. Why Human Oversight Remains Essential
Medical imaging AI should generally be understood as a component of a clinical decision-making system rather than an isolated replacement for professional judgment.
This does not mean that every AI application must operate identically.
The required level of human oversight depends on:
- intended use,
- clinical risk,
- degree of automation,
- reversibility of error,
- severity of potential harm,
- available confirmatory evidence,
- regulatory status,
- workflow context.
A low-risk quantitative tool and an autonomous diagnostic system should not be governed in exactly the same way.
Human oversight should therefore be risk proportional.
The clinician must understand:
- what the AI is intended to do;
- what information it receives;
- what output it produces;
- what the output does not mean;
- when the AI should not be trusted;
- how to override the output.
The FDA's transparency framework specifically emphasizes the performance of the human-AI team and communication of information necessary for users to understand intended use, workflow, benefits, risks, limitations, and model logic when available.
8. Automation Bias: When AI Becomes Too Persuasive
One of the most subtle risks is not an incorrect algorithm.
It is an incorrect human response to a correct or incorrect algorithm.
Automation bias occurs when people give excessive weight to automated recommendations.
In radiology, this could occur when an AI output appears authoritative and the reader unconsciously adjusts their interpretation toward it.
For example, an AI system may label an examination as low risk.
A radiologist may initially notice a subtle abnormality but subsequently discount it because the algorithm did not flag the finding.
This creates an important safety principle:
AI should inform clinical reasoning, not replace clinical reasoning unless the system has been specifically validated and authorized for an appropriate autonomous use case.
The safest interface is therefore not necessarily the one that displays the most information.
It is the one that supports appropriate human judgment at the right moment.
9. Explainability: Does Every AI System Need to Explain Its Decision?
Explainability is frequently discussed as though it were a single technical property.
In clinical practice, it is more useful to think about actionable transparency.
A radiologist may need to know:
- what task the algorithm performs;
- what population it was validated on;
- what its output means;
- what confidence or uncertainty information is available;
- what limitations exist;
- whether the image quality is acceptable;
- whether the current examination is within the intended use.
A heatmap may sometimes be useful.
But a visually compelling heatmap does not automatically explain why an algorithm made a decision.
Therefore, explainability should be evaluated according to its clinical purpose.
The goal is not simply to make a neural network “look understandable.”
The goal is to give the clinician sufficient information to use the AI output safely.
10. Medical Imaging AI Safety Begins with the Imaging Data
The AI lifecycle starts before inference.
Every layer introduces potential risk.
For example:
Acquisition risk
Incorrect protocol or inadequate image quality.
Data-transfer risk
Missing or corrupted studies.
Preprocessing risk
Incorrect resampling, cropping, normalization, or orientation.
Inference risk
Algorithmic error.
Integration risk
Incorrect patient-study association.
Workflow risk
Output delivered too late.
Interpretation risk
AI output misunderstood by the clinician.
Reporting risk
Important AI findings fail to reach the final report.
This is why Medical Imaging AI Safety should be viewed as end-to-end clinical systems engineering.
A typical clinical
architecture connecting imaging modalities, PACS/VNA infrastructure, AI
orchestration, inference services, radiologist workflows, structured reporting,
and the electronic health record.
11. Why Cybersecurity Is Part of AI Safety
Medical AI safety is also cybersecurity safety.
A clinical AI system interacts with:
- DICOM data,
- PACS,
- cloud or on-premise infrastructure,
- inference servers,
- APIs,
- identity systems,
- EHR systems,
- reporting platforms,
- network infrastructure.
A security compromise could affect confidentiality, integrity, or availability.
The most important concern for AI safety is not only unauthorized access.
It is also data and output integrity.
If the input image is manipulated, the AI may generate a clinically misleading result.
If the AI output is modified or incorrectly associated with another patient, the clinical consequences can be serious.
Therefore:
Cybersecurity controls are part of clinical safety controls.
Medical AI governance should consider authentication, authorization, encryption, audit logging, system segmentation, vulnerability management, secure software development, and incident response as components of the broader patient-safety architecture.
12. How Should Hospitals Govern Medical Imaging AI?
A hospital deploying one AI model may be able to manage it informally.
Managing dozens or hundreds of algorithms is fundamentally different.
An enterprise AI governance program should define:
- who owns each algorithm;
- intended clinical use;
- approved users;
- validation status;
- regulatory status;
- model version;
- data requirements;
- known limitations;
- monitoring requirements;
- incident-reporting pathway;
- update/change-management process;
- retirement criteria.
This converts AI from an unmanaged software purchase into a governed clinical capability.
A useful enterprise concept is an AI model registry.
Each model can have a lifecycle record containing:
| Governance element | Example |
|---|---|
| Model ID | Unique enterprise identifier |
| Clinical purpose | Intended use |
| Modality | CT / MRI / X-ray/ultrasound |
| Version | Current production version |
| Validation | Internal/external/clinical |
| Population | Intended population |
| Performance | Validated metrics |
| Limitations | Known failure modes |
| Regulatory status | Applicable status |
| Owner | Responsible clinical/technical team |
| Monitoring | Defined KPIs |
| Incident history | Safety events |
| Update policy | Change-control process |
| Retirement | Decommissioning criteria |
The purpose is accountability.
An algorithm without an owner is a clinical risk.
FIGURE 4. Enterprise Governance Framework for Medical Imaging AI.A
governed AI ecosystem maintains traceability, accountability, validation
status, monitoring requirements, incident history, and lifecycle decisions for
deployed algorithms
13. What Should Happen When AI and the Radiologist Disagree?
Disagreement is not automatically evidence that the AI is wrong.
Nor is it evidence that the radiologist is wrong.
The correct response depends on the clinical context.
A mature system should make disagreement observable.
For high-risk applications, organizations may establish escalation pathways for:
- repeated AI-radiologist disagreement;
- unexpected false-negative patterns;
- unexpected false-positive clusters;
- subgroup performance concerns;
- image-quality failures;
- unusual output distributions;
- suspected cybersecurity events.
The objective is not to eliminate disagreement.
The objective is to learn from disagreement.
This turns the clinical environment into a feedback mechanism for safety surveillance.
14. What Does a Safe AI Clinical Workflow Look Like?
This architecture emphasizes a critical distinction:
The AI output is an intermediate clinical artifact, not automatically the final diagnosis.
For certain validated applications, the role may be more autonomous.
But autonomy must be explicitly established through evidence, intended use, risk assessment, validation, and appropriate governance.
It should never be assumed merely because the algorithm performs well in a benchmark dataset.
15. How Should AI Performance Be Monitored After Deployment?
A hospital should define monitoring indicators before deployment.
Potential categories include:
Technical indicators
- inference availability;
- processing latency;
- system failures;
- data-transfer failures;
- invalid inputs.
Model indicators
- output distribution;
- confidence distribution;
- performance estimates where ground truth becomes available;
- subgroup performance;
- drift indicators.
Clinical indicators
- radiologist-AI agreement;
- false-positive burden;
- false-negative signals;
- downstream testing;
- clinically significant missed findings.
Workflow indicators
- time to AI result;
- percentage of studies processed;
- percentage of outputs reviewed;
- alert response;
- report integration.
Safety indicators
- incidents;
- near misses;
- unexpected behavior;
- complaints;
- cybersecurity events.
The exact metrics should be determined by the clinical application.
There is no universal monitoring dashboard that is appropriate for every medical imaging AI system.
16. What Happens When a Model Changes?
AI updates introduce another safety challenge.
A software update may change:
- model weights,
- preprocessing,
- threshold,
- output format,
- inference infrastructure,
- intended population,
- performance characteristics.
Therefore, version control is essential.
A production AI system should have a traceable relationship between:
patient examination → model version → AI output → clinician interpretation → clinical action
This traceability becomes especially important when investigating an unexpected clinical event.
The FUTURE-AI framework explicitly includes traceability as one of its six core principles and addresses governance across the AI lifecycle.
17. Can Medical Imaging AI Ever Be Completely Safe?
No clinical technology is completely risk-free.
The goal of Medical Imaging AI Safety is therefore not to promise zero errors.
The objective is to:
- identify foreseeable risks;
- reduce preventable failures;
- detect failures early;
- prevent unsafe automation;
- maintain human oversight;
- monitor performance;
- document incidents;
- learn from failures;
- update controls;
- retire systems that no longer meet their intended safety requirements.
This is fundamentally a risk-management discipline.
18. What Does Trustworthy Medical Imaging AI Look Like?
Trustworthy AI is not simply accurate AI.
A trustworthy Medical Imaging AI system should be:
Fair
Performance should be evaluated across relevant populations and clinical contexts.
Universal
The system should have clearly defined applicability and should not silently be used outside its validated environment.
Traceable
The organization should know what model version produced what output and under which conditions.
Usable
The system should fit the clinical workflow rather than forcing clinicians to work around it.
Robust
The system should be appropriately tested against foreseeable variation and failure modes.
Explainable
Users should receive meaningful information about intended use, limitations, outputs, and—where appropriate—the basis of decisions.
These six concepts correspond closely to the FUTURE-AI international consensus framework.
A 2026 review specifically focused on trustworthy AI in medical imaging translates the FUTURE-AI principles into imaging-specific practices across design, development, validation, deployment, and monitoring.
19. The Future: From Model-Centric AI to Safety-Centric AI
The next stage of Medical Imaging AI will not be defined only by larger models or higher benchmark performance.
The more consequential development may be the transition from model-centric AI to system-centric AI.
In the model-centric paradigm:
“How good is the algorithm?”
In the system-centric paradigm:
“How safely does the complete clinical system perform?”
That system includes:
- data;
- imaging equipment;
- PACS;
- AI orchestration;
- inference;
- workflow;
- radiologists;
- reporting;
- EHR;
- cybersecurity;
- governance;
- monitoring.
This is particularly important as hospitals move toward multimodal AI and foundation-model-based systems.
The WHO has emphasized that healthcare AI requires governance and oversight designed to promote safety, equity, and accountability, while its more recent guidance addresses emerging large multimodal models and their implications for healthcare.
The future of safe imaging AI therefore depends not only on better algorithms, but also on better institutions.
20. Radiologist's Practical Checklist for Medical Imaging AI Safety
Before adopting an AI system, ask:
Clinical purpose
- What exact clinical problem does the AI solve?
- Is the intended use clearly defined?
- What decisions may be influenced by its output?
Data
- What population was used for development?
- What scanners and protocols were represented?
- Are image-quality limitations known?
Validation
- Was the system externally validated?
- Was validation performed in a population similar to ours?
- Were clinically meaningful endpoints assessed?
Bias
- Were relevant subgroups evaluated?
- Are known data gaps documented?
Workflow
- Where does the AI output appear?
- When does it appear?
- Who reviews it?
- What happens when the AI is unavailable?
Human oversight
- Can clinicians override the output?
- Are limitations visible?
- Is automation bias addressed?
Cybersecurity
- How is data protected?
- Are access and audit controls implemented?
- What happens during a security incident?
Monitoring
- What performance indicators are monitored?
- How is model drift detected?
- Who receives safety alerts?
Governance
- Who owns the model?
- Who approves updates?
- When is the model revalidated?
- What triggers retirement?
21. Key Takeaways
Medical Imaging AI Safety is not synonymous with model accuracy.
A clinically safe AI ecosystem requires:
- robust validation;
- representative datasets;
- bias assessment;
- external evaluation;
- workflow validation;
- human oversight;
- meaningful transparency;
- cybersecurity;
- model version control;
- continuous monitoring;
- incident management;
- organizational accountability.
The most important shift is conceptual.
AI safety should be designed into the clinical workflow rather than added after deployment.
The safest radiology AI system is therefore not necessarily the most sophisticated algorithm.
It is the system in which the algorithm, data, infrastructure, clinicians, workflow, governance, and monitoring work together in a controlled and auditable manner.
For radiologists and healthcare organizations, the central question is no longer simply whether AI can recognize an abnormality.
It is whether the entire clinical environment can use that capability safely, consistently, transparently, and responsibly.
22. Conclusion
Medical imaging is one of the most data-rich areas of medicine, making it an important environment for clinical AI.
But imaging AI also illustrates a fundamental truth about healthcare technology:
A model becomes a clinical technology only when it enters a clinical system.
At that point, algorithmic performance becomes only one component of safety.
The real safety architecture includes data quality, validation, fairness, robustness, explainability, workflow design, human oversight, cybersecurity, lifecycle governance, and continuous monitoring.
International guidance is increasingly moving in this direction. WHO emphasizes safe, ethical, equitable, and accountable AI governance, while FDA and international medical-device guidance increasingly address transparency, bias, lifecycle risk management, and good machine-learning practice.
For medical imaging, this means that the future should not be defined by AI replacing the radiologist.
A more realistic and clinically responsible objective is to build trustworthy human-AI systems in which machines perform what they are validated to perform, clinicians retain appropriate judgment, and institutions continuously monitor whether the combined system remains safe.
That is the foundation of Medical Imaging AI Safety.
TABLES
Table 1. Medical Imaging AI Safety Risk Framework
| Domain | Primary Question | Safety Control |
|---|---|---|
| Data | Are inputs appropriate? | Data-quality and eligibility checks |
| Bias | Does performance vary across groups? | Subgroup validation |
| Robustness | Does performance survive real-world variation? | External validation and stress testing |
| Workflow | Does AI fit clinical practice? | Workflow simulation and clinical evaluation |
| Human factors | Could users overtrust AI? | Human-centered interface design |
| Explainability | Can users understand appropriate use? | Transparent intended use and limitations |
| Cybersecurity | Can data/output integrity be protected? | Security controls and auditability |
| Drift | Does performance change over time? | Continuous monitoring |
| Governance | Who is accountable? | Model registry and lifecycle ownership |
Table 2. Validation Across the AI Lifecycle
| Stage | Main Objective |
|---|---|
| Development | Establish technical performance |
| Internal validation | Test generalization within related data |
| External validation | Test performance in different environments |
| Clinical validation | Determine clinical utility and safety |
| Workflow validation | Assess integration with clinical practice |
| Deployment | Controlled introduction |
| Post-market monitoring | Detect performance and safety changes |
| Revalidation | Confirm continued suitability after major changes |
| Retirement | Remove systems that no longer meet requirements |
Table 3. AI vs Clinical System Safety
| Question | Model-Centric View | System-Centric View |
|---|---|---|
| Accuracy | Is the model accurate? | Does it improve clinical performance? |
| Data | Is the training set large? | Is the operational population appropriate? |
| Workflow | Does inference work? | Does the result arrive at the right time? |
| Human factors | Is output understandable? | Does the interface support safe judgment? |
| Monitoring | Does the server run? | Is clinical performance maintained? |
| Governance | Is the product approved? | Is the entire lifecycle controlled? |
FAQ
What is Medical Imaging AI Safety?
Medical Imaging AI Safety refers to the systematic management of risks associated with using artificial intelligence in medical imaging. It includes model validation, bias assessment, robustness, workflow integration, human oversight, transparency, cybersecurity, monitoring, governance, and lifecycle management.
Why is AI accuracy not enough for safe radiology AI?
Accuracy is usually measured under specific test conditions. Real clinical environments vary by patient population, scanner, protocol, image quality, disease prevalence, and workflow. A model can therefore perform well in testing while behaving differently after deployment.
What is model drift in medical imaging AI?
Model drift describes changes in an AI system's operational environment or performance over time that may reduce reliability. Scanner upgrades, protocol changes, population shifts, and changes in disease prevalence can all affect how an AI system performs.
How can hospitals reduce bias in medical imaging AI?
Hospitals should evaluate whether the AI was validated across relevant patient populations, imaging environments, scanners, and clinical settings. Subgroup performance and data limitations should be documented rather than relying only on aggregate performance.
Should radiologists always verify AI results?
The appropriate level of human oversight depends on the intended use, clinical risk, validation evidence, regulatory framework, and degree of automation. For many decision-support applications, radiologist review remains an important safety mechanism.
Why is cybersecurity part of AI safety?
AI depends on clinical data and connected infrastructure. Compromise of imaging data, patient identity, AI outputs, or system availability can create clinical as well as information-security risks.
How often should medical imaging AI be revalidated?
There is no universal interval appropriate for every AI system. Revalidation should be risk-based and should consider significant model updates, changes in patient populations, imaging equipment, acquisition protocols, workflow, or observed performance.
REFERENCES
- Lekadir K, Frangi AF, Porras AR, et al. FUTURE-AI: international consensus guideline for trustworthy and deployable artificial intelligence in healthcare. BMJ. 2025;388:e081554. doi:10.1136/bmj-2024-081554.
- Kondylakis H, Osuala R, Puig-Bosch X, et al. A Review of Methods for Trustworthy AI in Medical Imaging: The FUTURE-AI Guidelines. IEEE Journal of Biomedical and Health Informatics. 2026;30(3):2299-2315. doi:10.1109/JBHI.2025.3614546.
- Sounderajah V, Guni A, Liu X, et al. The STARD-AI reporting guideline for diagnostic accuracy studies using artificial intelligence. Nature Medicine. 2025;31(10):3283-3289. doi:10.1038/s41591-025-03953-8.
- Rivera SC, Liu X, Chan AW, et al. Guidelines for clinical trial protocols for interventions involving artificial intelligence: the SPIRIT-AI Extension. Nature Medicine. 2020;26:1351-1363. doi:10.1038/s41591-020-1037-7.
- Cruz Rivera S, Liu X, Chan AW, et al. Reporting guidelines for clinical trials of artificial intelligence interventions: the SPIRIT-AI and CONSORT-AI guidelines. Trials. 2020;21:912. doi:10.1186/s13063-020-04951-6.
- World Health Organization. Ethics and governance of artificial intelligence for health. WHO Guidance. 2021.
- World Health Organization. Ethics and governance of artificial intelligence for health: Guidance on large multi-modal models. WHO. 2025.
- U.S. Food and Drug Administration. Transparency for Machine Learning-Enabled Medical Devices: Guiding Principles. FDA.
- U.S. Food and Drug Administration. Artificial Intelligence-Enabled Device Software Functions: Lifecycle Management and Marketing Submission Recommendations. Draft Guidance. 2025.
- U.S. FDA, Health Canada, MHRA/IMDRF. Good Machine Learning Practice for Medical Device Development: Guiding Principles. Updated international guidance, 2025.
Comments
Post a Comment