The Future of Clinical Judgment: Diagnostic Sovereignty in the Age of AI
How artificial intelligence is reshaping medical cognition, training, and professional autonomy
Imagine this…
In the not-so-distant future, a patient arrives at the emergency department with chest pain that began 30 minutes ago. He reports that the pain started after a heated argument with his boss.
A third-year resident evaluates him and opens the clinical AI assistant.
The system suggests structured questions:
When did the pain begin?
Does it radiate?
Any cardiovascular history?
The patient answers. Routine labs are ordered. The ECG shows no clear changes. Biomarkers are still within normal range.
After integrating the data, the AI concludes:
“90% probability of anxiety crisis triggered by occupational stress. Recommendation: discharge with outpatient follow-up.”
The resident validates the recommendation. The patient goes home.
Hours later, an ambulance returns. A man in cardiac arrest. A paramedic is performing chest compressions as they rush into the resuscitation bay.
The resident recognizes him.
It is the same patient.
The senior physician enters, reviews the case quickly, and says in a firm voice:
— It was an evolving myocardial infarction. He should have been admitted.
The resident responds, almost defensively:
— I followed the AI’s recommendation.
The senior looks at him:
— In medicine, nothing is absolute. You must learn to suspect.
Minutes later, the patient dies.
And the question is no longer technological.
It is epistemic.
The large-scale integration of artificial intelligence into medicine promises efficiency, precision, and reduction of human error. And in many cases, it delivers.
But we are beginning to observe something more subtle and more structural: a transformation in the architecture of clinical reasoning itself.
Between 2023 and 2026, research from institutions such as MIT, Harvard, and Stanford has explored how the use of large language models and clinical decision support systems does not merely enhance short-term task performance, but may also reshape cognitive patterns, memory mechanisms, and metacognitive disposition.
To analyze this phenomenon, I propose an operational concept:
Systemic epistemic fragility: the progressive reduction in a professional system’s collective capacity to:
Detect algorithmic error when it occurs.
Generate diagnostic hypotheses outside the space suggested by AI.
Tolerate and manage uncertainty without computational support.
Transmit tacit knowledge intergenerationally through deliberate practice.
If this fragility consolidates, the problem will not be that AI makes mistakes.
The problem will be that we no longer know how to recognize when it does.
A second, less discussed risk concerns the homogenization of clinical judgment.
In machine learning, mode collapse describes a phenomenon in which a generative model converges toward average solutions and loses diversity in its outputs.
In medicine, we may face an analogous phenomenon:
Clinical mode collapse: diagnostic convergence toward statistically dominant patterns when multiple institutions rely on the same models trained on similar datasets, progressively reducing interpretive diversity and diminishing detection of rare or atypical phenotypes.
The central question of this investigation is not whether AI is good or bad.
It is whether we are preserving diagnostic sovereignty while integrating it.
Because when the system fails, the question will be simple:
Will there still be someone able to think without a safety net?
Analysis of Cognitive Architecture in the Era of Large Language Models
If the emergency room scenario feels uncomfortable, it is because it does not merely describe a clinical error.
It describes cognition.
To understand what may be unfolding in medicine, we must first examine what has already been observed in education, where large language models have become a live laboratory for cognitive transformation
In 2025, researchers at the MIT Media Lab conducted an experiment known as “Your Brain on ChatGPT.” Using electroencephalography (EEG), they monitored the brain activity of participants writing essays under three conditions: no assistance, search engine assistance, and full AI assistance.
The findings were striking.
Participants using AI assistance showed up to a 55% reduction in neural connectivity in frequency bands associated with deep thinking and internal monitoring. At the same time, writing speed increased by approximately 60%.
Performance improved.
Cognitive engagement decreased.
Even more concerning was memory retention. A significant proportion of participants who relied heavily on AI struggled to recall passages they had written only minutes earlier.
This pattern reflects a well-established psychological mechanism known as cognitive offloading: the delegation of reasoning, synthesis, and memory retrieval to an external system.
Cognitive offloading is not inherently harmful. Humans have always externalized cognition, from writing to calculators.
But here the shift is qualitative.
When a system does not merely store information but generates structured reasoning on our behalf, the brain may begin to reorganize its allocation of effort. The discomfort required for deep encoding, hypothesis generation, and conceptual integration is reduced.
And in learning, discomfort is not noise. It is signal.
This reveals a tension that is often overlooked: short-term performance optimization versus long-term cognitive resilience.
The question is not whether AI makes us faster, the question is whether it changes how much uncertainty we are willing to tolerate before delegating the effort.
In clinical medicine, that tolerance is not optional, it is the substrate of expertise.
Because expertise is not the ability to recognize common patterns when everything aligns, it is the ability to remain cognitively active when nothing aligns.
If large language models subtly reduce metacognitive vigilance, diminish internal simulation of alternatives, or weaken memory consolidation, the downstream impact on medical training could be profound.
Residents may produce correct answers more frequently.
But they may construct fewer diagnostic pathways internally.
And when a pathway has not been built, it cannot be reconstructed under pressure.
This is where systemic epistemic fragility begins.
Not in visible collapse.
But in incremental delegation.
The Emergence of Metacognitive Laziness and Professional Deskilling
Cognitive change does not begin with error.
It begins with comfort.
As AI systems become more fluent, more precise, and more seamlessly embedded into workflows, something subtle shifts in the user. The effort required to question an answer decreases. The impulse to scrutinize it weakens.
This has been described as metacognitive laziness: a progressive reduction in the willingness to critically evaluate, actively challenge, and internally simulate alternatives when an external system offers a coherent solution
The danger is not ignorance, It is passivity.
When responses are instantaneous, structured, and delivered with linguistic confidence, the cognitive friction that normally drives deeper inquiry diminishes. The discomfort that forces us to ask, “What else could this be?” softens.
And without friction, critical thinking atrophies, in professional environments, this dynamic evolves into what has historically been called deskilling.
Deskilling is not new. Across industries, repeated delegation of complex tasks to automated systems has led to erosion of independent competence. The principle is both organizational and neurobiological: functions that are not exercised lose structural robustness.
In medicine, this takes on a different dimension.
Expertise is not merely accumulated knowledge. It is pattern recognition built through repeated exposure, error correction, and active reasoning under uncertainty. It develops through cognitive effort.
If early formative tasks, image interpretation, differential generation, structured clinical reasoning, are consistently mediated by AI, the developmental trajectory of expertise may shift.
A resident who repeatedly validates AI-generated drafts instead of constructing independent differential diagnoses may become more efficient.
But efficiency is not depth, this dynamic is further amplified by automation bias: the systematic tendency to favor suggestions from automated systems even when contradictory evidence exists.
Automation bias becomes particularly dangerous in high-complexity tasks. It gives rise to what could be described as the “verification paradox”: the more complex the case, the more likely the professional is to rely on AI precisely because they feel unable to independently verify the output.
Errors then take two forms:
Errors of commission: acting on an incorrect AI recommendation.
Errors of omission: failing to act because the system did not generate an alert.
Neither type of error arises purely from incompetence, they arise from over-delegation.
The central issue is not whether AI can outperform humans in specific tasks. In many domains, it already does.
The issue is whether repeated reliance is silently redefining the cognitive baseline of future professionals.
When reasoning shifts from generative to supervisory, something fundamental changes. The clinician no longer builds the argument, they audit it.
And auditing is cognitively lighter than constructing.
Over time, this may redefine what it means to “know”, n ot knowing how to derive from first principles, but knowing how to validate what another system proposes.
This is how systemic epistemic fragility deepens, not through spectacular technological failure, but through the gradual erosion of cognitive initiative.
Impact on Medical Education and Clinical Workflows
Medicine is not merely a body of knowledge.
It is a training architecture.
For decades, medical education has followed a progressive logic: residents learn by solving low- and medium-complexity cases under supervision. They make controlled mistakes. They formulate incomplete hypotheses. They err. They correct. They repeat.
Through that process, what are known as illness scripts are constructed, mental structures that integrate symptoms, pathophysiology, probabilities, and contextual experience.
The uncomfortable question is this:
What happens when AI becomes the first reader of the image, the first generator of differential diagnoses, the first synthesizer of the case?
In specialties such as radiology, pathology, and dermatology, AI systems already match or exceed human accuracy in specific detection tasks
But expert formation does not depend solely on final accuracy.
It depends on the intermediate cognitive process.
If the resident no longer generates hypotheses from scratch because the system already provides a ranked list, their exposure to uncertainty decreases. And with it, the active construction of independent diagnostic pathways diminishes.
The issue is not that AI assists, the issue is whether it replaces the generative phase of learning.
A qualitative study in Europe captured a recurring concern among senior physicians: they feel protected by having trained before the widespread integration of AI. Yet they worry that newer generations may not develop the same diagnostic “muscle” if formative tasks become fully automated
Here, a structural risk emerges.
The transition from junior to senior physician may be altered if the tasks through which early expertise is acquired, systematic review, active doubt, exhaustive exploration, are externalized.
And when that transition is disrupted, the system loses more than efficiency, it loses transmission.
Medicine has historically functioned as an intergenerational apprenticeship model. Tacit knowledge, calibrated intuition, the subtle “something doesn’t fit,” the quiet suspicion, is not taught through manuals. It is transmitted in practice.
If residents validate machine-generated drafts rather than construct reasoning in front of their mentors, a gap of expertise begins to form.
That gap is invisible when everything works.
It becomes visible when the system fails.
This is where the notion of diagnostic sovereignty becomes central.
Diagnostic sovereignty does not mean rejecting technology.
It means preserving the autonomous capacity of the medical system to sustain and evolve its knowledge without structural dependency on external computational support.
If clinical reasoning shifts from argument construction to algorithm supervision, the profession may undergo a silent transition:
From physicians who think with AI
to physicians who merely validate what AI thinks.
And that difference is deeper than it appears.
Because when the algorithm fails to recognize the pattern, someone must.
The question is whether that someone will still exist.
Evidence of Automation Bias in Medical Environments
So far, we have discussed cognitive architecture and transformation in training.
But this phenomenon is not merely theoretical.
There is already empirical evidence showing how algorithmic recommendations modify clinical judgment.
Automation bias is not an abstract hypothesis. It is a documented behavioral pattern: the tendency to favor suggestions from automated systems even when contradictory information is available
In 2025, a randomized clinical trial evaluated the diagnostic performance of physicians exposed to recommendations generated by a large language model.
The findings were revealing.
When the system provided accurate recommendations, diagnostic accuracy reached approximately 84.9%.
But when the system introduced incorrect suggestions, accuracy dropped to around 73.3%.
A difference of more than ten percentage points, what was most concerning was not the model’s error, It was human deference to the model’s error
Even physicians trained in AI usage showed vulnerability when confronted with incorrect algorithmic outputs. The mere presence of an AI-generated recommendation reshaped subsequent reasoning.
In other words, AI does not simply add information.
It alters the cognitive frame within which information is interpreted.
Additional studies in radiology have shown that when a system suggests an incorrect reading, some residents reduce their independent search for confirmatory signs. In gastroenterology, prolonged use of algorithmic assistance has been associated with reduced manual vigilance. In critical care, clinicians have been observed dismissing divergent clinical data in favor of automated sepsis alerts
These findings do not imply that AI is inherently dangerous, they imply that human–AI interaction modifies human behavior.
Here emerges what we might call the complexity paradox.
The more complex the case, the more tempting it becomes to defer to the system.
But the more complex the case, the higher the potential cost of an incorrect recommendation.
Systemic epistemic fragility does not arise solely from technological failure, It arises when professionals progressively internalize the algorithm as the starting point of reasoning rather than its complement.
If human reasoning begins to align itself around algorithmic outputs instead of actively interrogating them, medicine may enter a phase in which interpretive diversity diminishes.
And this leads us to the next layer of the problem, not merely individual dependency, but collective convergence.
The Challenge of Rare Diseases and Out-of-Distribution Presentations
Artificial intelligence has demonstrated extraordinary potential in analyzing large volumes of biomedical data.
In rare diseases, for instance, deep learning systems have identified unexpected therapeutic associations after screening thousands of molecules and genetic profiles. Models trained on massive datasets have, in certain contexts, surpassed human capacity to integrate complex genomic information.
This is the luminous face of clinical AI, but there is a less visible reverse side.
Machine learning models perform optimally within the statistical space on which they were trained. They are highly efficient when the presented case belongs to the known data distribution.
The problem emerges when the case does not fit.
In real-world medicine, patients are rarely statistical averages.
They are atypical combinations.
Incomplete phenotypes.
Hybrid presentations.
Noisy data.
These scenarios are technically referred to as out-of-distribution (OOD) cases.
Evidence suggests that large language models and other AI systems struggle significantly with recognizing misleading cues, updating hypotheses in the face of ambiguous new information, or reinterpreting data that does not obviously alter the initial management plan
In simple terms:
They are excellent at recalling learned patterns.
They are less robust when navigating dynamic ambiguity.
This is where the concept of clinical mode collapse becomes particularly relevant, in machine learning, mode collapse occurs when a model converges toward average solutions and loses diversity in its outputs.
In medicine, an analogous phenomenon may emerge if multiple institutions adopt:
The same models
Trained on similar datasets
Optimized under similar performance objectives
The result could be diagnostic convergence toward dominant statistical patterns, a progressive reduction in interpretive diversity and consequently, diminished sensitivity to rare or atypical presentations.
This is not merely a technical risk, it is a systemic one.
The historical strength of medicine has been its ability to recognize the unexpected.
If clinical practice increasingly aligns around statistically frequent patterns, rare diseases may become even more invisible.
And when the algorithm fails to recognize a pattern because it has not encountered it sufficiently in training, the only remaining safeguard is human judgment trained in suspicion.
But such judgment must have been exercised, here, systemic epistemic fragility reaches its most critical expression:
Not when the model fails in a common case, but when the system as a whole loses the capacity to detect the improbable.
Medicine does not collapse when it fails at the obvious, it collapses when it ceases to recognize the rare.
Systemic Epistemic Fragility and the Erosion of Tacit Knowledge
So far, we have described individual components of the problem: cognitive offloading, metacognitive laziness, automation bias, diagnostic convergence.
Together, however, they configure something broader, not an isolated error, but a structural transformation of the medical knowledge ecosystem.
Here I formalize the central concept of this analysis:
Systemic epistemic fragility is the progressive reduction of a professional system’s collective capacity to:
Detect algorithmic errors when they occur.
Generate diagnostic hypotheses outside the space suggested by AI.
Tolerate uncertainty without computational support.
Transmit tacit knowledge intergenerationally through deliberate practice.
Fragility does not imply immediate incompetence, It implies loss of cognitive resilience.
An epistemically fragile system may function efficiently under normal conditions, but it becomes vulnerable in unforeseen scenarios, technological failures, or abrupt changes in the clinical environment.
The most delicate component of this fragility is the erosion of tacit knowledge.
Explicit knowledge can be documented, digitized, transferred, tacit knowledge cannot.
It is calibrated intuition built through thousands of micro-decisions.
It is the “something doesn’t fit” before sufficient data are available.
It is suspicion emerging before the algorithm assigns a probability.
Historically, this knowledge has been transmitted through supervised practice, case discussion, and repeated exposure to ambiguity, if AI disrupts this intergenerational chain by automating the tasks through which early intuition is acquired a silent gap begins to form.
That gap is not detected by productivity metrics, It does not appear in performance dashboards, It manifests when the system needs to improvise.
A complementary phenomenon may be described as confidence laundering. Large language models present outputs with a uniform tone of authority, even when underlying uncertainty is high, when statistical probability is delivered with narrative fluency, the human brain tends to interpret it as implicit certainty.
If clinicians do not understand the underlying generative architecture, an epistemic dissociation occurs:
The source of knowledge is no longer interrogated, It is accepted.
And when acceptance replaces inquiry, the scientific framework weakens, modern medicine was built on systematic doubt, if technological integration reduces the practice of doubt, we are not merely facing a digital revolution.
We are facing an epistemological transformation.
The question is not whether AI will make mistakes, the question is whether the system will retain the structural capacity to question it.
Diagnostic resilience does not depend on technological perfection, It depends on the active coexistence between human suspicion and algorithmic calculation.
And that coexistence does not emerge by default, It must be designed.
Counteranalysis: AI as an Amplifier of Cognitive Capacity
So far, the analysis has been deliberately critical, but it would be intellectually dishonest to ignore the other side of the phenomenon.
Artificial intelligence does not only introduce risks, It can also amplify human capability when integrated deliberately.
Several studies suggest that certain modes of human–AI interaction, particularly those that make the model’s reasoning process explicit, can enhance clinical reflection rather than weaken it
For example, systems that expose their chain of reasoning allow clinicians not only to receive a conclusion, but to examine the inferential pathway behind it.
In that context, AI can function as:
A generator of alternative hypotheses
A detector of human cognitive bias
A real-time probabilistic simulator
When the model does not replace reasoning but instead challenges it, a different dynamic emerges.
Instead of cognitive offloading, cognitive expansion may occur.
Instead of metacognitive laziness, assisted metacognition may develop.
The key is not the presence of AI, It is the design of the interaction.
If clinicians treat AI output as a “first draft” that must be systematically questioned, compared, and reconstructed, the process may strengthen reasoning rather than erode it.
Similarly, automating administrative or large-scale data-processing tasks can free cognitive bandwidth for higher-level activities: patient communication, contextual integration, ethical deliberation.
Even Bayesian-based systems can serve as probabilistic correctives to human confirmation bias, recalibrating hypotheses in real time.
In this scenario, AI does not reduce diagnostic sovereignty, It reinforces it.
But this reinforcement does not emerge automatically, It requires institutional design, explicit training, and clear operational boundaries. Technology alone does not determine the outcome, the system that adopts it does.
Prospective Scenario 2035–2045 and Governance Frameworks
Let us imagine the year 2040. Artificial intelligence is fully integrated into hospital systems. Algorithms perform initial triage, generate prioritized differentials, suggest therapeutic plans, and monitor outcomes in real time. Medicine is faster, more efficient, and more standardized.
Yet from this integration, two distinct futures may emerge.
In the first scenario, the clinician progressively becomes an algorithmic supervisor whose primary function is to validate recommendations generated by highly optimized systems. Average precision improves and variability decreases, but interpretive diversity diminishes, along with the capacity for diagnostic improvisation. The system functions smoothly as long as the environment remains stable. However, when confronted with emerging diseases, out-of-distribution presentations, or systemic technological failures, structural dependency becomes visible. The problem is not a single error. It is the erosion of collective resilience.
In the second scenario, AI is integrated as a tool of expansion rather than substitution. Systems are deliberately designed to display explicit uncertainty levels, expose their inferential reasoning, compel the generation of alternative hypotheses, and simulate failure scenarios. The clinician does not lose autonomy but exercises it under new conditions. Human–AI interaction becomes deliberative rather than passive delegation.
The difference between these futures is not technological. It is institutional. If we accept that AI integration reshapes cognitive architecture and medical training structures, its adoption cannot be guided solely by efficiency metrics. It must incorporate deliberate design for epistemic resilience.
Within this context, three governance pillars become central:
Explicitly defining boundaries of algorithmic autonomy according to clinical risk.
Implementing “guard band” strategies that include failure simulations and mandatory manual-mode training.
Ensuring structural transparency and auditability in deployed systems. Without transparency, there is no critical deliberation. Without critical deliberation, there is no diagnostic sovereignty.
The large-scale deployment of AI in medicine is not merely a technological challenge. It represents a redesign of the profession’s epistemic contract. And if that redesign is not deliberate, it will be accidental. And what is accidental rarely favors resilience.
Introduction of the Model
If we accept that AI integration is not merely a technological enhancement but a transformation of the cognitive architecture of medical practice, then the response cannot be limited to warnings or superficial adjustments.
It requires structural redesign.
It requires a deliberate architecture that preserves intellectual autonomy while integrating algorithmic power.
From this analysis, I propose what I call:
The Diagnostic Sovereignty Framework: A Path to Cognitive Resilience in the Age of AI
This framework does not seek to oppose artificial intelligence, but to structure its integration in a way that strengthens, rather than weakens, human diagnostic capacity.
Its objective is to preserve diagnostic sovereignty through the active development of cognitive resilience across three interconnected levels:
Medical education
Human–AI interaction
Institutional governance
This is not about slowing innovation.
It is about designing it with epistemic awareness.
Conclusion
Artificial intelligence is not the problem.
The problem is how we choose to integrate it.
Evidence suggests that algorithmic systems can improve performance, speed, and consistency. Yet it also indicates that they reshape cognitive patterns, influence metacognitive disposition, and may reduce interpretive diversity if not carefully designed.
The question is not whether AI will make mistakes.
The question is whether medicine will retain the structural capacity to recognize them.
The history of medicine has been a history of disciplined doubt, of calibrated suspicion and first-principles reasoning in the face of uncertainty.
If the AI era transforms that habit into passive validation, we will lose more than professional autonomy.
We will lose resilience.
The future of medicine does not depend on rejecting artificial intelligence.
It depends on designing its integration in a way that preserves diagnostic sovereignty and strengthens cognitive resilience.
That is the purpose of the Diagnostic Sovereignty Framework.
I am actively developing this framework as an applicable architecture for medical education, clinical practice, and hospital governance in AI-assisted environments.
If this conversation resonates with you, let us continue building it.
References:
Mineo L. Is AI dulling our minds? Harvard Gazette [Internet]. 2025 Nov 13 [cited 2026 Feb 20]. Available from: https://news.harvard.edu/gazette/story/2025/11/is-ai-dulling-our-minds/
EDUCAUSE Review [Internet]. [cited 2026 Feb 20]. The Paradox of AI Assistance: Better Results, Worse Thinking. Available from: https://er.educause.edu/articles/2025/12/the-paradox-of-ai-assistance-better-results-worse-thinking
Roxin I. Generative AI: the risk of cognitive atrophy. Polytechnique Insights [Internet]. 2025 Jul 3 [cited 2026 Feb 20]. Available from: https://www.polytechnique-insights.com/en/columns/neuroscience/generative-ai-the-risk-of-cognitive-atrophy/
Shukla P, Bui P, Levy SS, Kowalski M, Baigelenov A, Parsons P. De-skilling, Cognitive Offloading, and Misplaced Responsibilities: Potential Ironies of AI-Assisted Design. In: Proceedings of the Extended Abstracts of the CHI Conference on Human Factors in Computing Systems [Internet]. 2025 [cited 2026 Feb 20]. p. 1–7. Available from: http://arxiv.org/abs/2503.03924 doi:10.1145/3706599.3719931
Risko E, Gilbert S. Cognitive Offloading. Trends in cognitive sciences. 2016 Aug 1;20. doi:10.1016/j.tics.2016.07.002
Sunday O. Behavioral Impacts of AI Reliance in Diagnostics: Balancing Automation with Skill Retention. Epidemiology and Health Data Insights. 2025 Sep 8;1(3):ehdi011. doi:10.63946/ehdi/16894
Natali C, Marconi L, Duran L, Cabitza F. AI-induced Deskilling in Medicine: A Mixed-Method Review and Research Agenda for Healthcare and Beyond. Artificial Intelligence Review. 2025 Aug 27;58. doi:10.1007/s10462-025-11352-1
Qazi IA, Ali A, Khawaja AU, Akhtar MJ, Sheikh AZ, Alizai MH. Automation Bias in Large Language Model Assisted Diagnostic Reasoning Among AI-Trained Physicians [Internet]. medRxiv; 2025 [cited 2026 Feb 20]. p. 2025.08.23.25334280. Available from: https://www.medrxiv.org/content/10.1101/2025.08.23.25334280v2doi:10.1101/2025.08.23.25334280
Wang X, He D, Jin C. AI-driven enhancements in rare disease diagnosis and support system optimization. Intractable Rare Dis Res. 2025 Nov 30;14(4):306–8. doi:10.5582/irdr.2025.01079 PubMed PMID: 41341910; PubMed Central PMCID: PMC12672146.
Rutherford G. Doctors still outperform AI in clinical reasoning, study shows [Internet]. [cited 2026 Feb 20]. Available from: https://www.ualberta.ca/en/folio/2025/11/doctors-still-outperform-ai-in-clinical-reasoning.html
Sunday O. Behavioral Impacts of AI Reliance in Diagnostics: Balancing Automation with Skill Retention. Epidemiology and Health Data Insights. 2025 Sep 8;1(3):ehdi011. doi:10.63946/ehdi/16894
Epistemic Justice as a Condition for Meaningful Human Control Over Medical AI | springerprofessional.de [Internet]. [cited 2026 Feb 20]. Available from: https://link.springer.com/article/10.1007/s11023-026-09762-3
Breaking the Chain of Knowledge Transfer: AI Shadows Implicit, Explicit and Tacit Exchange. IIS. 2025. doi:10.48009/3_iis_2025_2025_108
Ide E. Automation, AI, and the Intergenerational Transmission of Knowledge [Internet]. arXiv; 2025 [cited 2026 Feb 20]. Available from: http://arxiv.org/abs/2507.16078 doi:10.48550/arXiv.2507.16078
Kelly M. The Epistemic Suite: A Post-Foundational Diagnostic Methodology for Assessing AI Knowledge Claims.
Rahdari B, Shaikh S, Chen JH, Gerstenberg T, Raj S. From Retrieving Information to Reasoning with AI: Exploring Different Interaction Modalities to Support Human-AI Coordination in Clinical Decision-Making [Internet]. arXiv; 2026 [cited 2026 Feb 20]. Available from: http://arxiv.org/abs/2601.22338doi:10.48550/arXiv.2601.22338
Greengrass CJ. Transforming clinical reasoning—the role of AI in supporting human cognitive limitations. Front Digit Health. 7:1715440. doi:10.3389/fdgth.2025.1715440 PubMed PMID: 41561162; PubMed Central PMCID: PMC12813117.
iatroX [Internet]. 2025 [cited 2026 Feb 20]. The deskilling dilemma: will clinical AI erode or enhance medical expertise? (UK, 2025) | iatroX Clinical AI Insights. Available from: https://www.iatrox.com/blog/clinical-ai-deskilling-evidence-and-strategies-for-uk-doctors-2025
Enthusiast EH. Your Brain on ChatGPT: Accumulation of Cognitive Debt when Using an AI Assistant for Essay Writing…. Medium [Internet]. 2025 Jun 23 [cited 2026 Feb 20]. Available from: https://medium.com/@EleventhHourEnthusiast/your-brain-on-chatgpt-accumulation-of-cognitive-debt-when-using-an-ai-assistant-for-essay-writing-253fb48d0863
Harvard Warns of AI’s Cognitive Impact. AI CERTs News [Internet]. [cited 2026 Feb 20]. Available from: https://www.aicerts.ai/news/harvard-warns-of-ais-cognitive-impact/
Warm E. ICE Blog [Internet]. 2025 [cited 2026 Feb 20]. Deskilling and Automation Bias: A Cautionary Tale for Health Professions Educators. Available from: https://icenet.blog/2025/08/26/deskilling-and-automation-bias-a-cautionary-tale-for-health-professions-educators/
Shukla P, Bui P, Levy SS, Kowalski M, Baigelenov A, Parsons P. De-skilling, Cognitive Offloading, and Misplaced Responsibilities: Potential Ironies of AI-Assisted Design. Proceedings of the Extended Abstracts of the CHI Conference on Human Factors in Computing Systems. 2025 Apr 26;1–7. doi:10.1145/3706599.3719931
Nilsson C. The artificial intelligence (AI) competence paradox: how AI reshapes clinical expertise. Transforming Government: People, Process and Policy. 2025 Aug 12. doi:10.1108/TG-02-2025-0048
From Code to Conscience: An Ethical Framework for Healthcare AI | Edmond & Lily Safra Center for Ethics [Internet]. 2025 [cited 2026 Feb 20]. Available from: https://www.ethics.harvard.edu/news/2025/11/code-conscience-ethical-framework-healthcare-ai-0




