
Or: What happens on the day the tool fails
Within a single month, three warnings about artificial intelligence (AI) and medical training reached three different audiences. A Nature Medicine Perspective named never-skilling: the risk that trainees who rely on AI during formative years fail to develop independent clinical reasoning.1 The clinical trade press described educators struggling to decide which skills to protect. Surgeon-educators called for AI literacy to become mandatory core curriculum rather than fortunate exposure. The three converge on a real problem. They also share an omission that determines whether any proposed fix will hold.
The omission is an engineering concept that operative medicine already lives by: redundancy. In the operating room, equipment failure is not an unexpected event. It is a rehearsed contingency. Anesthesia depends on the capnograph and the pulse oximeter, and the competent anesthesiologist still knows how to monitor by hand when those devices fail. The never-skilling Perspective offers the precordial stethoscope as an example of a skill safely retired by technological progress.1 The example proves the opposite of what it is asked to prove. Manual monitoring was safe to retire as a routine precisely because the underlying competency was retained as a fallback. Capnography replaced the need to auscultate continuously on every case. It did not replace the ability to do so when the monitor quits. The routine was delegated. The competency was kept, on call, as a redundant channel.
This distinction is the whole question, and it has a clean rule. A skill is safe to retire as routine only when the underlying competency is retained as a redundant channel against the foreseeable failure of the tool that replaced it. Retiring the routine is delegation. Retiring the competency is abdication. The two look identical on a normal day and diverge completely on the day the tool fails.
Applied to AI, the rule exposes an asymmetry the current debate has not named. For physiological monitoring, redundancy is mechanical: when one device fails, another covers it. For clinical reasoning, there is no redundant device. The physician is the only redundant channel. There is no second trusted reasoner at the bedside, no backup model whose output a clinician would accept unexamined when the first is wrong or unavailable. AI-assisted clinical reasoning is therefore not analogous to the pulse oximeter, which a second instrument can back up. It is analogous to manual flight: the human is the redundancy, which is exactly why the human must remain skilled. Aviation already encodes this. Pilots maintain manual proficiency alongside automation because automation failure is a planned contingency, not a remote one.2
The consequence reverses the dominant framing. AI is widely assumed to lower what a trainee must master by offloading cognitive work. It does the opposite. The former standard was singular: perform the skill. The new standard is dual. The clinician must still perform the skill, because the clinician is the fallback when the tool fails, and must additionally understand what the AI is doing and how and why, because that understanding is the only means of detecting that the tool is wrong. Offloading executes the task. It does not retire the competency. The promised efficiency dividend, less for the trainee to learn, is inverted: competent AI use requires a more capable clinician, not a cheaper one.
This reframes never-skilling in the language of systems safety. Deskilling degrades a redundant channel a clinician once possessed.3 Never-skilling means the redundant channel was never built. In any other safety-critical domain, commissioning a single-channel system for a life-critical function is not a training shortcut to be optimized. It is a design defect. Evidence already shows that unrestricted AI assistance can raise assisted performance while degrading independent capability,4 and that current models express confidence without the metacognition to flag their own errors.5 A trainee without an independent reasoning architecture cannot verify such an output. The trainee can only accept it. The verify-and-trust standard proposed for clinical AI presupposes a competency to verify against.6 Never-skilling removes the thing verification requires.
The accountability stakes follow directly. When AI-influenced reasoning produces harm, the developer is shielded by a disclaimer while the physician holds the liability, the license, and the patient. The US Food and Drug Administration excludes educational software from the clinical decision support oversight it applies elsewhere,7 leaving the formation of physician competence ungoverned at the point where it is most consequential. A physician trained to operate AI without the competence to audit it is positioned to cosign decisions they cannot evaluate while bearing full responsibility for them. The asymmetry of consequence should determine the asymmetry of authority, and with it the competency standard.
None of this argues against AI in medical education. It argues for sequence. The Nature Medicine framework is correct to place an AI-independent foundational phase before AI-integrated practice.1 The organizing principle for that sequence is redundancy. Build the physician first. Add the tool second. Teach the task as delegable and the understanding as never delegable. Assess the AI-unavailable case as a rehearsed contingency, the way surgical training rehearses conversion to an open procedure, rather than as a remote edge case. Mandating AI literacy without this sequence does not close the competency gap educators have identified. It certifies a more credentialed version of it.
The standard is not whether a clinician can use AI. It is whether the clinician remains the redundancy when AI fails. Optimize the clinician. Do not, in the name of efficiency, build one who cannot.
References
- Ke Y, Jin L, Ong JCL, et al. AI-induced never-skilling in medical education. Nat Med. Published online May 22, 2026. doi:10.1038/s41591-026-04438-y
- Ruskin KJ, Corvin C, Rice SC, Winter SR. Autopilots in the operating room: safe use of automated medical technology. Anesthesiology. 2020;133(3):653-665. doi:10.1097/ALN.0000000000003385
- Budzyń K, Romańczyk M, Kitala D, et al. Endoscopist deskilling risk after exposure to artificial intelligence in colonoscopy: a multicentre, observational study. Lancet Gastroenterol Hepatol. 2025;10(10):896-903. doi:10.1016/S2468-1253(25)00133-5
- Bastani H, Bastani O, Sungu A, Ge H, Kabakcı Ö, Mariman R. Generative AI without guardrails can harm learning: evidence from high school mathematics. Proc Natl Acad Sci U S A. 2025;122(26):e2422633122. doi:10.1073/pnas.2422633122
- Griot M, Hemptinne C, Vanderdonckt J, Yuksel D. Large language models lack essential metacognition for reliable medical reasoning. Nat Commun. 2025;16:642. doi:10.1038/s41467-024-55628-6
- Abdulnour RE, Gin B, Boscardin CK. Educational strategies for clinical supervision of artificial intelligence use. N Engl J Med. 2025;393(8):786-797. doi:10.1056/NEJMra2503232
- US Food and Drug Administration. Clinical Decision Support Software: Guidance for Industry and FDA Staff. FDA-2017-D-6569. January 2026.


