A retrospective observational study published in Lancet Gastroenterology & Hepatology in 2025 that reports a decline in the adenoma detection rate (ADR) of standard, non-AI colonoscopy among endoscopists after they began regular use of AI polyp-detection tools. The analysis is nested in the ACCEPT randomised trial at four Polish endoscopy centres and is described by the authors as evidence that continuous AI exposure might have a negative effect on endoscopist behaviour. It is the first peer-reviewed paper to document a patient-relevant clinical endpoint declining after clinicians began working alongside AI.
Citation: Budzyń K, Romańczyk M, Kitala D, et al. Endoscopist deskilling risk after exposure to artificial intelligence in colonoscopy: a multicentre, observational study. Lancet Gastroenterol Hepatol. 2025;10(10):896–903. DOI: 10.1016/S2468-1253(25)00133-5. PMID: 40816301.
Confidence note: This summary works from the PubMed abstract only; the Lancet full text is paywalled. Claims about per-centre variation, proposed mechanisms, and the authors' own limitations framing cannot be verified from the abstract and should not be added without retrieving the full text.
Study design
The study is a retrospective, observational cohort analysis nested in the ACCEPT randomised trial (Artificial Intelligence in Colonoscopy for Cancer Prevention), conducted at four endoscopy centres in Poland (Source: https://www.thelancet.com/journals/langas/article/PIIS2468-1253(25)00133-5/fulltext). ACCEPT introduced AI polyp-detection tools at the end of 2021; after rollout, individual procedures were randomised (AI on vs. AI off) by date.
The analysis covers colonoscopies performed between September 8, 2021 and March 9, 2022, comparing three months before the AI rollout against three months after it. It does not measure AI-assisted performance; instead it isolates only the non-AI colonoscopies from both phases (n=1443 total; 795 pre-AI, 648 post-AI) to measure the counterfactual skill level of endoscopists working without the tool. Patients with intensive anticoagulant use, pregnancy, a history of colorectal resection, or inflammatory bowel disease were excluded.
The primary outcome was the change in ADR of standard, non-AI colonoscopy before versus after AI exposure. ADR was modelled with multivariable logistic regression adjusting for patient sex and age. ADR is a validated quality indicator linked to post-colonoscopy colorectal cancer incidence and mortality.
Findings
ADR of standard (non-AI) colonoscopy fell from 28.4% (226/795) pre-AI to 22.4% (145/648) post-AI, an absolute difference of −6.0 percentage points (95% CI: −10.5 to −1.6; p = 0.0089). In the multivariable logistic regression, AI exposure was associated with an odds ratio of 0.69 (95% CI 0.53–0.89) for detecting an adenoma when AI was not in use.
Other independent predictors of ADR in the adjusted model were male versus female patient sex (OR 1.78, 1.38–2.30) and patient age ≥60 versus <60 (OR 3.60, 2.74–4.72).
The authors' abstract interpretation is cautious: "Continuous exposure to AI might reduce the ADR of standard non-AI assisted colonoscopy, suggesting a negative effect on endoscopist behaviour." They do not claim causation from a retrospective design. The abstract notes the finding is compatible with several mechanisms — automation bias, reduced vigilance when the safety net is absent, attenuation of deliberate search behaviour, or a shift in attentional allocation — but does not adjudicate between them, and the full-text discussion is not available here.
Where AI-adoption studies have mostly measured combined human-plus-AI performance and found it superior, this study isolates the human-only counterfactual and reports that it degraded.
Limitations
The limitations below are drawn from the abstract; the full text was not consulted.
- Observational and retrospective. The pre/post comparison is not randomised at the endoscopist level. Secular trends, seasonal case-mix shifts, or other 2021–2022 changes at these centres could contribute.
- Short window. Three months before versus three months after; the long-run trajectory is unclear.
- Patient-level, not endoscopist-level, primary unit. The abstract does not report how many individual endoscopists contributed to each phase or the distribution of effect across them.
- Single health system and trial context. All four centres were ACCEPT-enrolled Polish sites. Generalisability to other settings, AI tools, and specialties is untested.
- Unmeasured confounders. Changes in referral patterns, bowel-prep protocols, scope hardware, or withdrawal-time norms during the rollout period are not addressed in the abstract.
- Paywall caveat. Mechanism discussion and further sensitivity analyses may be present in the full text and are not reflected here.
Context
The study is positioned as the first peer-reviewed, patient-endpoint evidence for AI-induced deskilling. Earlier deskilling discussions relied on ergonomics, psychology, and analogy to aviation autopilot and radiology reading lists; this analysis moves the question from theoretical to empirical in at least one clinical domain.
The finding sits alongside, rather than against, the task-level augmentation literature. Studies such as Brynjolfsson, Li, Raymond (2023) find that AI lifts baseline-low performers while combined output rises. The Lancet result is consistent with both — combined human-plus-AI performance likely improved during ACCEPT — while indicating that unassisted human skill can degrade over the same period. The policy question raised is what happens when AI is unavailable: downtime, contraindicated patients, new settings without the tool, or regulatory removal. If continuous AI use erodes unassisted ADR within months, training pathways (residency, certifying exams, continuing medical education) may need to preserve regular unaided practice, with parallels drawn to pilot "hand-flying" minimums and surgical simulation.
The result connects to the workforce literature on generative AI and cognitive offloading: Generative AI and the Nature of Work on task-composition shifts and Applying AI to Rebuild Middle-Class Jobs on whether AI rebuilds or hollows out expertise. It also provides an outcome-level reference point for survey sentiment such as the AMA 2026 survey, in which 88% of physicians reported concern about AI-induced deskilling; that survey captures sentiment, while this paper offers outcome-level corroboration in one task (AI Deskilling).
Relationships
- supports: AI Deskilling — first peer-reviewed outcome-based evidence for the concept
- related: AI and Productivity — complicates, rather than contradicts, the augmentation-then-displacement frame by isolating the unassisted-human counterfactual
- related: Applying AI to Rebuild Middle-Class Jobs — Autor's question of whether AI rebuilds or erodes expertise finds a negative data point here in a high-skill medical task
- related: Generative AI and the Nature of Work — task-composition and cognitive-offloading mechanisms are plausible candidates for the underlying effect
- related: AI Labor Disruption — deskilling is a distinct workforce risk from displacement and augmentation