The authors define “near future” as 5–10 years (page 2), yet the conclusion (page 28) states AI “will not replace physicians in the foreseeable future” and mentions “future directions” requiring longitudinal studies, standardized frameworks, and regulatory alignment. Is there a logical inconsistency between a 5–10 year technological cycle and the conclusion that replacement remains structurally impossible? If not, what specific milestones would need to be achieved within this window to challenge the conclusion?
The manuscript cites “Table 2” on page 4 as summarizing task-level vs. physician-level replacement, but then Table 2 on page 19 presents performance metrics. Additionally, Table 4 is introduced on page 21 as summarizing representative studies, yet Table 2 on page 19 already serves this purpose. Why are these tables not clearly distinguished, and is there redundancy that undermines the paper’s organizational clarity?
Section 7.1 states that “even a system with 90% accuracy necessarily implies systematic harm to a nontrivial number of individuals” and uses this as an argument against autonomous AI. However, the authors’ own Table 4 shows many systems with AUC > 0.90, and clinical benchmarks for human physicians in certain tasks (e.g., mammography interpretation) show comparable or lower sensitivity. Does this argument unintentionally imply that current human practice is also unacceptably error-prone? If not, why is statistical error framed differently for AI than for humans, especially given that AI performance in some domains exceeds human benchmarks?