1. The screening process led to a remarkably high attrition rate (5,731 abstracts to 25 included papers), raising a significant concern about selection bias. The authors’ exclusion of “Kirkpatrick level 1” studies (perceptions) and their focus on “objective or measurable learning outcomes” seems to have filtered out a vast amount of context-rich data that is crucial for a realist review. How do the authors justify this stringent exclusion criterion given the realist methodology’s focus on understanding mechanisms (e.g., motivation, engagement) which are often best illuminated by qualitative data on student and faculty perceptions?
2. The study’s findings are heavily influenced by a homogeneous sample, predominantly from U.S. medical programs. Given that the review itself identifies “culture and language” (Programme Theory 7) as a key mechanism impacting success, how do the authors account for the potential that their programme theories are culturally specific? The absence of studies from non-Western, non-English-speaking contexts (with only one each from China, South Korea, and Indonesia) significantly limits the transferability and generalizability of their conclusions.
3. The discussion of “institutional support” (Programme Theory 1) appears to overlook a fundamental contradiction in the evidence presented. The authors argue that institutional support leads to better outcomes, yet many of the studies they cite as evidence for this theory (e.g., Nordquist 2012) actually reported no significant improvement in knowledge or engagement. How do the authors reconcile this, and does this suggest that factors beyond simple recognition of faculty time (such as curriculum integration or pedagogical design) are more critical than their programme theory suggests?