The first Delphi round had a low completion rate of only 39% (28 out of 72 invitees). Given the exceptionally long average survey completion time of 71 minutes (with responses ranging up to 118 minutes), is it not highly likely that the resulting consensus is biased towards a self-selected group of highly motivated specialists? How can the steering group be confident that this significant attrition did not systematically exclude important perspectives (e.g., from primary care or those with less time), thereby undermining the generalizability of the framework?
The first-round survey was excessively burdensome, featuring up to 176 closed questions due to the integrated mapping exercise. The paper even concedes this may have caused in-survey dropout.Can the authors justify their decision to combine the importance rating and the level-of-delivery mapping in a single, lengthy first round, rather than conducting the mapping as a more focused, second-phase survey after the core capability statements had been agreed upon?