H1: AI and human agreement (primary outcome)
The Kairos assessor will show statistically significant agreement with expert human judges on developmental stage ratings from natural language transcripts, measured by quadratic weighted kappa (κw). The pre-registered success criterion requires both the point estimate and the lower bound of the 95% confidence interval to exceed κw ≥ 0.40. This matches the substantial agreement threshold on Fleiss’s benchmark scale [24].
H2: Confidence and accuracy calibration
The AI assessor’s self-reported confidence scores will show a statistically significant negative Spearman rank correlation with the AI’s own absolute rating error relative to human consensus ground truth. This would confirm that higher confidence goes with lower error. The pre-registered success criterion is a one-tailed p < 0.10.
H3: Confidence convergence
The AI assessor’s self-reported confidence will show a statistically significant positive Spearman rank correlation with the human expert panel’s own confidence ratings. This would confirm that the AI’s uncertainty tracks a shared sense of how hard each case is. The pre-registered success criterion is a one-tailed p < 0.10.
For these two secondary calibration hypotheses, the pre-registered threshold of one-tailed p < 0.10 is more lenient than the conventional 0.05. This bar was set in advance because H2 and H3 are directional secondary checks in a pilot study, intended to guide the design of a larger replication rather than to establish the calibration effects on their own. The primary hypothesis (H1) uses a much stricter criterion that does not depend on this threshold.