A recent study published by researchers at Dartmouth College has shown that an AI tutoring platform called Phosphor can have a significant impact on student performance [1]. The platform uses large language models to grade formative assessments and provide feedback to students. In a pilot study, students who fully engaged with the platform showed a 0.71 to 1.30 standard deviation improvement in final exam performance compared to baseline expectations [2]. The study found that the constructed-response format with AI grading was the driver of this improvement, with no significant effect from multiple-choice-only quizzes [3].

The researchers note that this was an observational study, not a randomized controlled trial, and that there are limitations to the findings [4]. However, the results suggest that AI-powered personalized instruction can be highly effective in enhancing student learning outcomes [5]. The study's findings are consistent with the emerging vision of intelligent textbooks, which integrate AI-powered assessment and feedback into instructional content [6]. The use of large language models for grading and feedback has shown promising alignment with human raters [7], and the study's results suggest that this technology can be successfully deployed in an institutional setting [8].

The researchers plan to conduct further studies to strengthen the causal claim about the effectiveness of the AI tutor and to replicate the findings in additional university gateway courses [9].