Two college chatbot studies show AI helps with tasks, not persistence
Two July 2026 working papers point to a narrower use for higher-ed AI. Course-specific chatbot outreach improved grades and tutoring use, while a four-year administrative bot helped students complete time-sensitive tasks without detectable gains in academic performance or persistence.

Two new higher-ed chatbot studies released in July 2026 argue for a narrower, less magical view of AI student support. A new NBER working paper reports that course-specific chatbot outreach in large undergraduate classes improved final grades and students’ engagement with supports such as tutoring. A separate July 2026 EdWorkingPaper following an administrative text-messaging chatbot over four years found gains in completing time-sensitive tasks, but no detectable effects on academic performance or persistence. Read together, the papers suggest that chatbots can help colleges solve specific problems; they do not show that AI, by itself, fixes retention. (nber.org)
That distinction matters as colleges head into the fall under heavy pressure to do something with AI. The 2026 EDUCAUSE Horizon Report says AI is already reshaping academic support, assessment, and student-faculty relationships, while both new chatbot studies arrive as working papers rather than final journal articles, meaning leaders should treat them as meaningful evidence but not the last word. The NBER paper explicitly notes that it has not been peer reviewed. (library.educause.edu)
Stronger inside the course
The clearest positive academic signal comes from the course-level study, Let’s Chat: Leveraging Chatbot Outreach for Improved Course Performance. The paper reports preregistered randomized trials of a non-generative AI chatbot in large-enrollment undergraduate courses, with the experiments registered in the Registry of Efficacy and Effectiveness Studies. Public descriptions of the study identify the setting as Georgia State University and the courses as Introduction to American Government and Principles of Microeconomics. (nber.org)
The headline result is not that students suddenly had an always-on AI tutor. It is that proactive, course-specific messaging improved final grades and increased engagement with academic supports such as tutoring. The paper says effects were generally consistent across student groups, with one notable exception: women in one microeconomics course saw a larger gain, earning final grades seven percentage points higher than women in the control group. That is a promising instructional result, especially because it came from a targeted system tied to a particular course rather than from a campuswide general assistant. (nber.org)
Earlier public summaries of the same Georgia State line of work help explain why the intervention may have worked. They describe the chatbot as supporting academic task navigation in introductory courses and point to increased use of supplemental instruction opportunities, alongside stronger course performance. In other words, the bot appears to have been useful less as a substitute teacher than as a timed prompt that kept students connected to the work and to human help already available around the course. (eric.ed.gov)
Administrative bots helped with deadlines, not persistence
The July 2026 administrative study lands in a different place. In Sustaining AI-Enabled Student Support: A Four-Year Implementation and Impact Study, Catherine Mata, Emily Russell, and Lindsay Page examine an AI-enabled text-messaging chatbot at a large urban public university over a four-year randomized controlled trial, alongside implementation evidence from system observation and administrator discussions. Their conclusion is specific: the chatbot improved completion of time-sensitive administrative tasks, students stayed receptive over time, and the strongest implementation conditions were centralized ownership and flexible communication. But the study found no detectable effects on academic performance or persistence. (edworkingpapers.com)
That is not a trivial result. Colleges often lose students through mundane but consequential friction: holds, forms, registration barriers, and other deadlines that are easy to miss and hard to untangle. A related 2024 NBER paper from this research line found that chatbot outreach was most effective when it focused on discrete administrative processes such as financial aid forms and registration holds, especially when tasks were acute, required, and relevant to the student receiving the message. The new four-year paper is consistent with that logic. It shows that reducing administrative friction is plausible; what it does not show is that better task completion automatically turns into stronger grades or reenrollment. (nber.org)
The implementation lesson may matter more than the AI label
The deeper lesson across both papers is about problem selection and ownership. When the task was tightly bounded and close to the action — get this form in, clear this hold, use this tutoring option, pay attention to this course deadline — chatbots helped. When the hoped-for outcome was broad and distal, such as persistence or overall academic transformation, the evidence was much weaker. That is partly a finding from the papers and partly an inference from putting them side by side, but it is a grounded one. Both studies point toward workflows, timing, and relevance as the real mechanism of action. (edworkingpapers.com)
For college operators, that should reframe procurement conversations. A campuswide chatbot can be sustainable only if someone owns it centrally and can keep information current across offices, the four-year study argues. But the stronger academic effects in the Georgia State work came when the bot was course-specific, meaning its content, cadence, and referrals were tied to a professor’s syllabus and the supports surrounding a particular class. That creates a governance trade-off: centralization may help maintain infrastructure, while course-level specificity may be what actually moves grades. Buying a general AI assistant and hoping persistence rises is a much looser theory of change than either paper supports. (edworkingpapers.com)
There is also a useful caution here for institutions blending every AI product into one bucket. The course paper explicitly studied a non-generative chatbot, and the administrative paper focused on an AI-enabled text-messaging system built for targeted outreach. Neither paper tested the kind of broad, open-ended generative AI assistants now being marketed for everything from advising to writing help. In a year when EDUCAUSE says AI is reshaping student support and teaching, these studies are a reminder that evidence for narrow conversational systems should not be casually stretched into claims about general-purpose campus AI. (nber.org)
What remains uncertain is just as important as what changed. The evidence comes from specific institutional contexts, with overlapping research teams and, in the NBER course paper, proprietary Georgia State data. The studies are more rigorous than much of the AI-in-higher-ed marketing now circulating, but they do not yet establish how well similar tools will work in other course formats, other student populations, or newer LLM-based systems. (nber.org)
The next thing worth watching is whether the course-specific model holds up at larger scale and across more types of institutions. A U.S. Department of Education project abstract for TEACH ME says Georgia State and partners plan randomized trials of course-specific chatbot communication in gateway math and English courses across four diverse sites, involving more than 21,000 students and including cost analysis and implementation study. If colleges want evidence that can actually guide the next round of AI spending, that kind of multi-site follow-up — not another generic promise that chatbots boost student success — is what should matter most. (ed.gov)


