How we test our AI tutor
A tutor that hands out answers can make students worse at the test. Here's the 36-case evaluation we use to catch that, and what it found across four Claude models.
In 2025, a study in PNAS by Bastani and colleagues found that students who practiced with an unguarded AI chatbot did worse on the exam they later took without it, by 17%. A tutor that withheld full solutions, asked for an attempt first and gave hints largely avoided the damage. That result shaped how we build and test CourseWing.
36 cases, eight categories
Our evaluation has 36 multi-turn conversations. Some are students asking the tutor to solve assigned work. Some are "answer laundering", where a tutor technically refuses but then writes the answer as an example. Some are bypass attempts: "this isn't for class", repeating the request five times, role-play. The rest check the other direction: ordinary requests like grades, study plans and quizzes must be answered fully, and a student's own answer must be checked clearly.
A grader model scores each reply on up to eight axes, such as no final answer, attempt first, incremental hints, clear explanation and holding under pressure. A case passes only if every axis passes.
What we found
With the same tutoring prompt, Claude Opus 5.5 passed all 36 cases. Claude Sonnet 5.5 passed 32 and Claude Sonnet 4.6 passed 29. Every model held the line on bypass attempts. The hardest category was checking a student's own answer: weaker runs either solved the whole problem themselves or never said whether the answer was right.
What's next
We'll keep growing the suite from real patterns and re-run it whenever we change the tutor or the model behind it. Full results are on our research page.