ChatGPT Aces Intro Psych Exams, Stumbles in Advanced Courses, Challenging Academic Integrity

Join our daily newsletter for breaking news, product launches and deals, research breakdowns, and other industry-leading AI coverage

Join Now

ChatGPT outperforms undergrads in introductory psychology courses but struggles in higher-level classes, raising questions about the impact of AI on academic integrity and the effectiveness of AI detection tools.

Study tests ChatGPT’s performance on university psychology exams: Researchers at the University of Reading conducted an experiment where they submitted ChatGPT-generated answers to exam questions in five undergraduate psychology modules, spanning all three years of study:

The AI-generated submissions included both short 200-word answers and longer 1,500-word essays, with minimal editing or formatting applied to the AI-produced content.
Out of the 63 AI-generated submissions, 94% went undetected by markers, and nearly 84% received better grades than a randomly selected group of human students taking the same exams.

AI detection tools fall short in real-world scenarios: Despite claims of high accuracy in detecting AI-generated content, tools like OpenAI’s GPTZero and Turnitin’s AI writing detection system performed poorly when applied to the study’s AI-generated submissions:

OpenAI’s GPTZero has a reported 26% success rate in flagging AI-generated text as “likely” AI, with a concerning 9% false positive rate.
Turnitin’s system, while claiming 97% accuracy in detecting ChatGPT and GPT-3 content in lab settings, performed significantly worse when tested against the study’s real-world submissions.

Challenges for educators in the age of AI: The study’s findings highlight the difficulties faced by educators in identifying and addressing the use of AI-generated content in academic work:

Markers were surprised by the quality of the AI-generated submissions, with some being flagged not for being too repetitive or robotic, but for being too good.
The high success rate of ChatGPT in introductory-level courses raises concerns about the impact of AI on academic integrity and the need for effective detection methods.

Analyzing deeper: While ChatGPT’s strong performance in introductory psychology courses is notable, its struggles in higher-level classes suggest that AI may not yet be a comprehensive replacement for human knowledge and critical thinking skills. However, the study’s findings underscore the urgent need for educators to adapt to the rapidly evolving landscape of AI in academia, developing more robust detection tools and rethinking assessment methods to ensure academic integrity in the face of increasingly sophisticated AI systems.

ChatGPT outperforms undergrads in intro-level courses, falls short later

Ars Technica

Menu

ChatGPT Aces Intro Psych Exams, Stumbles in Advanced Courses, Challenging Academic Integrity

Recent News

Time Partners with OpenAI, Joining Growing Trend of Media Companies Embracing AI

AI Uncovers EV Adoption Barriers, Sparking New Climate Research Opportunities

AI Accelerates Disease Diagnosis: Earlier Detection, Novel Biomarkers, and Personalized Insights

Join the revolution

CO/AI

Resources

Join the revolution

Menu

Welcome

ChatGPT Aces Intro Psych Exams, Stumbles in Advanced Courses, Challenging Academic Integrity

Recent News

Time Partners with OpenAI, Joining Growing Trend of Media Companies Embracing AI

AI Uncovers EV Adoption Barriers, Sparking New Climate Research Opportunities

AI Accelerates Disease Diagnosis: Earlier Detection, Novel Biomarkers, and Personalized Insights

Join the revolution

CO/AI

Resources

Join the revolution