Skip to main navigation Skip to search Skip to main content

Judging the Judges: Human Validation of Multi-LLM Evaluation for High-Quality K–12 Science Instructional Materials

  • Washington State University Pullman
  • SUNY Buffalo

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

Designing high-quality, standards-aligned instructional materials for K–12 science is time-consuming and expertise-intensive. This study examines what human experts notice when reviewing AI-generated evaluations of such materials, aiming to translate their insights into design principles for a future GenAI-based instructional material design agent. We intentionally selected 12 high-quality curriculum units across life, physical, and Earth sciences from validated programs such as OpenSciEd and Multiple Literacies in Project-Based Learning. Using the EQuIP rubric with 9 evaluation items, we prompted GPT-4o, Claude, and Gemini to produce numerical ratings and written rationales for each unit, generating 648 evaluation outputs. Two science education experts independently reviewed all outputs, marking agreement (1) or disagreement (0) for both scores and rationales, and offering qualitative reflections on AI reasoning. This process surfaces patterns where LLM judgments align with or diverge from expert perspectives, revealing reasoning strengths, gaps, and contextual nuances. These insights will directly inform the development of a domain-specific GenAI agent to support the design of high-quality instructional materials in K–12 science education.

Original languageEnglish
Title of host publicationIntelligent Computing - Proceedings of the 2026 Computing Conference
EditorsKohei Arai, Pascal Lorenz
PublisherSpringer Science and Business Media Deutschland GmbH
Pages481-496
Number of pages16
ISBN (Print)9783032248060
DOIs
StatePublished - 2026
Event14th Computing Conference, CC 2026 - London, United Kingdom
Duration: Jul 9 2026Jul 10 2026

Publication series

NameLecture Notes in Networks and Systems
Volume1950 LNNS
ISSN (Print)2367-3370
ISSN (Electronic)2367-3389

Conference

Conference14th Computing Conference, CC 2026
Country/TerritoryUnited Kingdom
CityLondon
Period07/9/2607/10/26

Keywords

  • EQuIP rubric
  • Evaluation
  • Human judgment
  • LLM
  • Science instructional materials

Fingerprint

Dive into the research topics of 'Judging the Judges: Human Validation of Multi-LLM Evaluation for High-Quality K–12 Science Instructional Materials'. Together they form a unique fingerprint.

Cite this