TY - GEN
T1 - Judging the Judges
T2 - 14th Computing Conference, CC 2026
AU - He, Peng
AU - Li, Zhaohui
AU - Wang, Zeyuan
AU - Xiong, Jinjun
AU - Li, Tingting
N1 - Publisher Copyright:
© The Author(s), under exclusive license to Springer Nature Switzerland AG 2026.
PY - 2026
Y1 - 2026
N2 - Designing high-quality, standards-aligned instructional materials for K–12 science is time-consuming and expertise-intensive. This study examines what human experts notice when reviewing AI-generated evaluations of such materials, aiming to translate their insights into design principles for a future GenAI-based instructional material design agent. We intentionally selected 12 high-quality curriculum units across life, physical, and Earth sciences from validated programs such as OpenSciEd and Multiple Literacies in Project-Based Learning. Using the EQuIP rubric with 9 evaluation items, we prompted GPT-4o, Claude, and Gemini to produce numerical ratings and written rationales for each unit, generating 648 evaluation outputs. Two science education experts independently reviewed all outputs, marking agreement (1) or disagreement (0) for both scores and rationales, and offering qualitative reflections on AI reasoning. This process surfaces patterns where LLM judgments align with or diverge from expert perspectives, revealing reasoning strengths, gaps, and contextual nuances. These insights will directly inform the development of a domain-specific GenAI agent to support the design of high-quality instructional materials in K–12 science education.
AB - Designing high-quality, standards-aligned instructional materials for K–12 science is time-consuming and expertise-intensive. This study examines what human experts notice when reviewing AI-generated evaluations of such materials, aiming to translate their insights into design principles for a future GenAI-based instructional material design agent. We intentionally selected 12 high-quality curriculum units across life, physical, and Earth sciences from validated programs such as OpenSciEd and Multiple Literacies in Project-Based Learning. Using the EQuIP rubric with 9 evaluation items, we prompted GPT-4o, Claude, and Gemini to produce numerical ratings and written rationales for each unit, generating 648 evaluation outputs. Two science education experts independently reviewed all outputs, marking agreement (1) or disagreement (0) for both scores and rationales, and offering qualitative reflections on AI reasoning. This process surfaces patterns where LLM judgments align with or diverge from expert perspectives, revealing reasoning strengths, gaps, and contextual nuances. These insights will directly inform the development of a domain-specific GenAI agent to support the design of high-quality instructional materials in K–12 science education.
KW - EQuIP rubric
KW - Evaluation
KW - Human judgment
KW - LLM
KW - Science instructional materials
UR - https://www.scopus.com/pages/publications/105045695492
U2 - 10.1007/978-3-032-24807-7_31
DO - 10.1007/978-3-032-24807-7_31
M3 - Conference contribution
AN - SCOPUS:105045695492
SN - 9783032248060
T3 - Lecture Notes in Networks and Systems
SP - 481
EP - 496
BT - Intelligent Computing - Proceedings of the 2026 Computing Conference
A2 - Arai, Kohei
A2 - Lorenz, Pascal
PB - Springer Science and Business Media Deutschland GmbH
Y2 - 9 July 2026 through 10 July 2026
ER -