人工智能與數字教育觀察
AI and Digital Education Spectator

希伯來語化學考試:三個AI模型仍遜於中學生

研究以120道希伯來語高中化學公開試選擇題,測試ChatGPT-4o、Claude 3.5 Sonnet及Gemini 1.5 Pro,並與逾13.9萬名考生的資料比較。三個模型的整體表現均顯著低於學生,圖像及多步推理題尤其困難。研究再由化學教育專家分析最難題目的學科錯誤。結果對應受測版本及特定語言試卷,不能直接代表更新模型的能力。

論文原文

Benchmarking AI on Standard Chemistry Exams: LLMs Still Underperform Compared to High School Students
研究及來源資料
研究來自哪裏?
Weizmann Institute of Science
地域視角
世界各地研究 · 以色列
DOI
10.1007/s10956-026-10310-y
期刊/研討會
Journal of Science Education and Technology
類型
一百二十道高中化學題;三模型及既有考生成績比較
研究/內容地區
以色列
學習領域
科學教育
原文日期
2026-03-28
最近核實
2026-09-11
研究來源分類
期刊
作者
Elad Yacobson、Yael Schleifer、Ziva Bar-Dov、Shelley Rap、Ron Blonder、Giora Alexandron
學習階段
中學
研究領域
數學與科學、教學與評估
研究對象
人工智能系統、文獻及資料
分類說明
學生表現來自既有考試資料;研究直接測試模型,並非新的學生教學實驗。