Abstract

Automatic Short Answer Scoring (ASAS) is central to NLP for Education. However, openly available benchmarks remain scarce, and existing datasets largely address how well students answer a question directly rather than how well they master underlying concepts (knowledge elements) such as thermal energy or epistemic activities (skills) such as reasoning or claim. To address this gap, we introduce Alice, a large-scale, rubric-based German ASAS dataset that is pedagogically aligned and comprises three subtasks: (i) learning performance (Alice-LP), (ii) knowledge elements (Alice-KE), and (iii) skills (Alice-SK). We further formulate rubric-based ASAS as a rubric-retrieval task and benchmark the dataset with a range of language models, from encoder-only models to lightweight LLMs. We also benchmark the dataset with zero-shot prompting via LLMs and a standard classification baseline. The experiments show that LLMs, in particular, struggle to score knowledge elements and skills in the zero-shot setting. They also indicate that rubric text is often useful, especially for Alice-KE and Alice-SK, while on Alice-LP gains over sample-solution-focused inputs are more modest and vary by model and input format.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Sun, Z., Gombert, S., Lossjew, J., Wyrwich, T., Czinczel, B. K., Bednorz, D., Kubsch, M., Neumann, K., & Drachsler, H. (2026). Alice: A Large-Scale German Benchmark for Rubric-Based Multi-Dimensional Automatic Short Answer Scoring. https://omanscience.com/en/articles/alice-a-large-scale-german-benchmark-for-rubric-based-multi-dimensional-automatic-short-answer-scoring

MLA 9

Sun, Zhifan, et al. "Alice: A Large-Scale German Benchmark for Rubric-Based Multi-Dimensional Automatic Short Answer Scoring." https://omanscience.com/en/articles/alice-a-large-scale-german-benchmark-for-rubric-based-multi-dimensional-automatic-short-answer-scoring.

Chicago (author–date)

Sun, Zhifan, Sebastian Gombert, Jannik Lossjew, Tobias Wyrwich, Berrit Katharina Czinczel, David Bednorz, Marcus Kubsch, Knut Neumann, and Hendrik Drachsler. 2026. "Alice: A Large-Scale German Benchmark for Rubric-Based Multi-Dimensional Automatic Short Answer Scoring." https://omanscience.com/en/articles/alice-a-large-scale-german-benchmark-for-rubric-based-multi-dimensional-automatic-short-answer-scoring.

Harvard

Sun, Z., Gombert, S., Lossjew, J., Wyrwich, T., Czinczel, B. K., Bednorz, D., Kubsch, M., Neumann, K. and Drachsler, H. (2026) 'Alice: A Large-Scale German Benchmark for Rubric-Based Multi-Dimensional Automatic Short Answer Scoring', Available at: https://omanscience.com/en/articles/alice-a-large-scale-german-benchmark-for-rubric-based-multi-dimensional-automatic-short-answer-scoring.

Vancouver

Sun Z, Gombert S, Lossjew J, Wyrwich T, Czinczel BK, Bednorz D, et al. Alice: A Large-Scale German Benchmark for Rubric-Based Multi-Dimensional Automatic Short Answer Scoring. https://omanscience.com/en/articles/alice-a-large-scale-german-benchmark-for-rubric-based-multi-dimensional-automatic-short-answer-scoring

IEEE

Z. Sun, S. Gombert, J. Lossjew, T. Wyrwich, B. K. Czinczel, D. Bednorz, M. Kubsch, K. Neumann, and H. Drachsler, "Alice: A Large-Scale German Benchmark for Rubric-Based Multi-Dimensional Automatic Short Answer Scoring," https://omanscience.com/en/articles/alice-a-large-scale-german-benchmark-for-rubric-based-multi-dimensional-automatic-short-answer-scoring.