Psychologie in Erziehung und Unterricht
3
0342-183X
Ernst Reinhardt Verlag, GmbH & Co. KG München
10.2378/peu2026.art13d
3_073_2026_3/3_073_2026_3.pdf71
2026
733
Empirische Arbeit: One Model to Score Them All? On Suitability, Stability, and Synergy in Automated Essay Evaluation
71
2026
Andrea Horbach
Daniel Mora Melanchthon
Nils-Jonathan Schaller
Stefan Keller
Jennifer Meyer
Thorben Jansen
Writing is a key educational competence whose development depends on feedback. Automated Essay Scoring (AES) is increasingly used to support feedback generation, yet most prior research has neglected score stability across repeated runs. However, inconsistent scoring can undermine trust in AES and limit its educational value. We evaluate feature-based logistic regression models, transformer-based neural models, and generative large language models on 4,593 EFL essays from the MEWS dataset. Beyond accuracy, we analyze prediction variability across multiple runs and examine agreement patterns within and across model families. Results show that feature-based models remain competitive, while LLMs achieve high accuracy. Despite strong average performance, GPT-5 exhibits substantial variability across runs. Across models, agreement patterns reveal that different families succeed and fail on different item subsets. Our findings underline stability as a crucial dimension for deploying AES models in educational contexts and highlight the need for careful model selection and potentially model combination.
3_073_2026_3_0005
n Empirische Arbeit One Model to Score Them All? On Suitability, Stability, and Synergy in Automated Essay Evaluation Andrea Horbach 1 , Daniel Mora Melanchthon 1 , Nils-Jonathan Schaller 1 , Stefan Keller 2 , Jennifer Meyer 3 , Thorben Jansen 1 1 Leibniz-Institut für die Pädagogik der Naturwissenschaften und Mathematik, Kiel 2 Pädagogische Hochschule Zürich 3 Universität Wien Summary: Writing is a key educational competence whose development depends on feedback. Automated Essay Scoring (AES) is increasingly used to support feedback generation, yet most prior research has neglected score stability across repeated runs. However, inconsistent scoring can undermine trust in AES and limit its educational value. We evaluate feature-based logistic regression models, transformer-based neural models, and generative large language models on 4,593 EFL essays from the MEWS dataset. Beyond accuracy, we analyze prediction variability across multiple runs and examine agreement patterns within and across model families. Results show that feature-based models remain competitive, while LLMs achieve high accuracy. Despite strong average performance, GPT-5 exhibits substantial variability across runs. Across models, agreement patterns reveal that different families succeed and fail on different item subsets. Our findings underline stability as a crucial dimension for deploying AES models in educational contexts and highlight the need for careful model selection and potentially model combination. Keywords: Automated Essay Scoring, Natural Language Processing, Model Stability, Educational Assessment Ein Modell für alle Essays? Zur Eignung, Stabilität und Kombination automatisierter Bewertungsmodelle Zusammenfassung: Das Schreiben ist eine wichtige Bildungskompetenz, deren Entwicklung von Feedback abhängt. Automatisierte Essaybewertung (AES) wird zunehmend zur Unterstützung der Feedbackgenerierung eingesetzt, doch die Stabilität der Bewertungen bei wiederholten Durchläufen wurde bislang kaum untersucht. Inkonsistente Bewertungen können jedoch das Vertrauen in AES untergraben und dessen pädagogischen Wert einschränken. Wir bewerten featurebasierte logistische Regressionsmodelle, transformatorbasierte neuronale Modelle und generative große Sprachmodelle anhand von 4.593 EFL-Essays aus dem MEWS-Datensatz. Außer Genauigkeit analysieren wir die Variabilität der Vorhersagen über mehrere Durchläufe hinweg und untersuchen Übereinstimmungsmuster innerhalb und zwischen Modellfamilien. Die Ergebnisse zeigen, dass featurebasierte Modelle wettbewerbsfähig bleiben, während LLMs eine hohe Gesamtgenauigkeit erzielen. Trotz starker durchschnittlicher Leistung weist GPT-5 erhebliche Schwankungen zwischen den Durchläufen auf. Über alle Modelle hinweg zeigen Übereinstimmungsmuster, dass verschiedene Modellfamilien bei unterschiedlichen Texten erfolgreich sind. Unsere Ergebnisse unterstreichen Stabilität als entscheidende Dimension für den Einsatz von AES-Modellen im Bildungskontext und betonen die Notwendigkeit einer sorgfältigen Modellauswahl und möglicherweise einer Modellkombination. Schlüsselbegriffe: Automatisches Essayscoring, Natural Language Processing, Modellstabilität, Assessment im Bildungsbereich Dieser Beitrag steht open access online unter https: / / dx.doi.org/ 10.2378/ peu2026.art13d Psychologie in Erziehung und Unterricht, 2026, 73, 170 -183 DOI 10.2378/ peu2026.art13d © Ernst Reinhardt Verlag On Suitability, Stability, and Synergy in Automated Essay Evaluation 171 Writing is a target educational practice through which students develop language, content knowledge, and higher-order thinking (Graham, Kiuhara & MacKay, 2020). Learning to write is a gradual process that depends on practice and on feedback that helps learners reflect on and revise their texts. In classroom contexts, however, providing timely and consistent feedback on student writing is demanding for educators. Automated Essay Scoring (AES) has emerged as a technological response to support feedback generation. AES refers to the task of assigning a numeric or categorical score to a learner-written text. These texts often comprise a few hundred words and scores are assigned holistically (rating the essay as a whole) or trait-based, focusing on aspects such as language, content or organization. From a learning analytics perspective, this situates AES as a tool for generating measures of learner competencies (Gašević, Dawson & Siemens, 2015; Siemens & Baker, 2012), with implications for fairness, transparency, and stability in educational decision-making (Pardo & Siemens, 2014). Beyond scoring alone, automated assessment approaches are increasingly used as a basis for generating more informative feedback for learners and for studying how such feedback can be integrated into classroom practice (Gombert et al., 2024; Horbach et al., 2022). In recent decades, machine-learning-based models have been employed for this task, using a variety of linguistic indicators as features, while over the past 10 years, neural models became popular (Dong, Zhang & Yang, 2017; Taghipour & Ng, 2016; Uto, Xie & Ueno, 2020). Since the increasing popularity of generative large language models, they have been widely used for educational tasks with some promising results in few-shot scenarios for AES as well (Hou, Ciuba & Li, 2025; Mansour, Albatarni, Eltanbouly & Elsayed, 2024; Pack, Barrett & Escalante, 2024). When evaluating the suitability of technology to support feedback generation, accuracy is only one dimension. Other critical factors include fairness, bias, explainability, and stability (Zesch, Horbach & Zehner, 2023), with stability being essential for fostering trust in both the technology and the assessment decisions it supports. In this paper, we focus on the extent to which AES assigns the same score to the same student text when applied repeatedly (i. e., model stability), both within and across approaches. Specifically, we ask to what extent the performance of scoring models varies between different models (feature-based, neural, generative LLMs) as well as across multiple runs of the same model. In particular, our paper makes the following contributions: ◾ We present and analyze experiments on a large dataset of authentic argumentative learner essays in English (the MEWS dataset [Keller et al., 2024]) with two individual prompts and more than 4000 essays overall containing essays written by EFL learners with German as L1. ◾ We compare feature-based logistic regression, transformer-based and genLLM-based models. ◾ We especially analyze how often these models agree or not, whether essays that are hard for one model are also problematic for the other models and the other way round and whether there are items that are particularly easy or hard for all models. ◾ Finally, these analyses also answer the question whether one might benefit from a careful combination of models. The paper is organized as follows: we first discuss related work on AES and teacher assessment, we then introduce data and evaluation method and finally, we present our results. We conclude with a discussion. Related Work Teachers’ assessments of students’ writing vary when multiple teachers assess the same texts (e. g., Birkel & Birkel, 2002) or when teachers assess the same text multiple times (Li, 2022). Hence, research on writing assessment has long focused on the reliability of teachers’ assess- 172 Andrea Horbach et al. ments. By contrast, AES research has largely operationalized “quality” as single-run predictive performance against a human gold standard, implicitly assuming that a model’s output is stable once accuracy is acceptable. For a general review of essay scoring approaches, see Bai & Stede, 2023; Ke & Ng, 2019; Uto, 2021. Here, we focus on three strands most relevant to our study. Feature-based, neural, and LLMbased AES. Feature-based models have long dominated AES (Attali & Burstein, 2006; Meyer, Jansen, Fleckenstein, Keller & Köller, 2020), before neural architectures such as LSTMs and transformers (Dong, Zhang & Yang, 2017; Horbach, Scholten-Akoun, Ding & Zesch, 2017; Taghipour Ng, 2016) shifted the field towards end-to-end learning. More recently, research has increasingly explored fine-tuned pretrained language models and longer-context transformer architectures for essay scoring, including work on transformer-based AES, implicit rubric modeling, long-context scoring, and improved generalizability across prompts (Bexte, Ding & Horbach, 2025; Fiacco, Adamson & Rose, 2023; Ormerod & Kehat, 2025). Generative LLMs, in turn, have been used for AES with rubricbased prompting and few-shot setups (Hou, Ciuba & Li, 2025; Mansour, Albatarni, Eltanbouly & Elsayed, 2024; Pack, Barrett & Escalante, 2024). The idea of using ensembles of machine learning models instead of a single model has, for example, been explored by Ormerod et al. (2023), showing that combining models can improve performance. Prior work has emphasized not only scoring performance but also dimensions such as fairness and bias (Andersen, Mang, Goldhammer & Zehner, 2025; Loukina, Madnani & Zechner, 2019; Schaller, Ding, Horbach, Meyer & Jansen, 2024), as well as validity (Attali, 2013) and reliability, an issue not only important when evaluating automated scoring but also human raters (Brookhart et al., 2016). Model stability is a topic that received substantial attention in the machine learning community (see Summers & Dinneen, 2021). In the context of AES, García-Varela, Nussbaum, Mendoza, Martínez-Troncoso and Bekerman (2025) showed that ChatGPT’s scoring consistency improves only when strict rubrics and output constraints are applied. Similarly, Gaggioli et al. (2025) reported that repeated runs on the same essay often yield divergent results. These findings indicate that stability is an emerging concern in AES research, yet systematic comparisons across model families remain rare. We try to close this gap with our study. Method Data We use the MEWS data set (Measuring Writing Skills in English as a Second Language; Fleckenstein, Meyer, Jansen, Keller & Köller, 2020; Rupp, Casabianca, Krüger, Keller & Köller, 2019). This data set contains essays written by English-as-a-foreign-language learners with a German L1 background from Germany and Switzerland at the upper secondary level. They responded to prompts from the TOEFL language proficiency test. The two prompts were as follows. The first task, “Teacher”, was: “A teacher’s ability to relate well with students is more important than excellent knowledge of the subject being taught.” The second task, called “Advertising”, asked the learners to react to the statement “Television advertising directed toward young children (aged two to five) should not be allowed.” Students were asked to write an essay in the agree-disagree format. Table 1 shows one such essay as an example. TV Advertising Television has two effect on the lieving young childeren. It can use for reach to informations and also can destroied their manner. The young childeren can give informations with watch television such that maybe with read books they can not reach to those. In constract, the television can get the bad effect on their manner. For example, they see the evil movie and however destroied family communication. My opinion television is not interesting. Also, the young childeren spend very time for watch television such that is not enough useful for their. Table 1: An example essay from the MEWS dataset, taken from Rupp et al., 2019. On Suitability, Stability, and Synergy in Automated Essay Evaluation 173 We include a total of 4593 essays with an average length of 339 tokens per text answering two individual writing prompts. The texts have been scored by expert raters according to a number of categories, we use in our experiments the holistic scores, ranging on a scale from 0 to 5, in steps of 0.5. Annotators always assigned full points, half points arose from a disagreement between two annotators. For evaluation, we construct a test set of 500 essays, sampled in a stratified manner to obtain a distribution that is as balanced as possible across score levels. However, since the original distribution in the full dataset is skewed, the resulting test set is also not perfectly balanced (see Table 2). Scoring Models We experiment with three model types: First, we leverage feature-based models using a logistic regression as their backbone by means of a scikit-learn implementation (Pedregosa et al., 2011) in its base configuration. We use a set of over 200 linguistic features proved to be effective on this dataset by Lohmann et al. (2024). All features are normalized using Min-Max scaling, which transforms the variables to a fixed range between 0 and 1. Second, we use a DistilBERT model (distilbertbase-uncased; Sanh, Debut, Chaumond & Wolf, 2019) where we fine-tune a regression head - either using the model on its own or as a hybrid model by combining the max pooled representation of the final hidden layer from the DistilBERT encoder with a feedforward network that processes the linguistic features. The two representations are concatenated and passed through a shared regression head composed of two linear layers. Third, we explore several commercial generative LLM variants via their API, namely ChatGPT-4.1 (gpt-4.1-2025-04-14), ChatGPT-4o mini and ChatGPT-5 (gpt-5-2025-08-07). We use a few-shot scenario, where we prompt the model to take on the role of an expert evaluator and provide it with the same rubric description that was used by the human annotators of the dataset. Additionally, all models get the same two example essays, one with a score of 2 and one with a score of 4 as part of the rubric description. The models were instructed to assign half-point scores (e. g., 2.5, 3.5) when essays fell short of fully meeting the criteria for the next higher score. For GPT-4.1 and GPT-4o mini, we use a low-temperature setting (temperature = 0.1) to reduce stochastic variation in the outputs, a maximum output length of 2048 tokens, and the default value for top_p (top_p = 1.0). For GPT-5, generation parameters such as temperature and top_p are not user-configurable via the API. The structure of the prompt can be found in the appendix. Please note that we formulate the prediction task as a regression problem rather than a classification problem. The target labels lie on an ordered rating scale with meaningful distances between adjacent score points (including half-point increments). A regression formulation allows the model to exploit this ordinal structure and capture similarities between neighboring score levels, which is particularly beneficial for low-frequency score categories at the extremes of the distribution. We also experimented with classification-based formulations in preliminary experiments. However, these yielded consistently lower performance in our setting, likely due to the skewed label distribution and the sparse representation of extreme score categories. Therefore, we focus on regression-based models throughout this study. Machine Learning Setup For the genAI models, we use few-shot prompting methods as described above on the test data set. All other machine learning models are trained on data from the MEWS dataset in the following way: As the dataset is highly skewed, using the remainder of the dataset after splitting away the test data would leave us with no instances of the lowest and highest scores for training. Therefore, we split the test data section into 5 stratified subsections and train a model on the training data section plus four of the five test data sections and then test on the fifth test data section and repeat this process five times. Thus, we ensure that we can report performance on the full test data while still using every instance either for training or for testing in a single run. label 0.0 1.0 1.5 2.0 2.5 3.0 3.5 4.0 4.5 5.0 test 10 31 35 62 62 62 61 61 61 55 train + test 10 31 35 340 540 1462 951 888 244 92 Table 2: Label distribution in the full MEWS dataset and in the test section. 174 Andrea Horbach et al. As we are interested in model stability, we generate different model variants in the following ways: For generative LLM-based models we simply run the model several times. We also measure the influence of using few-shot examples for generative LLMs by resampling these few-shot items several times. For the regression-based model, different runs arise from the fitted parameters of the Min-Max scaler and the multicollinearity removal procedure, introducing slight variability in the preprocessing step. For the transformer-based and hybrid models, we use different weight initializations for the different runs. Evaluation Metrics We assess general model performance by a number of established evaluation metrics, namely quadratically weighted kappa (QWK; Cohen, 1968) and Pearson correlation. Experimental Studies In this section, we present our evaluation experiments starting with performance and then measuring model stability and the potential for model combination. Baseline Scoring Performance We first establish model performance as the basis for further analyses. We measure QWK and Pearson correlation for the individual models and report average performance over 10 runs as well as values for the best and worst runs to get a first hint of model stability. We also evaluate performance if we aggregate predictions per item using either majority voting (with ties broken randomly) or by averaging over individual runs. Table 3 summarizes the results. Overall, all models perform reasonably well with a performance range of 0.67 to 0.79 QWK and a Pearson correlation of 0.8 to 0.87. It is especially noteworthy that the computationally cheapest model, the Logistic Regression, is among the best-performing models. We also see, that best and worst model runs are not far from each other per model. When comparing aggregation methods, differences are small with a slight advantage for averaging. Thus we use the averaged model predictions in further analyses when we need a single outcome per model. Name QWK (Runs) QWK (Avg) QWK (MV) Pearson (Runs) Pearson (Avg) Pearson (MV) GPT-5 GPT-41 GPT-4o LogRes DistilBert Hybrid 0.672 [0.665, 0.683] 0.785 [0.748, 0.809] 0.708 [0.703, 0.715] 0.790 [0.780, 0.796] 0.673 [0.639, 0.690] 0.779 [0.757, 0.797] 0.675 0.791 0.711 0.793 0.674 0.783 0.673 0.795 0.710 0.786 0.673 0.790 0.829 [0.814, 0.836] 0.863 [0.848, 0.874] 0.797 [0.792, 0.804] 0.817 [0.809, 0.824] 0.807 [0.796, 0.820] 0.851 [0.827, 0.869] 0.839 0.867 0.802 0.820 0.809 0.856 0.831 0.868 0.798 0.812 0.809 0.859 Table 3: Baseline model performance measure in QWK and SD across 10 runs. 0.0 0.5 1.0 1.5 2.0 2.5 3.0 3.5 4.0 4.5 5.0 Mean Score GPT-5 GPT-4.1 GPT-4o LogRes DistilBert Hybrid True Labels 0.7 1.4 1.5 2.0 0.0 0.0 2.0 0.0 0.2 0.2 1.6 0.0 0.1 0.0 7.8 4.3 3.1 1.3 0.2 1.3 6.2 3.7 2.1 0.8 2.9 2.1 5.4 7.0 27.0 16.6 7.7 6.5 6.3 8.1 12.4 13.3 18.5 21.2 13.9 15.7 14.1 12.4 30.7 24.5 21.8 23.6 26.5 19.9 12.4 8.1 12.1 32.0 22.9 28.3 25.5 12.2 8.1 20.0 11.6 16.1 18.2 21.0 12.2 0.7 0.2 0.0 7.1 2.6 4.4 12.2 0.0 0.1 0.0 2.2 0.1 0.1 11.0 2.566 2.842 2.967 3.122 3.188 3.146 3.111 Table 4: Distribution of predicted labels (percentages) per model and mean score. On Suitability, Stability, and Synergy in Automated Essay Evaluation 175 max agree gpt 5 gpt 4o gpt 4.1 LogRes DistilBert Hybrid 10 9 8 7 6 5 4 3 2 1 28.46 14.63 13.03 14.63 16.03 10.62 2.61 0.00 0.00 0.00 81.36 5.61 3.21 2.61 4.61 2.61 0.00 0.00 0.00 0.00 77.15 7.62 5.41 3.01 5.61 1.00 0.20 0.00 0.00 0.00 78.76 6.81 4.61 4.01 3.81 2.00 0.00 0.00 0.00 0.00 54.31 17.03 9.22 7.62 8.62 3.21 0.00 0.00 0.00 0.00 49.10 17.23 8.42 8.22 10.02 6.01 0.80 0.20 0.00 0.00 Table 5: Distribution of maximum number of agreeing runs (out of 10) per item, by model. # correct runs within ±0.5 gpt 5 gpt 4o gpt 4.1 LogRes DistilBert Hybrid 0 1 2 3 4 5 6 7 8 9 10 29.86 6.01 3.01 3.01 2.61 2.00 2.40 3.21 3.21 6.61 38.08 34.07 1.80 0.60 0.40 0.80 0.40 0.40 0.20 0.40 0.60 60.32 26.25 1.20 0.80 0.00 1.00 0.80 1.00 0.80 1.00 0.80 66.33 19.64 1.00 1.40 1.00 0.40 1.00 1.20 0.80 0.60 0.60 72.34 25.25 2.61 2.40 1.40 2.00 1.00 1.20 2.40 2.20 3.41 56.11 15.03 2.20 0.80 0.80 1.00 2.40 2.61 2.20 2.20 5.41 65.33 Table 6: Distribution of benchmark predictions across 10 runs per item, allowing ±0.5 tolerance, by model. Name QWK Pearson GPT-5 GPT-4.1 GPT-4o LogRes DistilBert Hybrid Ensemble (Avg) Ensemble (MV) Oracle 0.675 0.791 0.711 0.793 0.674 0.783 0.786 0.795 0.932 0.839 0.867 0.802 0.820 0.809 0.856 0.892 0.855 0.954 Table 7: Model performance when using model ensembles and an Oracle condition choosing the optimal model per item. max agree Percentage 0 1 1.5 2 2.5 3 3.5 4 4.5 5 1 2 3 4 5 6 0.2 % 9.8 % 42.4 % 26.6 % 17.4 % 3.6 % 0 3 5 2 0 0 1 5 16 7 1 1 0 5 16 10 4 0 0 3 28 21 9 1 0 5 24 19 12 2 0 7 19 21 12 3 0 2 30 19 9 1 0 2 28 16 12 3 0 10 29 10 9 3 0 7 17 8 19 4 Table 8: Distribution of maximum number of agreeing models per item and the true score distribution for the items in each category. 176 Andrea Horbach et al. To get a better understanding of the differences between models, Table 4 shows the frequency distribution (percent-ages) of individual labels per model and also the frequency distribution for the true labels (i. e., values per line sum up to 100). We see that genAI models tend to underestimate scores, with GPT-5 assigning a score that is more than 0.5 points lower on average than the actual gold standard score. By contrast, nongenerative models’ average predictions are closer to the ground truth, sometimes rather slightly overestimating student performance. Another key finding is that all models are reluctant to assign the highest possible scores of 4.5 and 5.0. Model Diversity Next, we turn to our main interest: the diversity and stability of results. Agreement across runs Our first analysis examines the stability of each model by asking: For a given item, how many of the 10 runs of the same model produced identical scores? Using the run definition introduced in the Method section, this refers to repeated training runs for logistic regression, DistilBERT, and the hybrid model with slightly varying preprocessing or initialization conditions, and to repeated calls for the generative LLM-based models. Table 5 reports the distribution of the maximum number of agreeing runs per item. The first row shows the percentage of items for which all 10 runs agreed, whereas the last row shows the percentage of items where no two runs produced the same score. Results indicate substantial differences between models: GPT-5, and to a lesser extent the DistilBERT-based models, display considerably more variability across runs. Thus, in a real-life educational setting, students might get different scores for the same text. Agreement with the benchmark score The previous analysis only considered whether the models agreed or not, not whether they actually agreed on the benchmark score provided by human annotators. We therefore next examine how many of the 10 runs per item and model predicted the benchmark score (allowing a tolerance of ±0.5 points). Table 6 shows the results. For example, the first row indicates that for 29.9 % of all items all 10 GPT-5 runs predicted the wrong score (with a tolerance of 0.5), for 6 % of all items (second row) only one GPT-5 run was correct and so on. For 38.1 % of all items all 10 GPT-5 runs predicted the benchmark score within the tolerance margin. We observe a similar pattern as before: GPT-4.1, GPT-4o, and logistic regression achieve the highest proportions of consistently correct predictions, while the hybrid model also performs well, indicating that this model often differs from the benchmark score by half a point. Potential for Model Combination Having established stability measures within models, we now compare predictions across models. For that, we use the average of the 10 runs of a model as this models’ prediction. Prediction Aggregation We explore two aggregation methods as discussed before for the combination of model runs: averaging and majority voting per item. Table 7 shows that we do not gain much from combining these models. We also evaluate an oracle condition: What would the performance be if an oracle told us which single model to ask for its score? I. e., what is the potential in a careful model combination? Unsurprisingly, such a scenario (which would be unrealistic in reallife) comes with a performance boost. Agreement between pairs of models To get a better idea of which models agree with each other, Figure 1 shows the agreement between pairs of models. The three generative and the three non-generative models form blocks of slightly higher pairwise agreement, while GPT-5 has a particularly low agreement with these models. On Suitability, Stability, and Synergy in Automated Essay Evaluation 177 Figure 1: Confusion matrices between models. (a) Pairwise QWK (avg predictions) 1.0 0.8 0.6 0.4 0.2 0.0 Gold GPT-5 GPT-4.1 GPT-4o LogRes DistilBert Hybrid Gold GPT-5 GPT-4.1 GPT-4o LogRes DistilBert Hybrid 1.000 0.675 0.791 0.711 0.793 0.674 0.783 0.675 1.000 0.856 0.763 0.623 0.566 0.630 0.791 0.856 1.000 0.895 0.745 0.680 0.754 0.711 0.763 0.895 1.000 0.707 0.670 0.716 0.793 0.623 0.745 0.707 1.000 0.754 0.869 0.674 0.566 0.680 0.670 0.754 1.000 0.907 0.783 0.630 0.754 0.716 0.869 0.907 1.000 (b) Pairwise Pearson (avg predictions) 1.00 0.75 0.50 0.25 0.00 -0.25 -0.50 -0.75 -1.00 Gold GPT-5 GPT-4.1 GPT-4o LogRes DistilBert Hybrid Gold GPT-5 GPT-4.1 GPT-4o LogRes DistilBert Hybrid 1.000 0.839 0.867 0.802 0.820 0.809 0.856 0.839 1.000 0.906 0.860 0.756 0.780 0.792 0.867 0.906 1.000 0.911 0.781 0.776 0.807 0.802 0.860 0.911 1.000 0.735 0.712 0.736 0.820 0.756 0.781 0.735 1.000 0.813 0.886 0.809 0.780 0.776 0.712 0.813 1.000 0.925 0.856 0.792 0.807 0.736 0.886 0.925 1.000 178 Andrea Horbach et al. Figure 2: Top: frequency of coalitions of models predicting the same score. Bottom: only coalitions with correct predictions. Agreement coalitions - ALL clusters counted (incl. Singletons) Agreement coalitions - ONLY CORRECT clusters (incl. Singletons) On Suitability, Stability, and Synergy in Automated Essay Evaluation 179 Agreement across models Similar to the agreement across runs discussed before, we now evaluate how many of the 6 models produced identical scores for a given item. Table 8 reports the distribution of the maximum number of agreeing models. For the majority of items, only 3 or 4 models agree while there seems to be a tendency that models rather agree on items with higher scores. To further investigate model agreement, Figure 2 presents UpSet plots (Lex, Gehlenborg, Strobelt, Vuillemot & Pfister, 2014), which visualize how frequently certain groups of models predict the same score. The upper panel shows all model coalitions, specifying for how many items a certain group of one or more models agreed on a score, while the lower panel shows only coalitions that produced the correct prediction and their frequency. These plots complement the confusion matrix analysis by revealing which subsets of models tend to align. For example, GPT-5 often disagrees with other models: in 204 runs, it produces a different prediction than the other models, but only for 31 runs, it is correct in doing so. Also, strong agreement is observed within coalitions of generative versus non-generative models. This highlights that different model families exhibit distinct strengths and weaknesses, underlining the importance of model choice in practical applications. Discussion and Conclusion The main finding of this study is that models with comparable levels of predictive accuracy can nevertheless differ in stability. In particular, logistic regression and GPT-4.1 achieved strong accuracy with relatively consistent results across runs, whereas GPT-5 and DistilBERT models displayed greater variability. This indicates that, for practical use, AES models should be evaluated not only for accuracy, but also for stability. As our study shows, accuracy on its own can be a deceptive concept because it can be unstable. Across model families, some items were particularly easy or difficult across all models, while others divided the models, indicating complementary strengths. Although ensemble methods offered only marginal improvements, an oracle analysis suggested that carefully selecting models per item could yield substantial gains. Future work should thus concentrate on the problem of which model is most suitable for a particular item. Our results underline the importance of looking beyond accuracy when evaluating AES models. Stability across runs and across models is critical for educational use, where inconsistency can undermine trust in automated assessments. For future studies of applications in practice, this implies that researchers should include a research phase in which they determine which model is most suitable for the task at hand. Moreover, these findings should be set in relation to ethical implications, including fairness across learner groups (Schaller et al., 2024), transparency of model decisions (Hall, Seyam & Dunlap, 2024), and the cost (Oketch, Lalor, Yang & Abbasi, 2025) and environmental impact (Faiz et al., 2023) of large-scale deployment. Beyond model selection, our study has various implications for pedagogical practice. When using such tools to support learning, it is reasonable to provide learners with the best-working model. They need information that is as accurate as possible to gain information on how to improve. However, it is equally important to instill learners with a sense that they should not trust such results blindly, but rather that there is a degree of uncertainty behind these computational models. For this purpose, it may be fruitful to expose them to score distributions derived from multiple models and multiple model runs. Variability and distributions in these scores can make them aware that automated assessments of their writings are always approximations. What is equally, if not more important in teaching writing is equipping learners with a sense of agency: They need to decide actively what they want to achieve in their writing and how they want to pursue this. Relying on automated 180 Andrea Horbach et al. scores for feedback and assessment can be helpful in this process, but human agency should always be at the forefront of this process. Limitations Our study is limited in several ways. First, we analyze a single dataset with essays from two prompts only, which restricts the generalizability of our findings. Second, we focus on holistic scores, whereas trait-based evaluation may reveal additional insights into model behavior. Third, our set of models is not exhaustive: we investigate one feature-based baseline, one transformer model, a hybrid model (all using large amounts of training data) and selected commercial LLMs in a few-shot setup without any fine-tuning of the models. We do not systematically vary architectures, hyperparameters, or prompting strategies. While we also explored smaller open-weight LLMs in preliminary experiments, their performance in our setup was substantially lower. A systematic comparison of openand closedweight LLMs remains an important direction for future work. Finally, our stability analyses cover variability across runs within a given setup, but not longitudinal changes in models over time (e. g., API updates to commercial LLMs). Ethics Statement This study does not involve the collection of new data or user studies. All analyses are based on the MEWS dataset, which does not contain personally identifiable information. To the best of our knowledge, the dataset providers ensured appropriate data protection and ethical handling. Consequently, no additional ethical approval was required for this work. Nevertheless, the educational use of Automated Essay Scoring raises broader ethical considerations. Model variability, bias against certain learner groups, and the transparency of automated decisions may affect the fairness and trustworthiness of such systems. While our study does not directly address these aspects, our focus on stability contributes to the responsible evaluation of AES models in educational contexts. References Andersen, N., Mang, J., Goldhammer, F. & Zehner, F. (2025). Algorithmic fairness in automatic short answer scoring. International Journal of Artificial Intelligence in Education, 35, 3128 - 3165. https: / / doi.org/ 10.1007/ s40593-025-00495-5 Attali, Y. (2013). Validity and reliability of automated essay scoring. In M. D. Shermis & J. Burstein (Eds.), Handbook of automated essay evaluation (pp. 181-198). Routledge. https: / / doi.org/ 10.4324/ 9780203122761-19 Attali, Y. & Burstein, J. (2006). Automated essay scoring with e-rater® V.2. The Journal of Technology, Learning, and Assessment, 4 (3), 1 - 29. Bai, X. & Stede, M. (2023). A survey of current machine learning approaches to student free-text evaluation for intelligent tutoring. International Journal of Artificial Intelligence in Education, 33 (4), 992 - 1030. https: / / doi. org/ 10.1007/ s40593-022-00323-0 Bexte, M., Ding, Y. & Horbach, A. (2025). Increasing the generalizability of similarity-based essay scoring through cross-prompt training. In E. Kochmar, B. Alhafni, M. Bexte, J. Burstein, A. Horbach, R. Laarmann-Quante, … Z. Yuan (Eds.), Proceedings of the 20th Workshop on Innovative Use of NLP for Building Educational Applications (BEA 2025) (pp. 225 - 236). Association for Computational Linguistics. https: / / doi.org/ 10.18653/ v1/ 2025.bea-1.17 Birkel, P. & Birkel, C. (2002). Wie einig sind sich Lehrer bei der Aufsatzbeurteilung? Eine Replikationsstudie zur Untersuchung von Rudolf Weiss. Psychologie in Erziehung und Unterricht, 49 (3), 219 - 224. Brookhart, S. M., Guskey, T. R., Bowers, A. J., McMillan, J. H., Smith, J. K., Smith, L. F., … Welsh, M. E. (2016). A century of grading research: Meaning and value in the most common educational measure. Review of Educational Research, 86 (4), 803 - 848. https: / / doi.org/ 10.3102/ 0034654316672069 Cohen, J. (1968). Weighted kappa: Nominal scale agreement provision for scaled disagreement or partial credit. Psychological Bulletin, 70 (4), 213 - 220. https: / / doi.org/ 10.1037/ h0026256 Dong, F., Zhang, Y. & Yang, J. (2017). Attention-based recurrent convolutional neural network for automatic essay scoring. In: Proceedings of the 21st Conference on Computational Natural Language Learning (CoNLL 2017) (pp. 153 - 162). Association for Computational Linguistics. https: / / doi.org/ 10.18653/ v1/ K17-1016 Faiz, A., Kaneda, S., Wang, R., Osi, R., Sharma, P., Chen, F. & Jiang, L. (2023). LLMCarbon: Modeling the endto-end carbon footprint of large language models. arXiv. https: / / arxiv.org/ abs/ 2309.14393 Fiacco, J., Adamson, D. & Rose, C. (2023). Towards extracting and understanding the implicit rubrics of transformer based automatic essay scoring models. In E. Kochmar, J. Burstein, A. Horbach, R. Laarmann- Quante, N. Madnani, A. Tack, … T. Zesch (Eds.), Proceedings of the 18th Workshop on Innovative Use of NLP for Building Educational Applications (BEA 2023) (pp. 232 - 241). Association for Computational Linguistics. https: / / doi.org/ 10.18653/ v1/ 2023.bea-1.20 Fleckenstein, J., Meyer, J., Jansen, T., Keller, S. & Köller, O. (2020). Is a long essay always a good essay? The effect of text length on writing assessment. Frontiers in Psychology, 11, Article 562462. https: / / doi.org/ 10.33 89/ fpsyg.2020.562462 On Suitability, Stability, and Synergy in Automated Essay Evaluation 181 Gaggioli, A., Casaburi, G., Ercolani, L., Collovà, F., Torre, P. & Davide, F. (2025). Assessing the reliability and validity of large language models for automated assessment of student essays in higher education. arXiv. https: / / arxiv.org/ abs/ 2508.02442 García-Varela, F., Nussbaum, M., Mendoza, M., Martínez- Troncoso, C. & Bekerman, Z. (2025). ChatGPT as a stable and fair tool for automated essay scoring. Education Sciences, 15 (8), Article 946. https: / / doi.org/ 10.33 90/ educsci15080946 Gašević, D., Dawson, S. & Siemens, G. (2015). Let’s not forget: Learning analytics are about learning. TechTrends, 59 (1), 64 - 71. https: / / doi.org/ 10.1007/ s11528-014- 0822-x Gombert, S., Fink, A., Giorgashvili, T., Jivet, I., Di Mitri, D., Yau, J., … Drachsler, H. (2024). From the automated assessment of student essay content to highly informative feedback: A case study. International Journal of Artificial Intelligence in Education, 34 (4), 1378 - 1416. https: / / doi.org/ 10.1007/ s40593-023-00387-6 Graham, S., Kiuhara, S. A. & MacKay, M. (2020). The effects of writing on learning in science, social studies, and mathematics: A meta-analysis. Review of Educational Research, 90 (2), 179 - 226. https: / / doi.org/ 10.3102/ 00 34654320914744 Hall, E., Seyam, M. & Dunlap, D. (2024). Exploring explainability and transparency in automated essay scoring systems: A user-centered evaluation. In P. Zaphiris & A. Ioannou (Eds.), Learning and Collaboration Technologies (pp. 266 - 282). Springer. https: / / doi.org/ 10. 1007/ 978-3-031-61691-4_18 Horbach, A., Laarmann-Quante, R., Liebenow, L., Jansen, T., Keller, S., Meyer, J., … Fleckenstein, J. (2022). Bringing automatic scoring into the classroom: Measuring the impact of automated analytic feedback on student writing performance. In: Proceedings of the 11th Workshop on Natural Language Processing for Computer-Assisted Language Learning (NLP4CALL 2022) (pp. 72 - 83). LiU Electronic Press. https: / / doi.org/ 10. 3384/ ecp190008 Horbach, A., Scholten-Akoun, D., Ding, Y. & Zesch,T. (2017). Fine-grained essay scoring of a complex writing task for native speakers. In: Proceedings of the 12th Workshop on Innovative Use of NLP for Building Educational Applications (pp. 357 - 366). Association for Computational Linguistics. https: / / doi.org/ 10.18653/ v1/ W17-5040 Hou, Z., Ciuba, A. & Li, X. (2025). Improving LLM-based automatic essay scoring with linguistic features. In Z. Wang, S. Woodhead, M. Ananda, D. B. Mallick, J. Sharpnack & J. Burstein (Eds.), Proceedings of the Innovation and Responsibility in AI-Supported Education Workshop. Proceedings of Machine Learning Research, 273, 41 - 65. PMLR. https: / / proceedings. mlr.press/ v273/ hou25a.html Ke, Z. & Ng, V. (2019). Automated essay scoring: A survey of the state of the art. In: Proceedings of the 28th International Joint Conference on Artificial Intelligence (pp. 6300 - 6308). AAAI Press. https: / / doi.org/ 10.24963/ ijcai.2019/ 879 Keller, S. D., Lohmann, J. F., Trüb, R., Fleckenstein, J., Meyer, J., Jansen, T. & Möller, J. (2024). Language quality, content, structure: What analytic ratings tell us about EFL writing skills at upper secondary school level in Germany and Switzerland. Journal of Second Language Writing, 65, Article 101129. https: / / doi.org/ 10.1016/ j.jslw.2024.101129 Lex, A., Gehlenborg, N., Strobelt, H., Vuillemot, R. & Pfister, H. (2014). UpSet: Visualization of intersecting sets. IEEE Transactions on Visualization and Computer Graphics, 20 (12), 1983-1992. https: / / doi.org/ 10.1109/ TVCG.2014.2346248 Li, W. (2022). Scoring rubric reliability and internal validity in rater-mediated EFL writing assessment: Insights from many-facet Rasch measurement. Reading and Writing, 35 (10), 2409 - 2431. https: / / doi.org/ 10.1007/ s11145-022-10279-1 Lohmann, J. F., Junge, F., Möller, J., Fleckenstein, J., Trüb, R., Keller, S., Jansen, T. & Horbach, A. (2025). Neural networks or linguistic features? Comparing different machine-learning approaches for automated assessment of text quality traits among L1and L2-learners’ argumentative essays. International Journal of Artificial Intelligence in Education, 35, 1178 - 1217. https: / / doi. org/ 10.1007/ s40593-024-00426-w Loukina, A., Madnani, N. & Zechner, K. (2019). The many dimensions of algorithmic fairness in educational applications. In: Proceedings of the Fourteenth Workshop on Innovative Use of NLP for Building Educational Applications (pp. 1 - 10). Association for Computational Linguistics. https: / / doi.org/ 10.18653/ v1/ W19- 4401 Mansour, W., Albatarni, S., Eltanbouly, S. & Elsayed, T. (2024). Can large language models automatically score proficiency of written essays? arXiv. https: / / arxiv.org/ abs/ 2403.06149 Meyer, J., Jansen, T., Fleckenstein, J., Keller, S. & Köller, O. (2020). Machine learning im Bildungskontext: Evidenz für die Genauigkeit der automatisierten Beurteilung von Essays im Fach Englisch. Zeitschrift für Pädagogische Psychologie, 34 (3 - 4), 159 - 171. https: / / doi.org/ 10.1024/ 1010-0652/ a000296 Oketch, K., Lalor, J. P., Yang, Y. & Abbasi, A. (2025). Bridging the LLM accessibility divide? Performance, fairness, and cost of closed versus open LLMs for automated essay scoring. arXiv. https: / / arxiv.org/ abs/ 2503.118 27 Ormerod, C., Lottridge, S., Harris, A. E., Patel, M., van Wamelen, P., Kodeswaran, B., Woolf, S. & Young, M. (2023). Automated short answer scoring using an ensemble of neural networks and latent semantic analysis classifiers. International Journal of Artificial Intelligence in Education, 33 (3), 467 - 496. https: / / doi.org/ 10.10 07/ s40593-022-00294-2 Ormerod, C. & Kehat, G. (2025). Long context automated essay scoring with language models. In: Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con) - Volume 1: Full Papers (pp. 35 - 42). Pack, A., Barrett, A. & Escalante, J. (2024). Large language models and automated essay scoring of English language learner writing: Insights into validity and reliability. Computers and Education: Artificial Intelligence, 6, Article 100234. https: / / doi.org/ 10.1016/ j.caeai.2024. 100234 Pardo, A. & Siemens, G. (2014). Ethical and privacy principles for learning analytics. British Journal of Educational Technology, 45 (3), 438 - 450. https: / / doi.org/ 10. 1111/ bjet.12152 Pedregosa, F., Varoquaux, G., Gramfort, A., Michel, V., Thirion, B., Grisel, O., … Weiss, R. (2011). Scikit-learn: Machine learning in Python. Journal of Machine Learning Research, 12, 2825 - 2830. 182 Andrea Horbach et al. Rupp, A. A., Casabianca, J. M., Krüger, M., Keller, S. & Köller, O. (2019). Automated essay scoring at scale: A case study in Switzerland and Germany. ETS Research Report Series, 2019 (1), 1 - 23. https: / / doi.org/ 10.1002/ ets2.12249 Sanh, V., Debut, L., Chaumond, J. & Wolf, T. (2019). DistilBERT: A distilled version of BERT: Smaller, faster, cheaper and lighter. arXiv. https: / / arxiv.org/ abs/ 1910. 01108 Schaller, N.-J., Ding, Y., Horbach, A., Meyer, J. & Jansen, T. (2024). Fairness in automated essay scoring: A comparative analysis of algorithms on German learner essays from secondary education. In: Proceedings of the 19th Workshop on Innovative Use of NLP for Building Educational Applications (BEA 2024) (pp. 210 - 221). Association for Computational Linguistics. https: / / aclan thology.org/ 2024.bea-1.18/ Siemens, G. & Baker, R. S. J. d. (2012). Learning analytics and educational data mining: Towards communication and collaboration. In: Proceedings of the 2nd International Conference on Learning Analytics and Knowledge (pp. 252 - 254). ACM. https: / / doi.org/ 10.1145/ 2330 601.2330661 Summers, C. & Dinneen, M. J. (2021). Nondeterminism and instability in neural network optimization. In: Proceedings of the 38th International Conference on Machine Learning (pp. 9913 - 9922). PMLR. Taghipour, K. & Ng, H.T. (2016). A neural approach to automated essay scoring. In: Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing (pp. 1882 - 1891). Association for Computational Linguistics. https: / / doi.org/ 10.18653/ v1/ D16-1193 Uto, M. (2021). A review of deep-neural automated essay scoring models. Behaviormetrika, 48 (2), 459 - 484. https: / / doi.org/ 10.1007/ s41237-021-00142-y Uto, M., Xie, Y. & Ueno, M. (2020). Neural automated essay scoring incorporating handcrafted features. In: Proceedings of the 28th International Conference on Computational Linguistics (pp. 6077 - 6088). Association for Computational Linguistics. https: / / doi.org/ 10.186 53/ v1/ 2020.coling-main.535 Zesch, T., Horbach, A. & Zehner, F. (2023). To score or not to score: Factors influencing performance and feasibility of automatic content scoring of text responses. Educational Measurement: Issues and Practice, 42 (1), 44 - 58. https: / / doi.org/ 10.1111/ emip.12544 Andrea Horbach Daniel Mora Melanchthon Nils-Jonathan Schaller Leibniz-Institut für die Pädagogik der Naturwissenschaften und Mathematik Olshausenstr. 62 24118 Kiel E-Mail: horbach@leibniz-ipn.de mora@leibniz-ipn.de schaller@leibniz-ipn.de Stefan Keller Pädagogische Hochschule Zürich Lagerstr. 2 CH-8090 Zürich E-Mail: stefandaniel.keller@phzh.ch Jennifer Meyer Universität Wien Porzellangasse 4 A-1090 Wien E-Mail: jennifer.meyer@univie.ac.at Thorben Jansen Leibniz-Institut für die Pädagogik der Naturwissenschaften und Mathematik Olshausenstr. 62 24118 Kiel E-Mail: tjansen@leibniz-ipn.de Grundwissen ADHS fürs Studium Fragen zum Thema ADHS betreffen viele Studiengänge: Welche Symptome sind typisch? Wie diagnostiziert man ADHS? Welche Ursachen wurden erforscht - genetisch, neuropsychologisch, umweltbedingt? Wie entwickelt sich ADHS über die Lebensspanne? Neben diesen Themen werden insbesondere psychologische und medizinische Therapiemaßnahmen kritisch beleuchtet. Die ideale Seminarlektüre für Studierende in Psychologie, Pädagogik und Lehramt. 4., aktualisierte Auflage 2025. 194 Seiten. 16 Abb. 10 Tab. utb-M (978-3-8252-6547-2) kt a www.reinhardt-verlag.de On Suitability, Stability, and Synergy in Automated Essay Evaluation 183 Appendix The prompt structure used for the generative LLM models: # Essay Evaluation Prompt You are an expert evaluator tasked with scoring an essay based on a provided rubric. Your goal is to replicate the scoring decisions of a trained human grader. Carefully review the scoring rubric below. Your evaluation of the essay must be based solely on these criteria. --- ## PART 1: EVALUATION FRAMEWORK ### A. Scoring Rubric [SCORING RUBRIC INCLUDING SAMPLE ESSAYS INSERTED HERE] ### B. Guiding Principles **Specific Applications: ** * **Holistic Evaluation: ** Consider all aspects of the essay comprehensively rather than focusing on isolated elements * **Score Alignment: ** Ensure your score reflects the overall quality as described in the rubric criteria * **Consistency: ** Apply the same standards uniformly across all essays being evaluated Use a half-point score (e. g., 2.5, 3.5) when the essay clearly exceeds the expectations for the lower whole-number score, but falls short of fully meeting the criteria for the next higher score. Half scores are appropriate when the essay is clearly “in-between” levels - stronger than a 2 but not quite a 3, for example. --- ## PART 2: EVALUATION TASK ### 1. Read the Essay Carefully ### 2. Analyze According to the Rubric ### 3. Determine the Final Score Based on your analysis of the essay against the rubric criteria, assign a score from 0 to 5 (including half-point increments: 0.5, 1.5, 2.5, 3.5, 4.5) that best represents the overall quality of the work. ### **REQUIRED OUTPUT FORMAT** **Final Score: ** [Return only a single float between 0.0 and 5.0, including possible half-point scores: 0.0, 0.5, 1.0, 1.5, 2.0, 2.5, 3.0, 3.5, 4.0, 4.5, or 5.0.] --- ## PART 3: ESSAY TO EVALUATE [ESSAY TEXT INSERTED HERE]
