20 results for “grading”
CS papers onlyHybrid search: Keyword + semantic, ranked by combined score.ⓘ
Want pure semantic search? Try claim verification →
Hang Li, Fedor Filippov, Yuling Lin, Pengfei He +5 more
This paper investigates the vulnerability of LLM-based automatic grading systems to prompt injection (PI) attacks, demonstrating that current systems are highly susceptible to manipulation that can le…
Julien Benchek, Austin Bennett, Jasmin Kern, Ryan Stevens +7 more
APEX-Accounting benchmark is introduced to assess the capability of frontier models in performing accounting tasks. Claude-Fable-5 (Max) outperforms other models with 56.4% Mean Criteria@3.
This paper presents methods for ranking and unranking permutations avoiding a pattern of length three in lexicographic or colexicographic order.
Seojeong Park, Jiho Choi, Junyong Kang, Seonho Lee +2 more
The paper addresses Perceptual Judgment Bias in multimodal LLM judges by introducing a new dataset and a unified training framework that forces models to prioritize visual evidence over plausible text…
This paper provides translations between two graded coeffect systems, enabling the transfer of results and ideas between them.
The paper introduces an Item Response Theory (IRT)-based indicator that effectively identifies likely mislabeled items in existing LLM benchmarks, revealing systematic errors in labeling and model spe…
Haolin He, Renhe Sun, Zheqi Dai, Xingjian Du +15 more
This paper introduces Audio-Dependency Filtering (ADF) pipeline for Audio-Dependent Question Answering (ADQA) task in DCASE~2026, achieving top overall and sub-10B accuracy.
Junsoo Park, Youssef Medhat, Htet Phyo Wai, Ploy Thajchayapong +1 more
The paper proposes an interpretable, AI-driven decision layer that ranks course topics needing attention using multiple student and teacher signals, successfully identifying learning gaps before forma…
Valdemar Švábenský, Jan Vykopal, Sukrit Leelaluk, Pavel Čeleda +2 more
This paper compares two methods for assessing student teams in tabletop exercises using data from learning platforms and evaluates their validity and reliability.
This paper audits eight automatic scorers for attribution in LLM retrieval-augmented generation and finds that none of them transfer across datasets for generated-answer attribution.
The paper formalizes the concept of calibration for probabilistic label ranking, demonstrating that popular models are often poorly calibrated and that calibration captures a meaningful quality dimens…
Zhaoyang Jiang, Xuanqi Peng, Fei Teng, Zhizhong Fu +4 more
The paper demonstrates that while distilling large language models for medical QA can significantly improve final answer accuracy, this gain often comes at the cost of factual accuracy and detailed re…
This paper characterizes proper binary classification from positive-only samples, revealing a rich landscape that differs from standard PAC learning.
This paper extends gradient boosting to functions of vector inputs using a simple algorithm with histogram-based decision trees.
Swastik Roy, Rajkumar Pujari, Tharindu Kumarage, Charith Peris +4 more
PReMISE introduces a framework to audit and improve the quality of rubrics used to guide LLM judges, demonstrating that it can significantly increase judge accuracy and reduce the exploitability of re…
Wenhan Xiao, Ziwei Zhang, Chuanyue Yu, Xingcheng Fu +3 more
CRITIC-R1 introduces a structured critic framework that treats RAG critique as an explicit error diagnosis problem using reinforcement learning, significantly improving answer quality over strong RAG…
Tsvetomila Mihaylova, Jing Fan, Bita Akram, Narges Norouzi +3 more
This paper explores using Knowledge Components (KCs) as interpretable signals to understand assignment difficulty and student struggle in intro programming courses.
A new boosting algorithm that strong learns concept classes closed under O(log 1/γ)-XOR using O(log 1/ε) calls to a γ-advantage weak learner and additional samples, by connecting boosting with list-de…
The paper identifies a fundamental mismatch between standard pairwise ranking metrics (like AP and FPR-95) and the true assignment objective in multi-view object association, proposing a Sinkhorn-base…