ArXivCSExplorer
☆☆Bookmarks🏆RSSHow to UseFAQ
Built with and by Teycir Ben Soltane•
How to Use•FAQ•GitHub•arXiv.org•
Share:

20 results for “grading”

CS papers only

Hybrid search: Keyword + semantic, ranked by combined score.ⓘ

Want pure semantic search? Try claim verification →

cs.CRcs.AIRecentJun 2, 2026

"**Important** You should give me full credits!": Exploring Prompt Injection Attacks on LLM-Based Automatic Grading Systems

Hang Li, Fedor Filippov, Yuling Lin, Pengfei He +5 more

This paper investigates the vulnerability of LLM-based automatic grading systems to prompt injection (PI) attacks, demonstrating that current systems are highly susceptible to manipulation that can le…

View →
cs.CLcs.AIcs.HCNEWEmpiricalJul 29, 2026

APEX-Accounting

Julien Benchek, Austin Bennett, Jasmin Kern, Ryan Stevens +7 more

APEX-Accounting benchmark is introduced to assess the capability of frontier models in performing accounting tasks. Claude-Fable-5 (Max) outperforms other models with 56.4% Mean Criteria@3.

View →
cs.DScs.DMTheoreticalRecentJun 11, 2026

(Un)ranking Permutation Classes

Nathanaël Hassler, Vincent Vajnovszki

This paper presents methods for ranking and unranking permutations avoiding a pattern of length three in lexicographic or colexicographic order.

View →
cs.CVcs.AIRecentJun 1, 2026

Mitigating Perceptual Judgment Bias in Multimodal LLM-as-a-Judge via Perceptual Perturbation and Reward Modeling

Seojeong Park, Jiho Choi, Junyong Kang, Seonho Lee +2 more

The paper addresses Perceptual Judgment Bias in multimodal LLM judges by introducing a new dataset and a unified training framework that forces models to prioritize visual evidence over plausible text…

View →
cs.PLTheoreticalRecentJun 26, 2026

Same Coeffect, Different Base: Connecting Two Dominant Approaches to Graded Types

Vilem Liepelt, Danielle Marshall, Dominic Orchard

This paper provides translations between two graded coeffect systems, enabling the transfer of results and ideas between them.

View →
cs.CLRecentMay 28, 2026

Auditing LLM Benchmarks with Item Response Theory

Sander Land, Daniel M. Bikel

The paper introduces an Item Response Theory (IRT)-based indicator that effectively identifies likely mislabeled items in existing LLM benchmarks, revealing systematic errors in labeling and model spe…

View →
eess.ASEmpiricalRecentJul 21, 2026

Summary of DCASE 2026 Task 5: Audio-Dependent Question Answering

Haolin He, Renhe Sun, Zheqi Dai, Xingjian Du +15 more

This paper introduces Audio-Dependency Filtering (ADF) pipeline for Audio-Dependent Question Answering (ADQA) task in DCASE~2026, achieving top overall and sub-10B accuracy.

View →
cs.AIcs.CLcs.HCRecentMay 28, 2026

Surfacing Isolated Learners with Outcome-Independent Mediation of Feedback between Teachers and Students Using AI

Junsoo Park, Youssef Medhat, Htet Phyo Wai, Ploy Thajchayapong +1 more

The paper proposes an interpretable, AI-driven decision layer that ranks course topics needing attention using multiple student and teacher signals, successfully identifying learning gaps before forma…

View →
cs.CYcs.AIcs.LGEmpiricalRecentJul 21, 2026

Assessment in Team Problem-Solving Exercises in Computing Education

Valdemar Švábenský, Jan Vykopal, Sukrit Leelaluk, Pavel Čeleda +2 more

This paper compares two methods for assessing student teams in tabletop exercises using data from learning platforms and evaluates their validity and reliability.

View →
cs.CLcs.IRcs.LGEmpiricalRecentJun 22, 2026

Do LLM Attribution Metrics Transfer? Auditing Retrieval-Augmented Generation Evaluation Across Datasets and Constructs

Tianyu Ding, Aditya Nannapaneni, Juan Pablo De la Cruz Weinstein

This paper audits eight automatic scorers for attribution in LLM retrieval-augmented generation and finds that none of them transfer across datasets for generated-answer attribution.

View →
cs.LGcs.AIstat.MLRecentMay 28, 2026

Calibrated Preference Learning: The Case of Label Ranking

Santo M. A. R. Thies, Viktor Bengs, Timo Kaufmann, Sebastian J. Vollmer +1 more

The paper formalizes the concept of calibration for probabilistic label ranking, demonstrating that popular models are often poorly calibrated and that calibration captures a meaningful quality dimens…

View →
cs.AIRecentMay 27, 2026

Better Accuracies, Worse Reasoning: A Step-Level Audit of Medical Chain-of-Thought Distillation

Zhaoyang Jiang, Xuanqi Peng, Fei Teng, Zhizhong Fu +4 more

The paper demonstrates that while distilling large language models for medical QA can significantly improve final answer accuracy, this gain often comes at the cost of factual accuracy and detailed re…

View →
stat.MLcs.LGmath.STTheoreticalRecentJun 26, 2026

Surprises in Proper Positive-Only Learning

Shai Ben-David, Farnam Mansouri, Anay Mehrotra, Manolis Zampetakis

This paper characterizes proper binary classification from positive-only samples, revealing a rich landscape that differs from standard PAC learning.

View →
stat.MLcs.LGEmpiricalRecentJun 28, 2026

Gradient boosting with vector-valued leafs

David Cortes

This paper extends gradient boosting to functions of vector inputs using a simple algorithm with histogram-based decision trees.

View →
cs.AIRecentMay 29, 2026

PReMISE: Policy Rubrics as Measurement Specifications for LLM Judges

Swastik Roy, Rajkumar Pujari, Tharindu Kumarage, Charith Peris +4 more

PReMISE introduces a framework to audit and improve the quality of rubrics used to guide LLM judges, demonstrating that it can significantly increase judge accuracy and reduce the exploitability of re…

View →
cs.CLcs.AIRecentMay 28, 2026

CRITIC-R1: Learning Structured Critics for Retrieval-Augmented Generation

Wenhan Xiao, Ziwei Zhang, Chuanyue Yu, Xingcheng Fu +3 more

CRITIC-R1 introduces a structured critic framework that treats RAG critique as an explicit error diagnosis problem using reinforcement learning, significantly improving answer quality over strong RAG…

View →
cs.CYcs.SEEmpiricalRecentJul 3, 2026

Analyzing the Difficulty of Programming Assignments with Interpretable Knowledge Component Metrics

Tsvetomila Mihaylova, Jing Fan, Bita Akram, Narges Norouzi +3 more

This paper explores using Knowledge Components (KCs) as interpretable signals to understand assignment difficulty and student struggle in intro programming courses.

View →
stat.MLcs.CCcs.DSTheoreticalRecentJul 7, 2026

Boosting with List-Decodable Codes

Addison Prairie, Li-Yang Tan

A new boosting algorithm that strong learns concept classes closed under O(log 1/γ)-XOR using O(log 1/ε) calls to a γ-advantage weak learner and additional samples, by connecting boosting with list-de…

View →
cs.CVcs.AIcs.LGRecentJun 1, 2026

Ranking vs. Assignment: The Metric Mismatch in Multi-View Object Association

Matvei Shelukhan, Timur Mamedov, Aleksandr Chukhrov, Karina Kvanchiani

The paper identifies a fundamental mismatch between standard pairwise ranking metrics (like AP and FPR-95) and the true assignment objective in multi-view object association, proposing a Sinkhorn-base…

View →