20 results for “evaluation framework”
CS papers onlyHybrid search: Keyword + semantic, ranked by combined score.ⓘ
Want pure semantic search? Try claim verification →
Pengyu Zhu, Lijun Li, Yaxing Lyu, Qianxin Luo +7 more
The paper introduces a unified framework to fairly evaluate LLM agentic capabilities by standardizing diverse benchmarks and separating the effects of the LLM model from the surrounding framework and…
BADGER is a unified, production-grade evaluation framework that integrates text-to-SQL assessment with agentic behavior evaluation, significantly outperforming existing benchmarks on industry queries.
The paper applies an enhanced, practical version of the Argument-Based Validation (ABV) framework to assess eight Automated Essay Scoring (AES) models for French, demonstrating its value in understand…
LLM-FACETS introduces an open-source, privacy-preserving framework designed to enable non-technical domain experts and compliance officers to audit and evaluate the transparency and accountability of…
Ryan Burnell, Yumeya Yamamori, Orhan Firat, Kate Olszewska +9 more
The paper introduces a Cognitive Taxonomy and a rigorous evaluation protocol to provide an objective, multi-faceted framework for measuring system capabilities and tracking progress toward Artificial…
Aakash Pant, Kavya Shah, Apoorv Agnihotri, Sneha Nikam +2 more
The paper critiques current AI benchmarking practices for low-resource settings, arguing that evaluation must shift focus from isolated model performance to the holistic performance of the deployed sy…
The paper proposes a TEE-based architecture that enables external, auditable verification of AI-assisted grant evaluations without exposing the proprietary model, scoring logic, or intermediate reason…
This paper studies AI development frameworks for software engineering and proposes a six-dimension process taxonomy.
Aaron Chan, Tengfei Li, Tianyi Xiao, Angela Chen +2 more
The paper introduces LATTICE, a novel benchmark for evaluating how well crypto agents assist user decision-making, finding that different agents excel in different specific areas rather than having a…
This paper proposes a comprehensive, risk-based auditing framework designed to help internal and external auditors assess the cybersecurity risks posed by diverse IoT devices within corporate and indu…
The BEAMS initiative establishes comprehensive benchmarks and evaluates AI tools for modeling and simulation, finding that current AI tools excel at qualitative discussion tasks but struggle with comp…
Weiyi Chen, Shuaixiong Wang, Ziyun Gao, Kaichun Hu +4 more
The paper introduces TravelEval, a comprehensive, six-dimensional benchmarking framework that evaluates LLM-powered travel plans using realistic spatio-temporal simulation, revealing that current LLMs…
This paper identifies and validates eight core UX principles for human-AI agent interaction in the workplace using a multi-method approach.
Junjie Chen, Yuxi Dong, Haitao Li, Weihang Su +4 more
The paper introduces LongJudgeBench, a new benchmark designed to evaluate the reliability of LLM judges specifically for complex, long-form output evaluation, revealing significant instability gaps in…
The paper proposes a multi-axis evaluation framework for structured audio descriptions using a controlled perturbation testing protocol.
The paper introduces VIABLE, the first benchmark for evaluating Vision-Language Models (VLMs) as judges for Visually Impaired Assistance (VIA), finding that current models are largely unreliable and p…
A black-box evaluation framework is presented to assess Large Language Models' ability to generate Design Structure Matrices from technical documentation, using structural, classification, and stabili…
The paper describes the 2026 MOASEI Competition, which evaluates multi-agent decision-making under open-system conditions in wildfire fighting, cybersecurity, and ride-sharing domains.
Nhat Minh Le, Yisen Xu, Zhijie Wang, Tse-Hsun +1 more
This paper evaluates the performance of large language models on popular benchmarks and finds that only a small percentage of the performant implementations are significantly faster than canonical sol…