ArXivCSExplorer
☆☆Bookmarks🏆RSSHow to UseFAQ
Built with and by Teycir Ben Soltane•
How to Use•FAQ•GitHub•arXiv.org•
Share:

20 results for “evaluation framework”

CS papers only

Hybrid search: Keyword + semantic, ranked by combined score.ⓘ

Want pure semantic search? Try claim verification →

cs.AIRecentMay 27, 2026

A Unified Framework for the Evaluation of LLM Agentic Capabilities

Pengyu Zhu, Lijun Li, Yaxing Lyu, Qianxin Luo +7 more

The paper introduces a unified framework to fairly evaluate LLM agentic capabilities by standardizing diverse benchmarks and separating the effects of the LLM model from the surrounding framework and…

View →
cs.AIRecentJun 1, 2026

BADGER: Bridging Agentic and Deterministic Evaluation for Generative Enterprise Reasoning

Shannon Serrao, Soumitra Chatterjee, Dorina Strori, Abhishek Sharma +1 more

BADGER is a unified, production-grade evaluation framework that integrates text-to-SQL assessment with agentic behavior evaluation, significantly outperforming existing benchmarks on industry queries.

View →
cs.CLRecentJun 1, 2026

Automated Essay Scoring and Language Certification: Assessing Generalizability, Agreement and Validity for French

Rodrigo Wilkens, Rémi Cardon, Vincent Folny, Thomas François

The paper applies an enhanced, practical version of the Argument-Based Validation (ABV) framework to assess eight Automated Essay Scoring (AES) models for French, demonstrating its value in understand…

View →
cs.AIRecentMay 29, 2026

LLM-FACETS: A Privacy-Preserving Framework for Evaluating LLM Transparency and Accountability

Tom Lucas, Alessio Buscemi, Alfredo Capozucca, German Castignani +1 more

LLM-FACETS introduces an open-source, privacy-preserving framework designed to enable non-technical domain experts and compliance officers to audit and evaluate the transparency and accountability of…

View →
cs.AIRecentMay 27, 2026

Measuring Progress Toward AGI: A Cognitive Framework

Ryan Burnell, Yumeya Yamamori, Orhan Firat, Kate Olszewska +9 more

The paper introduces a Cognitive Taxonomy and a rigorous evaluation protocol to provide an objective, multi-faceted framework for measuring system capabilities and tracking progress toward Artificial…

View →
cs.AIRecentMay 27, 2026

Benchmarking AI for low-resource contexts: Thinking beyond leaderboards

Aakash Pant, Kavya Shah, Apoorv Agnihotri, Sneha Nikam +2 more

The paper critiques current AI benchmarking practices for low-resource settings, arguing that evaluation must shift focus from isolated model performance to the holistic performance of the deployed sy…

View →
cs.CRcs.AIcs.CYRecentApr 28, 2026

Making AI-Assisted Grant Evaluation Auditable without Exposing the Model

Kemal Bicakci

The paper proposes a TEE-based architecture that enables external, auditable verification of AI-assisted grant evaluations without exposing the proprietary model, scoring logic, or intermediate reason…

View →
cs.SEcs.AIRecentJun 3, 2026

From Prompt to Process: a Process Taxonomy and Comparative Assessment of Frameworks Supporting AI Software Development Agents

Sanderson Oliveira de Macedo

This paper studies AI development frameworks for software engineering and proposes a six-dimension process taxonomy.

View →
cs.CRcs.AIcs.CLRecentApr 29, 2026

LATTICE: Evaluating Decision Support Utility of Crypto Agents

Aaron Chan, Tengfei Li, Tianyi Xiao, Angela Chen +2 more

The paper introduces LATTICE, a novel benchmark for evaluating how well crypto agents assist user decision-making, finding that different agents excel in different specific areas rather than having a…

View →
cs.CRRecentMar 23, 2026

Framework for Risk-Based IoT Cybersecurity Audit Engagements

Danielle Hanson, Jeremy Straub

This paper proposes a comprehensive, risk-based auditing framework designed to help internal and external auditors assess the cybersecurity risks posed by diverse IoT devices within corporate and indu…

View →
cs.AIRecentMay 27, 2026

BEAMS: Benchmarking and Evaluating AI for Modeling and Simulation

Sara Metcalf, William Schoenberg

The BEAMS initiative establishes comprehensive benchmarks and evaluates AI tools for modeling and simulation, finding that current AI tools excel at qualitative discussion tasks but struggle with comp…

View →
cs.AIRecentMay 31, 2026

TravelEval: A Comprehensive Benchmarking Framework for Evaluating LLM-Powered Travel Planning Agents

Weiyi Chen, Shuaixiong Wang, Ziyun Gao, Kaichun Hu +4 more

The paper introduces TravelEval, a comprehensive, six-dimensional benchmarking framework that evaluates LLM-powered travel plans using realistic spatio-temporal simulation, revealing that current LLMs…

View →
cs.HCcs.AIEmpiricalRecentJul 22, 2026

A Framework of User Experience Principles for Human-AI Agent Interaction in the Workplace

Kathrin Paimann, Elizangela Valarini, Sebastian Juhl

This paper identifies and validates eight core UX principles for human-AI agent interaction in the workplace using a multi-method approach.

View →
cs.CLRecentJun 1, 2026

Benchmarking LLM-as-a-Judge for Long-Form Output Evaluation

Junjie Chen, Yuxi Dong, Haitao Li, Weihang Su +4 more

The paper introduces LongJudgeBench, a new benchmark designed to evaluate the reliability of LLM judges specifically for complex, long-form output evaluation, revealing significant instability gaps in…

View →
cs.CLcs.SDEmpiricalRecentJul 23, 2026

An Evaluation Framework for Structured Audio Captions Validated by Controlled Perturbations

Liang-Yuan Wu, Sripathi Sridhar, Mark Cartwright, Magdalena Fuentes

The paper proposes a multi-axis evaluation framework for structured audio descriptions using a controlled perturbation testing protocol.

View →
cs.CLcs.CVRecentMay 29, 2026

A Visually Impaired Assistance Benchmark for VLM-as-a-Judge Evaluation

Yi Zhao, Siqi Wang, Zhe Hu, Yushi Li +1 more

The paper introduces VIABLE, the first benchmark for evaluating Vision-Language Models (VLMs) as judges for Visually Impaired Assistance (VIA), finding that current models are largely unreliable and p…

View →
cs.AIcs.ARcs.CEEmpiricalRecentJul 7, 2026

Auto-DSM Under the Lens: A Black-Box Evaluation Framework for LLM-Based DSM Generation

Niels Potters, Theo Hofman

A black-box evaluation framework is presented to assess Large Language Models' ability to generate Design Structure Matrices from technical documentation, using structural, classification, and stabili…

View →
cs.MAEmpiricalRecentJul 3, 2026

Second MOASEI Competition at AAMAS'2026: A Technical Report

Ceferino Patino, Tyler J. Billings, Alireza Saleh Abadi, Daniel Redder +3 more

The paper describes the 2026 MOASEI Competition, which evaluates multi-agent decision-making under open-system conditions in wildfire fighting, cybersecurity, and ride-sharing domains.

View →
cs.SEEmpiricalRecentJul 8, 2026

Rethinking Code Performance Benchmarks for LLMs

Nhat Minh Le, Yisen Xu, Zhijie Wang, Tse-Hsun +1 more

This paper evaluates the performance of large language models on popular benchmarks and finds that only a small percentage of the performant implementations are significantly faster than canonical sol…

View →