ArXivCSExplorer
☆☆Bookmarks🏆RSSHow to UseFAQ
Built with and by Teycir Ben Soltane•
How to Use•FAQ•GitHub•arXiv.org•
Share:

20 results for “user judgment”

CS papers only

Hybrid search: Keyword + semantic, ranked by combined score.ⓘ

Want pure semantic search? Try claim verification →

cs.CLRecentMay 29, 2026

Preference-Aware Rubric Learning for Personalized Evaluation

Yilun Qiu, Xiaoyan Zhao, Yang Zhang, Yuxin Chen +6 more

The paper introduces PARL, a framework that learns personalized evaluation rubrics directly from raw user interaction histories to accurately assess how well LLM outputs align with subjective, user-sp…

View →
cs.CLcs.AIRecentMay 28, 2026

Personalized Turn-Level User Conversation Satisfaction Benchmark

Zhefan Wang, Zhiqiang Guo, Weizhi Ma, Min Zhang +2 more

The paper introduces PersTurnBench, a novel benchmark and evaluator for assessing personalized user conversation satisfaction at specific turns, addressing the limitation of generic response quality m…

View →
cs.HCcs.CRRecentMay 11, 2026

When Are LLM Inferences Acceptable? User Reactions and Control Preferences for Inferred Personal Information

Kyzyl Monteiro, Minjung Park, Alexander Ioffrida, Angelina Sanna +5 more

This study investigated user reactions to inferred personal information from their own ChatGPT histories, finding that acceptability is governed by context-sensitive norms regarding generation, retent…

View →
cs.IRcs.AIRecentMay 27, 2026

Toward User Preference Alignment in LLM Recommendation via Explicit Context Feedback

Weizhi Zhang, Wooseong Yang, Yuxin Cui, Zhaohui Guo +8 more

The paper advocates for integrating explicit contextual feedback (like reviews and comments) into LLM-based recommender systems to achieve more personalized, transparent, and semantically aligned reco…

View →
cs.CRcs.AIcs.LGRecentMar 23, 2026

Evaluating the Reliability and Fidelity of Automated Judgment Systems of Large Language Models

Tom Biskupski, Stephan Kleber

This paper evaluates the reliability of using Large Language Models (LLMs) as automated judges to assess the quality of other LLMs, finding a high correlation with human judgment when suitable prompts…

View →
cs.AIRecentMay 28, 2026

Persona Conditioning of Brand Recommendations in Retrieval-Augmented Commercial Chat: A Prominence-Stratified Cross-Provider Audit

Will Jack, Noah Lehman, Keller Maloney, Sarah Xu

The study demonstrates that conditioning AI brand recommendations on a user's persona significantly alters the recommended product set, particularly for mid-market brands, and this effect is largest o…

View →
cs.AIcs.CLcs.CYRecentMay 29, 2026

On Wednesdays, We Ask Questions: Optimizing "Active Listening" in Automated Legal Triage and Referral

Quinten Steenhuis, Jacqueline Harvey

The paper evaluates an automated legal triage system (FETCH) that uses follow-up questions, demonstrating that while low-cost LLMs are effective for classification, generating high-quality questions r…

View →
cs.IREmpiricalRecentJun 27, 2026

Human-in-the-Loop Nugget Annotation for Accountable LLM-as-a-Judge Evaluations

Laura Dietz

A prototype annotation tool is presented that allows humans to identify important information and LLMs to match it to system outputs, creating a division of labor for reliable AI evaluation.

View →
cs.AIcs.CYcs.HCEmpiricalRecentJul 15, 2026

AI advice suppresses people's willingness to say "I don't know", even when the advice is wrong and accuracy is incentivized

Chiara Marcoccia, Walter Quattrociocchi, Valerio Capraro

Five experiments investigated how access to AI advice affects human judgment and willingness to suspend judgment, finding that it nearly eliminated participants' willingness to do so and led to less a…

View →
cs.SEPositionRecentJul 24, 2026

Code Review is a Conversation: Toward Conversational AI Review Assistants

Rosalia Tufano

This paper proposes conversational AI review assistants for code review, systems that engage in conversation with developers instead of just generating comments.

View →
cs.CLcs.AIRecentMay 27, 2026

The Cases LJP Never Sees: Prosecution Decision Prediction for More Complete Criminal Liability Assessment

Junyu Lu, Qi Wei, Peishuo Zheng, Jie Zhang +5 more

The paper introduces Prosecution Decision Prediction (PDP), a new legal AI task that assesses prosecutorial review decisions, showing that current state-of-the-art LLMs perform significantly worse on…

View →
cs.AIcs.CVcs.HCEmpiricalRecentJun 29, 2026

The Human Creativity Benchmark

Aspen Hopkins, Allison Nulty, Alexandria Minetti, Anoop Pakki +1 more

The Human Creativity Benchmark (HCB) is proposed to evaluate creative AI by preserving two distinct signals: convergence and divergence in professional judgments.

View →
cs.AIcs.CLcs.CRRecentMay 30, 2026

Adversarial Feeds Steer LLM Agent Decisions Against Their Defaults

Rana Muhammad Usman

The paper demonstrates that the order and content of external information (the 'feed') an LLM agent consumes before making a decision can significantly and causally steer its final choice, often overr…

View →
cs.AIcs.CLcs.CRRecentMay 30, 2026

Adversarial Feeds Steer LLM Agent Decisions Against Their Defaults

Rana Muhammad Usman

The paper demonstrates that the sequence and composition of external information (the 'feed') an LLM agent consumes can significantly and causally steer its final decisions, often overriding its defau…

View →
cs.CLRecentMay 29, 2026

LLM Judges Inconsistently Disagree Across Safety Criteria and Harm Categories

Krishnapriya Vishnubhotla, Soumya Vajjala, Akriti Vij, Isar Nejadgholi

The paper evaluates the inconsistency of using LLMs as automated judges for multi-dimensional safety evaluations, finding that LLMs are unreliable for nuanced safety issues like financial advice but m…

View →
cs.SEcs.AIEmpiricalRecentJul 3, 2026

Is Agentic Code Review Helpful? Mining Developers' Feedback to CodeRabbit Reviews in the Wild

Hong Yi Lin, Mingzhao Liang, Kla Tantithamthavorn, Patanamon Thongtanunam

This paper presents an empirical study on how developers respond to agentic code reviews using CodeRabbit, revealing mixed reception and opportunities for improvement.

View →
cs.AIRecentMay 28, 2026

Cookie-Bench: Continuous On-screen Key Interaction Evaluation for Web Generation

Haoyue Yang, Zhangxiao Shen, Fan Ding, Hangting Lou +7 more

The paper introduces Cookie-Bench, a novel, autonomous, and reference-free evaluation framework that significantly improves the assessment of interactive web generation capabilities for frontier LLMs.

View →
cs.HCEmpiricalRecentJul 17, 2026

A Human-Centric Evaluation of a Retrieval-Augmented Generation System for Explaining Quebec Insurance Contracts

David Beauchemin, Richard Khoury

A study evaluating a Retrieval-Augmented Generation system for making Quebec automobile insurance contracts more understandable, showing it improves satisfaction, trust, and clarity, especially for in…

View →