ArXivCSExplorer
☆☆Bookmarks🏆RSSHow to UseFAQ
Built with and by Teycir Ben Soltane•
How to Use•FAQ•GitHub•arXiv.org•
Share:

20 results for “Discrepancy”

CS papers only

Hybrid search: Keyword + semantic, ranked by combined score.ⓘ

Want pure semantic search? Try claim verification →

cs.ITcs.CCcs.CRTheoreticalRecentJun 23, 2026

Discrepancy for Random Linear Codes

Dean Doron, Tal Leonov, Jonathan Mosheiff, Henrique Navas +2 more

This paper proves that random linear codes have nearly optimal discrepancy properties in various regimes, extending classical results and enabling new applications.

View →
cs.LGcs.AIRecentMay 29, 2026

Inconsistency-Aware Minimization: Improving Generalization with Unlabeled Data

Hee-Sung Kim, Hyeonseong Kim, Sungyoon Lee

The paper introduces Inconsistency-Aware Minimization (IAM), a novel training objective that uses a label-free measure called local inconsistency to improve model generalization, particularly in semi-…

View →
cs.LGcs.AIcs.CRRecentApr 28, 2026

Conditional misalignment: common interventions can hide emergent misalignment behind contextual triggers

Jan Dubiński, Jan Betley, Anna Sztyber-Betley, Daniel Tan +1 more

The paper introduces the concept of 'conditional misalignment,' demonstrating that common interventions designed to reduce emergent misalignment can fail by only masking misaligned behavior until the…

View →
cs.AIRecentMay 27, 2026

Better Accuracies, Worse Reasoning: A Step-Level Audit of Medical Chain-of-Thought Distillation

Zhaoyang Jiang, Xuanqi Peng, Fei Teng, Zhizhong Fu +4 more

The paper demonstrates that while distilling large language models for medical QA can significantly improve final answer accuracy, this gain often comes at the cost of factual accuracy and detailed re…

View →
cs.CLcs.AIEmpiricalRecentJun 12, 2026

Fodor and Pylyshyn's Systematicity Challenge Still Stands

Michael Goodale, Salvador Mascarenhas

This paper challenges the claim that neural networks have met the challenge of systematicity in language and thought as proposed by Fodor and Pylyshyn, demonstrating limitations in a recent neural net…

View →
stat.MLcs.LGmath.STTheoreticalRecentJun 26, 2026

Surprises in Proper Positive-Only Learning

Shai Ben-David, Farnam Mansouri, Anay Mehrotra, Manolis Zampetakis

This paper characterizes proper binary classification from positive-only samples, revealing a rich landscape that differs from standard PAC learning.

View →
cs.CVcs.AIRecentMay 28, 2026

Rethinking FID Through the Geometry of the Reference Dataset

Yunghee Lee, Byeonghyun Pak

The paper argues that the standard FID metric is unreliable because its performance depends significantly on the geometric structure and density of the reference dataset, not just the sample quality.

View →
cs.CLcs.AIcs.CYRecentMay 31, 2026

Implicit Geographic Inference in LLM Medical Triage: Language-Driven Disparities in Emergency Recommendations

Qi Han Wong

The study demonstrates that LLMs exhibit significant, language-driven disparities in medical triage recommendations, recommending emergency care more frequently for English and Arabic prompts, even wh…

View →
cs.AIcs.CLcs.LGRecentMay 31, 2026

An Enigma of Artificial Reason: Investigating the Production-Evaluation Gap in Large Reasoning Models

Mingzhong Sun, Teresa Yeo, Armando Solar-Lezama, Tan Zhi-Xuan

This paper investigates the production-evaluation gap in Large Reasoning Models (LRMs), finding that while LRMs excel at generating solutions, they struggle significantly to evaluate flawed reasoning,…

View →
cs.LGEmpiricalRecentJun 30, 2026

Calibration, Not Compilation: Detecting and Repairing Misspecified Probabilistic Programs Written by Language Models

Jian Xu, Delu Zeng, John Paisley, Qibin Zhao

This paper evaluates the effectiveness of Bayesian workflow for verifying statistical correctness of probabilistic programs written by language models, and compares it to unit tests and no feedback.

View →
cs.CLcs.AIcs.IRRecentMay 27, 2026

Same Question, Different Source, Different Answer: Auditing Source-Dependence in Medical Multi-Source RAG

Yubo Li, Rema Padman, Ramayya Krishnan

This paper introduces a framework to audit source-dependence in multi-source RAG systems, demonstrating that disagreement across institutional sources is a common and critical failure mode that curren…

View →
cs.AIEmpiricalRecentJul 16, 2026

Can We Trust Item Response Theory for AI Evaluation?

Han Jiang, Sunbeom Kwon, Jinwen Luo, Ziang Xiao +1 more

This paper evaluates the reliability of using item response theory (IRT) models for AI benchmarking, comparing four estimation tools under various simulation conditions.

View →
cs.CLRecentJun 2, 2026

Language Models Compare Quantities Using Number-specific and Unit-specific Heuristics

Mutsumi Sasaki, Go kamoda, Ryosuke Takahashi, Kosuke Sato +3 more

This study investigates how language models compare quantities with units, finding that they rely on a combination of separate heuristics for numerals and units rather than performing a precise, share…

View →
cs.LGcs.IREmpiricalRecentJun 10, 2026

DeMix: Debugging Training Data with Mixed Data Error Types by Investigating Influence Vectors

Jiale Deng, Yanyan Shen, Xiaogang Shi, Chai Junjun

This paper proposes DeMix, a novel framework for simultaneously diagnosing erroneous samples and their error types in machine learning models.

View →
cs.AIcs.CLRecentMay 27, 2026

The Importance of Being Statistically Earnest: A Critical Re-evaluation of GSM-Symbolic

Dominika Agnieszka Długosz, Arlindo Oliveira, Natalia Díaz-Rodríguez

The paper challenges the conclusion that LLMs lack reasoning by demonstrating that reported performance drops on GSM-Symbolic are often statistically weak and partially attributable to dataset biases,…

View →
cs.LGcs.CYRecentJun 1, 2026

Model Multiplicity and Predictive Arbitrariness in Recidivism Risk Assessment

Ashwin Singh, Carlos Castillo

The paper investigates predictive multiplicity and arbitrariness in recidivism risk assessment, finding that similarly accurate models often exhibit high predictive agreement, and proposes a simple po…

View →