ArXivCSExplorer
☆☆Bookmarks🏆RSSHow to UseFAQ
Built with and by Teycir Ben Soltane•
How to Use•FAQ•GitHub•arXiv.org•
Share:

20 results for “faithful calibration”

CS papers only

Hybrid search: Keyword + semantic, ranked by combined score.ⓘ

Want pure semantic search? Try claim verification →

cs.CLcs.AIRecentJun 2, 2026

Quantifying Faithful Confidence Expression in Large Reasoning Models

Areeb Gani, Asal Meskin, Gabrielle Kaili-May Liu, Arman Cohan

The paper introduces a novel framework to quantify faithful confidence expression (FC) in Large Reasoning Models (LRMs), finding that FC remains a significant and challenging reliability target for th…

View →
cs.LGcs.AIstat.MLRecentMay 28, 2026

CalArena: A Large-Scale Post-Hoc Calibration Benchmark

Eugène Berta, David Holzmüller, Francis Bach, Michael I. Jordan

The paper introduces CalArena, a large-scale, standardized benchmark covering nearly 2000 experiments to comprehensively evaluate post-hoc calibration methods, finding that smooth calibration function…

View →
cs.CLRecentMay 28, 2026

Counterfactual Graph for Multi-Agent LLM Calibration

Jiatan Huang, Mingchen Li, Ziming Li, Sunjae Kwon +2 more

The paper proposes CAGE-CAL, a counterfactual graph calibration framework, to accurately assess the reliability and detect over-confidence in multi-agent LLM systems after agents communicate.

View →
cs.LGEmpiricalRecentJun 30, 2026

Calibration, Not Compilation: Detecting and Repairing Misspecified Probabilistic Programs Written by Language Models

Jian Xu, Delu Zeng, John Paisley, Qibin Zhao

This paper evaluates the effectiveness of Bayesian workflow for verifying statistical correctness of probabilistic programs written by language models, and compares it to unit tests and no feedback.

View →
cs.AIcs.LGEmpiricalRecentJun 18, 2026

Toward Calibrated Mixture-of-Experts Under Distribution Shift

Gina Wong, Drew Prinster, Suchi Saria, Rama Chellappa +1 more

This paper studies the behavior of mixture-of-experts (MoE) models under distribution shift and proposes an adversarial reweighting method to improve their calibration.

View →
cs.LGcs.AIstat.MLRecentMay 28, 2026

Calibrated Preference Learning: The Case of Label Ranking

Santo M. A. R. Thies, Viktor Bengs, Timo Kaufmann, Sebastian J. Vollmer +1 more

The paper formalizes the concept of calibration for probabilistic label ranking, demonstrating that popular models are often poorly calibrated and that calibration captures a meaningful quality dimens…

View →
cs.CVEmpiricalRecentJun 18, 2026

Reliability-Aware Prototype Calibration for Frozen Pose-Flow Video Anomaly Detection

Ning Dong, Yingna Su, Xin Dong, Ziyun Jiao +2 more

This paper introduces Reliability-Aware Prototype Calibration (RPC), a post-hoc score calibration method for pose-flow video anomaly detectors in a frozen-detector setting.

View →
cs.LGRecentJun 1, 2026

Low-Pass Flow Matching

Francesco M. Ruscio, T. Konstantin Rusch

Low-Pass Flow Matching introduces a spectral bias into the flow matching process, allowing it to better model natural data by transitioning from a standard source spectrum to a frequency-decaying bias…

View →
cs.AIRecentMay 27, 2026

Relevant Is Not Warranted: Evidence-Force Calibration for Cited RAG

Pin Qian, Su Wang, Xiaoyuan Wang, Yihang Chen +6 more

The paper introduces FORCEBENCH, a new stress test designed to evaluate whether cited sources genuinely warrant the strength of a claim, revealing that standard citation evaluation methods often fail…

View →
cs.CVcs.AIcs.LGEmpiricalRecentJul 6, 2026

From Fixed to Free Cameras: Calibration-Free View-Robust Vision-Language-Action Model

Wenhao Li, Xueying Jiang, Quanhao Qian, Deli Zhao +3 more

This paper introduces Camera-Centric VLA, a new model for Vision-Language-Action policies that predicts camera-centric actions and hand-eye matrix, allowing the policy to figure out camera geometry on…

View →
eess.SPTheoreticalRecentJun 12, 2026

On Optimal Strategies for Joint Reciprocity Calibration in Distributed MIMO

Kohei Ueda, Anubhab Chowdhury, Koji Ishibashi, Erik G. Larsson

This paper compares the downlink spectral efficiency of multi-user large antenna systems with global and local calibration approaches, and demonstrates that global calibration outperforms local calibr…

View →
cs.AIcs.CLcs.LGEmpiricalRecentJul 2, 2026

Online Safety Monitoring for LLMs

Mona Schirmer, Metod Jazbec, Alexander Timans, Christian Naesseth +2 more

This paper proposes a simple real-time monitor for LLMs that turns an external verifier signal into an alarm decision by thresholding, showing competitiveness with advanced monitors in mathematical re…

View →
eess.SPEmpiricalRecentJul 27, 2026

On-Site Beam Calibration for RIS-Aided Positioning Systems

Mengting Li, Hui Chen, Sigurd S. Petersen, Alireza Pourafzal +4 more

This paper proposes a framework for on-site calibration of Reconfigurable Intelligent Surface (RIS) beam response models to reduce positioning error floor.

View →
cs.CRcs.LGRecentMay 23, 2026

CALIBURN: A Regime-Sensitivity Study of Operationally Calibrated Streaming Intrusion Detection

Michel A. Youssef

CALIBURN introduces a novel, five-component streaming pipeline for intrusion detection that allows operators to specify alerting behavior using cost and budget constraints, achieving state-of-the-art…

View →
cs.CVcs.AIcs.LGRecentJun 1, 2026

Ranking vs. Assignment: The Metric Mismatch in Multi-View Object Association

Matvei Shelukhan, Timur Mamedov, Aleksandr Chukhrov, Karina Kvanchiani

The paper identifies a fundamental mismatch between standard pairwise ranking metrics (like AP and FPR-95) and the true assignment objective in multi-view object association, proposing a Sinkhorn-base…

View →
cs.AIcs.DBcs.IRRecentMay 29, 2026

Vector Linking via Cross-Model Local Isometric Consistency

Ziying Chen, Yang Cao, He Sun, Beining Yang +1 more

The paper proposes a novel geometric embedding hashing method to recover object correspondences (vector links) between two embedding clouds generated by different black-box encoders using only a small…

View →
eess.SPcs.ITTheoreticalRecentJun 12, 2026

Repeater-Assisted Massive MIMO Downlink Performance with Calibration Errors

Kohei Ueda, Anubhab Chowdhury, Koji Ishibashi, Erik G. Larsson

This paper analyzes the effects of calibration errors on downlink beamforming in a repeater-assisted massive MIMO system and derives analytical expressions for the downlink spectral efficiency.

View →