ArXivCSExplorer
☆☆Bookmarks🏆RSSHow to UseFAQ
Built with and by Teycir Ben Soltane•
How to Use•FAQ•GitHub•arXiv.org•
Share:

20 results for “calibrated experts”

CS papers only

Hybrid search: Keyword + semantic, ranked by combined score.ⓘ

Want pure semantic search? Try claim verification →

cs.GTcs.DSTheoreticalRecentJul 9, 2026

Algorithmic Expert Aggregation

Wei Tang, Hanrui Zhang

This paper studies the problem of aggregating calibrated Bayesian experts into a new calibrated expert.

View →
cs.AIcs.LGEmpiricalRecentJun 18, 2026

Toward Calibrated Mixture-of-Experts Under Distribution Shift

Gina Wong, Drew Prinster, Suchi Saria, Rama Chellappa +1 more

This paper studies the behavior of mixture-of-experts (MoE) models under distribution shift and proposes an adversarial reweighting method to improve their calibration.

View →
cs.LGcs.AIstat.MLRecentMay 28, 2026

CalArena: A Large-Scale Post-Hoc Calibration Benchmark

Eugène Berta, David Holzmüller, Francis Bach, Michael I. Jordan

The paper introduces CalArena, a large-scale, standardized benchmark covering nearly 2000 experiments to comprehensively evaluate post-hoc calibration methods, finding that smooth calibration function…

View →
cs.LGcs.AIstat.MLRecentMay 28, 2026

Calibrated Preference Learning: The Case of Label Ranking

Santo M. A. R. Thies, Viktor Bengs, Timo Kaufmann, Sebastian J. Vollmer +1 more

The paper formalizes the concept of calibration for probabilistic label ranking, demonstrating that popular models are often poorly calibrated and that calibration captures a meaningful quality dimens…

View →
stat.MLcs.LGEmpiricalRecentJun 12, 2026

Gradient boosting for extremes: sampling theory and application to insurance

Stéphane Lhaut, Olivier Lopez

This paper develops statistical learning theory for gradient boosting in Peaks-over-Threshold modeling using Generalized Pareto distributions, deriving error bounds and reducing gradient correlation.

View →
cs.LGcs.AIRecentMay 27, 2026

Online Irregular Multivariate Time Series Forecasting via Uncertainty-Driven Dual-Expert Calibration

Haonan Wen, Hanyang Chen, Songhe Feng

The paper proposes Under-Cali, an uncertainty-driven dual-expert calibration framework, to achieve stable and efficient online forecasting for irregularly sampled multivariate time series.

View →
cs.CLstat.MERecentMay 31, 2026

A Finite-Calibration Regime Map for LLM Judge Panels

Bin Zhu, Yanghui Rao

The paper proposes a finite-calibration regime map to determine the optimal calibration method (low-dimensional stackers vs. joint tables) for LLM judge panels given limited human labeling budgets, sh…

View →
cs.CLRecentMay 28, 2026

Counterfactual Graph for Multi-Agent LLM Calibration

Jiatan Huang, Mingchen Li, Ziming Li, Sunjae Kwon +2 more

The paper proposes CAGE-CAL, a counterfactual graph calibration framework, to accurately assess the reliability and detect over-confidence in multi-agent LLM systems after agents communicate.

View →
cs.LGEmpiricalRecentJun 30, 2026

Calibration, Not Compilation: Detecting and Repairing Misspecified Probabilistic Programs Written by Language Models

Jian Xu, Delu Zeng, John Paisley, Qibin Zhao

This paper evaluates the effectiveness of Bayesian workflow for verifying statistical correctness of probabilistic programs written by language models, and compares it to unit tests and no feedback.

View →
cs.AIRecentMay 27, 2026

Relevant Is Not Warranted: Evidence-Force Calibration for Cited RAG

Pin Qian, Su Wang, Xiaoyuan Wang, Yihang Chen +6 more

The paper introduces FORCEBENCH, a new stress test designed to evaluate whether cited sources genuinely warrant the strength of a claim, revealing that standard citation evaluation methods often fail…

View →
cs.CLcs.AIRecentJun 2, 2026

Quantifying Faithful Confidence Expression in Large Reasoning Models

Areeb Gani, Asal Meskin, Gabrielle Kaili-May Liu, Arman Cohan

The paper introduces a novel framework to quantify faithful confidence expression (FC) in Large Reasoning Models (LRMs), finding that FC remains a significant and challenging reliability target for th…

View →
cs.AIEmpiricalRecentJul 24, 2026

The Regression Tax: Decomposing Why Skills Help and Hurt LLM Agents

Darshan Tank, Baran Nama

This paper measures the impact of procedural skills on LLM agents, distinguishing between improvements and regressions, and identifies causes of regression.

View →
cs.AIcs.CLcs.LGEmpiricalRecentJul 2, 2026

Online Safety Monitoring for LLMs

Mona Schirmer, Metod Jazbec, Alexander Timans, Christian Naesseth +2 more

This paper proposes a simple real-time monitor for LLMs that turns an external verifier signal into an alarm decision by thresholding, showing competitiveness with advanced monitors in mathematical re…

View →
cs.CRcs.LGRecentMay 23, 2026

CALIBURN: A Regime-Sensitivity Study of Operationally Calibrated Streaming Intrusion Detection

Michel A. Youssef

CALIBURN introduces a novel, five-component streaming pipeline for intrusion detection that allows operators to specify alerting behavior using cost and budget constraints, achieving state-of-the-art…

View →
cs.LGcs.AIRecentMay 28, 2026

DAMEL: Dual-Axis Multi-Expert Learning for Class-Imbalanced Learning

Hyuck Lee, Taemin Park, Heeyoung Kim

The paper proposes DAMEL, a dual-axis multi-expert learning algorithm that simultaneously reduces both prediction bias and variance in class-imbalanced learning by leveraging multiple experts across b…

View →
cs.CRcs.AIcs.MARecentMay 1, 2026

Skills as Verifiable Artifacts: A Trust Schema and a Biconditional Correctness Criterion for Human-in-the-Loop Agent Runtimes

Alfredo Metere

The paper proposes a trust schema and verification framework to ensure that agent skills, which augment LLMs, are rigorously verified before deployment, thereby making human-in-the-loop oversight scal…

View →