ArXivCSExplorer
☆☆Bookmarks🏆RSSHow to UseFAQ
Built with and by Teycir Ben Soltane•
How to Use•FAQ•GitHub•arXiv.org•
Share:

20 results for “Statistical models for ordinal data”

CS papers only

Hybrid search: Keyword + semantic, ranked by combined score.ⓘ

Want pure semantic search? Try claim verification →

cs.CLcs.LGRecentMay 29, 2026

Pairwise Reference Alignment as a Model-Level Ordinal Observable

Mujing Li

The paper provides a formal statistical and conceptual framework for defining and measuring 'pairwise reference alignment,' which quantifies how well a model's scoring function agrees with a given ref…

View →
cs.HCEmpiricalRecentJul 2, 2026

Adapting CCDF Plots for Visualizing Ordinal Regression Results

Abhraneel Sarma

The paper proposes using modified Complementary Cumulative Distribution Function (mCCDF) plots to visualize and communicate results of cumulative-link ordinal regression models for ordinal data.

View →
cs.AIEmpiricalRecentJul 16, 2026

Can We Trust Item Response Theory for AI Evaluation?

Han Jiang, Sunbeom Kwon, Jinwen Luo, Ziang Xiao +1 more

This paper evaluates the reliability of using item response theory (IRT) models for AI benchmarking, comparing four estimation tools under various simulation conditions.

View →
stat.MEcs.CYcs.HCEmpiricalRecentJun 23, 2026

When Surveys Become Conversations: Adaptive Matrix Validation for AI-Assisted Interviews

Tyler H. McCormick

This paper proposes Adaptive Matrix Validation (AMV), a method for validating AI-assisted interview data in surveys using statistical adjustment.

View →
cs.LGcs.AIRecentMay 29, 2026

From Rashomon Theory to PRAXIS: Efficient Decision Tree Rashomon Sets

Zakk Heile, Hayden McTavish, Varun Babbar, Margo Seltzer +1 more

The paper introduces PRAXIS, a novel algorithm that efficiently approximates the computation of 'Rashomon sets' for decision trees, significantly reducing memory and runtime complexity.

View →
cs.CLRecentMay 28, 2026

Auditing LLM Benchmarks with Item Response Theory

Sander Land, Daniel M. Bikel

The paper introduces an Item Response Theory (IRT)-based indicator that effectively identifies likely mislabeled items in existing LLM benchmarks, revealing systematic errors in labeling and model spe…

View →
stat.MLcs.LGEmpiricalRecentJun 12, 2026

Gradient boosting for extremes: sampling theory and application to insurance

Stéphane Lhaut, Olivier Lopez

This paper develops statistical learning theory for gradient boosting in Peaks-over-Threshold modeling using Generalized Pareto distributions, deriving error bounds and reducing gradient correlation.

View →
cs.AIcs.CLRecentMay 27, 2026

The Importance of Being Statistically Earnest: A Critical Re-evaluation of GSM-Symbolic

Dominika Agnieszka Długosz, Arlindo Oliveira, Natalia Díaz-Rodríguez

The paper challenges the conclusion that LLMs lack reasoning by demonstrating that reported performance drops on GSM-Symbolic are often statistically weak and partially attributable to dataset biases,…

View →
cs.DScs.DMTheoreticalRecentJun 11, 2026

(Un)ranking Permutation Classes

Nathanaël Hassler, Vincent Vajnovszki

This paper presents methods for ranking and unranking permutations avoiding a pattern of length three in lexicographic or colexicographic order.

View →
cs.DMcs.DSTheoreticalRecentJul 8, 2026

Ranking and Rank Aggregation with Matroid Prefix Constraints

Seiei Ando, Yu Yokoi

This paper studies ranking and aggregation under Kendall tau distance with matroid or flag matroid constraints on prefixes.

View →
cs.LGcs.AIstat.MLRecentMay 28, 2026

Calibrated Preference Learning: The Case of Label Ranking

Santo M. A. R. Thies, Viktor Bengs, Timo Kaufmann, Sebastian J. Vollmer +1 more

The paper formalizes the concept of calibration for probabilistic label ranking, demonstrating that popular models are often poorly calibrated and that calibration captures a meaningful quality dimens…

View →
stat.MLcs.LGstat.MEEmpiricalRecentJul 26, 2026

Distributional Split Criteria for Random Forests: Extensions, Shrinkage, and the Robustness of Mean Splitting

Silas Koemen

This paper introduces Distributional Random Forests, which replace mean-based CART splitting with criteria that compare full conditional response distributions in candidate children. The authors syste…

View →
cs.CLcs.AIRecentMay 28, 2026

Internal Representation, Not Clinical Knowledge: Where Apparent LLM Triage Failures Originate

David Fraile Navarro, Berardino Como, Jialei Sheng, Soundariya Ananthan +1 more

The paper investigates apparent LLM triage failures and concludes that the errors originate in the output format and decision process, rather than a deficiency in the model's underlying clinical knowl…

View →
cs.ITcs.LGmath.STTheoreticalRecentJul 3, 2026

Open Problem: Is Interaction Necessary for Order-Optimal 1-bit Mean Estimation?

Ivan Lau, Jonathan Scarlett

This paper investigates the necessity of interaction for order-optimal 1-bit mean estimation in nonparametric finite-moment classes.

View →
cs.CLRecentMay 29, 2026

Reliable Multilingual Orthopedic Decision Support from Clinical Narratives: Language-Aware Adaptation and Verification-Guided Deferral

Danish Ali, Li Xiaojian, Sundas Iqbal, Farrukh Zaidi

The paper introduces a reliability-oriented framework, IndicBERT-HPA, for multilingual orthopedic decision support from clinical narratives, achieving high performance and significantly improving reli…

View →
cs.CLcs.LGRecentJun 1, 2026

Machine Learning for Coding Retail Product Names to Consumer-Price Categories: A Rule-plus-Bag-of-Words Pipeline with Reliability-Weighted Human-in-the-Loop Labeling

Vladimir Beskorovainyi

The paper proposes a robust, multi-stage pipeline combining rule-based classification and machine learning to map noisy retail product names to standardized consumption categories, finding that simple…

View →
stat.MLcs.LGstat.MEEmpiricalRecentJul 2, 2026

Autorelevance function and other feature relevance measures for univariate time series

Julian Cardenas, Jamie Arjona, Pedro Delicado

The paper proposes methodologies to measure lag relevance in machine learning forecasting models using Ghost variables, Shapley values, and additive importance measures. It also introduces auto-releva…

View →