ArXivCSExplorer
☆☆Bookmarks🏆RSSHow to UseFAQ
Built with and by Teycir Ben Soltane•
How to Use•FAQ•GitHub•arXiv.org•
Share:

20 results for “Familiarity with stochastic systems”

CS papers only

Hybrid search: Keyword + semantic, ranked by combined score.ⓘ

Want pure semantic search? Try claim verification →

cs.LGcs.AImath.NARecentMay 28, 2026

Stochastic Lifting for Generating Trajectories of Stochastic Physical Systems

Jules Berman, Tobias Blickhan, Benjamin Peherstorfer

Stochastic Lifting is a novel technique that enhances the modeling of stochastic physical systems by introducing independent random labels to state transitions, allowing a single network to generate d…

View →
math.OCcs.LGcs.MATheoreticalRecentJul 1, 2026

Mean Field Reinforcement Learning

René Carmona, Mathieu Laurière

This paper introduces mean field reinforcement learning through Markov decision processes in large-population stochastic control, developing the necessary framework for representative-agent learning,…

View →
stat.MLcs.LGmath.STRecentJun 3, 2026

Bayesian learning for the stochastic shortest path problem

Chon Wai Ho, Sumeetpal S. Singh, Jiaqi Guo

The paper proposes a novel Bayesian framework to learn the optimal decision strategy for the stochastic shortest path problem by directly constructing the posterior beliefs for the action-value functi…

View →
stat.MLcs.LGmath.STEmpiricalRecentJul 26, 2026

Learning switched non-linear dynamical systems from a single trajectory

Sunny G. W. Wang, Hemant Tyagi

This paper derives non-asymptotic bounds on the prediction risk for learning switched non-linear dynamical systems, providing the first guarantees from a single trajectory.

View →
cs.LGcs.AIstat.MLRecentMay 28, 2026

The Sample Complexity of Multiclass and Sparse Contextual Bandits

Liad Erez, Fan Chen, Alon Cohen, Tomer Koren +3 more

The paper analyzes the sample complexity of contextual bandits in the $s$-sparse setting, achieving optimal sample bounds for identifying an $\epsilon$-optimal policy.

View →
cs.LGcs.AIRecentMay 30, 2026

Extending Causal Metamodeling to a non-Markovian Queue

Pracheta Amaranath, Anant Bhide, David Jensen, Peter Haas

The paper extends modular dynamic Bayesian networks (MDBNs) to model non-Markovian queues, providing the first causal metamodeling technique for such systems with significant speedup.

View →
stat.MLcs.LGmath.PRTheoreticalRecentJun 16, 2026

A Diffusion Approximation for Temporal-Difference Learning with Linear Features under Markovian Noise

M. Forzo, E. Monzio Compagnoni, A. Russo, A. Pacchiano

This paper introduces a stochastic differential equation approximation for linear Temporal Difference (TD) learning under Markovian noise, explaining the constant-stepsize error floor.

View →
cs.CCcs.AIcs.FLTheoreticalRecentJun 23, 2026

Token Complexity of Certifying Stochastic-Oracle Reliability

Jie Wang

This paper introduces a framework for certifying the reliability of stochastic oracles and derives bounds on the minimum expected token cost for reliable oracle certification.

View →
stat.MLcs.LGTheoreticalRecentJul 19, 2026

Non-Asymptotic Best Policy Identification Guarantees in Online Reinforcement Learning

Joseph Lazzaro, Alessio Russo, Aldo Pacchiano

This paper provides non-asytotic sample complexity guarantees for the Navigate and Stop algorithm in online tabular Reinforcement Learning, identifying additional attributes that affect the overall sa…

View →
cs.LGcs.AIRecentMay 29, 2026

From Rashomon Theory to PRAXIS: Efficient Decision Tree Rashomon Sets

Zakk Heile, Hayden McTavish, Varun Babbar, Margo Seltzer +1 more

The paper introduces PRAXIS, a novel algorithm that efficiently approximates the computation of 'Rashomon sets' for decision trees, significantly reducing memory and runtime complexity.

View →
cs.AIcs.CLRecentMay 27, 2026

The Importance of Being Statistically Earnest: A Critical Re-evaluation of GSM-Symbolic

Dominika Agnieszka Długosz, Arlindo Oliveira, Natalia Díaz-Rodríguez

The paper challenges the conclusion that LLMs lack reasoning by demonstrating that reported performance drops on GSM-Symbolic are often statistically weak and partially attributable to dataset biases,…

View →
stat.MLcs.LGTheoreticalRecentJul 3, 2026

A Hierarchy of Policy Learning Problems

Hamsa Bastani, Osbert Bastani, Shihan Chen

This paper provides a mathematical framework for studying different policy learning problems and shows reductions between them.

View →
cs.NITheoreticalRecentJul 21, 2026

Online Stochastic Matchings: Stability on Hypergraphs

Fabien Mathieu

This paper characterizes stabilizability of stochastic dynamic matching on hypergraphs using the arrival rates and incidence matrix, and proposes a stabilizing policy.

View →
cs.CLcs.AIEmpiricalRecentJun 30, 2026

Reinforcement Learning with Metacognitive Feedback Elicits Faithful Uncertainty Expression in LLMs

Gabrielle Kaili-May Liu, Avi Caciularu, Gal Yona, Idan Szpektor +1 more

This paper introduces two novel mechanisms, reinforcement learning with metacognitive feedback (RLMF) and metacognitive data selection, to enhance language model metacognition and achieve faithful cal…

View →
cs.DMmath.PRTheoreticalRecentJul 26, 2026

Stability in stochastic hypergraph matching I: necessary and sufficient criteria

Doan Dai Nguyen, Ana Bušić

This paper introduces online assignment policies for stochastic matching on hypergraphs, which are maximally stable and allow for the derivation of necessary and sufficient stability criteria.

View →
cs.LGcs.AIstat.MLRecentMay 29, 2026

Why Linear Recurrent Memory Works in Partially Observable Reinforcement Learning

Yike Zhao, Onno Eberhard, Malek Khammassi, Ali H. Sayed +1 more

This paper theoretically justifies the strong performance of linear recurrent neural networks as memory units in partially observable reinforcement learning by constructing specific linear filters tha…

View →
cs.LGcs.ITstat.MLTheoreticalRecentJun 28, 2026

Bayesian Best-Arm Identification with Abstention: A Polynomial-to-Exponential Phase Transition

Yuqi Huang, Yunlong Hou, Vincent Y. F. Tan

This paper analyzes the Bayesian fixed-budget best-arm identification problem with abstention, showing that it induces a phase transition from polynomial to exponential decay of error probability.

View →