20 results for “feedback loop”
CS papers onlyHybrid search: Keyword + semantic, ranked by combined score.ⓘ
Want pure semantic search? Try claim verification →
The paper proposes Self-Trained Verification (STV), a novel method that trains verifiers to catch self-generated errors by leveraging reference solutions, significantly boosting performance in both te…
The paper proposes CTRL-STEER, a closed-loop framework that adaptively adjusts intervention strength to stabilize concept regulation and improve task success in Vision-Language-Action models without r…
This paper surveys 1,250 arXiv papers on self-improving AI systems, categorizing them based on what they improve and the degree of loop closure. It identifies a distinctive feature of self-evaluation…
This paper proposes two strategies to improve feedback efficiency of reinforcement learning from human feedback (RLHF) in diffusion models.
Guangyuan Wu, Weining Cao, Zehui Tan, Yuan Yao +3 more
This paper introduces InvWeaver, a neuro-symbolic framework for synthesizing loop invariants in programs with multiple interacting loops.
This paper proposes the stable signal principle to explain the convergence behavior of retraining in performative prediction systems, revealing regularization as a natural force to control performativ…
This paper develops a convergence theory and runtime bound for closed-loop generative selection in computational drug discovery, showing that elitism makes the search absorbing and proving almost-sure…
This paper measures the lower bound for the shortest program generating a sequence, proving a conservation law and providing a deterministic engine to recover generating programs for certain sequences…
The paper introduces a Jacobian-based spectral audit to evaluate neural operators, demonstrating that standard prediction error metrics fail to capture crucial local dynamical structures and operator…
Pekka Malo, Lauri Viitasaari, Patrik Nummi, Antti Suominen +2 more
The paper introduces an operator calculus for population-based optimization methods, establishing a modular Lyapunov principle for their convergence analysis.
This paper proposes a new imitation learning algorithm called DistIL that uses distributional feedback to improve policy improvement and regret guarantees.
Chengleyang Lei, Wei Feng, Yanmin Wang, Yunfei Chen +3 more
This paper proposes a holistic design framework for a wireless-powered SC3 system using satellite RF signals, optimizing energy transfer, communication, and computing processes.
This paper presents a self-balancing sampler for sequential sampling that achieves faster convergence to a desired target law while maintaining unpredictability.
Yunsheng Zeng, Gen Li, Yuwei Miao, Xiandong Li +7 more
The paper proposes EAPO, an entropy-driven adaptive weighting method that dynamically adjusts the influence of positive samples during policy optimization to improve both response diversity and stabil…
Jek Huang, Jeffery Hsia, Jiayi Sun, Freddie Shi +2 more
This paper introduces Proof-or-Stop Lifecycle Control, a method that allows lifecycle transitions only when mechanically verifiable evidence is provided, and evaluates its implementation.