20 results for “Familiarity with fault tolerance concepts”
CS papers onlyHybrid search: Keyword + semantic, ranked by combined score.ⓘ
Want pure semantic search? Try claim verification →
Mikhail L. Arbuzov, Lee Mosbacker, Sisong Bei, Ziwei Dong +2 more
The paper reframes LLM reliability from an impossible universal problem to a manageable, local patch-based problem, showing that sufficient interventions can be found by focusing on recurring failure…
This paper introduces Epistemic Byzantine Fault Tolerance (EBFT), a fault-tolerance model for agentic infrastructure and post-deterministic distributed systems, addressing the Honest Quorum Problem an…
This paper introduces a framework for certifying the reliability of stochastic oracles and derives bounds on the minimum expected token cost for reliable oracle certification.
This paper systematically studies how soft errors propagate during Large Language Model (LLM) inference using a novel fault-injection framework, providing critical insights and mitigation strategies f…
The paper introduces RePoT, a method that significantly improves Program-of-Thought (PoT) planning by deterministically verifying the initial plan prefix and using a single LLM call to resume planning…
The authors resolve the contradiction between deterministic crash-tolerant consensus in a fully asynchronous environment and the FLP impossibility result by demonstrating that one protocol phase separ…
The paper introduces 'brittle safety,' a failure mode where aligned language models fail to adapt their safety behavior when a situational context changes, and proposes state-aware validation to detec…
The paper introduces Post-Deterministic Distributed Systems (PDDS) as a new model to coordinate autonomous infrastructure where participants, including stochastic agents, produce divergent reasoning p…
The paper details significant enhancements to the SONARR system's core logic, replacing restrictive Boolean logic with generic data type support and adding multi-compute capabilities to improve vulner…
Nizar Islah, Istabrak Abbes, Irina Rish, Sarath Chandar +1 more
This paper proposes a method to recover recoverability structure from failed traces of post-trained language models, enabling test-time routing and post-training analysis.
Yuhao Tan, Zhibang Yang, Fangkai Yang, Yuan Yao +8 more
This paper proposes CoHarden, a co-generation framework for automated program repair that uses a lax signal as an in-loop convergence criterion to prevent lax regressions.
KoAT is a tool that automatically infers complexity bounds and proves termination of integer programs using an alternating modular analysis approach and a portfolio of techniques.
This paper revisits N-version programming with AI coding agents and finds substantial common-mode failures but also practical benefits.
Xiaoyue Lu, Xianglin Yang, Haijun Liu, Jiahao Liu +3 more
The paper introduces POLARIS, a novel framework that systematically generates comprehensive and verifiable safety tests for LLMs by formalizing natural language policies into First-Order Logic and exp…
This paper presents a methodology for transitioning from Known Good Die (KGD) to Known Good Reliable Die (KGRD) screening in chiplet-based artificial intelligence systems-on-chips (SoCs) as a constrai…
This paper introduces SequenceFI, a non-intrusive framework for temporal fault injection in microservice systems, which observes message-level events, synthesizes temporal guards from traces, and achi…