ArXivCSExplorer
☆☆Bookmarks🏆RSSHow to UseFAQ
Built with and by Teycir Ben Soltane•
How to Use•FAQ•GitHub•arXiv.org•
Share:

20 results for “Familiarity with fault tolerance concepts”

CS papers only

Hybrid search: Keyword + semantic, ranked by combined score.ⓘ

Want pure semantic search? Try claim verification →

cs.CLcs.AIcs.LGRecentMay 28, 2026

The Architecture of Errors: From Universal Impossibility to Patch-Local LLM Reliability

Mikhail L. Arbuzov, Lee Mosbacker, Sisong Bei, Ziwei Dong +2 more

The paper reframes LLM reliability from an impossible universal problem to a manageable, local patch-based problem, showing that sufficient interventions can be found by focusing on recurring failure…

View →
cs.LGcs.DCcs.MATheoreticalRecentJul 17, 2026

The Honest Quorum Problem: Epistemic Byzantine Fault Tolerance for Agentic Infrastructure

Jun He, Deying Yu

This paper introduces Epistemic Byzantine Fault Tolerance (EBFT), a fault-tolerance model for agentic infrastructure and post-deterministic distributed systems, addressing the Honest Quorum Problem an…

View →
cs.CCcs.AIcs.FLTheoreticalRecentJun 23, 2026

Token Complexity of Certifying Stochastic-Oracle Reliability

Jie Wang

This paper introduces a framework for certifying the reliability of stochastic oracles and derives bounds on the minimum expected token cost for reliable oracle certification.

View →
cs.DCcs.AIRecentJun 1, 2026

Not All Errors Are Equal: A Systematic Study of Error Propagation in Large Language Model Inference

Yafan Huang, Sheng Di, Guanpeng Li

This paper systematically studies how soft errors propagate during Large Language Model (LLM) inference using a novel fault-injection framework, providing critical insights and mitigation strategies f…

View →
cs.SEcs.AIcs.CLRecentMay 28, 2026

REPOT: Recoverable Program-of-Thought via Checkpoint Repair

Parsa Mazaheri

The paper introduces RePoT, a method that significantly improves Program-of-Thought (PoT) planning by deterministically verifying the initial plan prefix and using a single LLM call to resume planning…

View →
cs.DCTheoreticalRecentJul 27, 2026

Consensus In Asynchrony: Strictly Formal

Ivan Klianev

The authors resolve the contradiction between deterministic crash-tolerant consensus in a fully asynchronous environment and the FLP impossibility result by demonstrating that one protocol phase separ…

View →
cs.AIRecentMay 27, 2026

When Context Flips, Safety Breaks: Diagnosing Brittle Safety in Aligned Language Models

Dasol Choi, Alex Kwon

The paper introduces 'brittle safety,' a failure mode where aligned language models fail to adapt their safety behavior when a situational context changes, and proposes state-aware validation to detec…

View →
cs.LGcs.AIcs.DCRecentJun 1, 2026

Post-Deterministic Distributed Systems: A New Foundation for Trustworthy Autonomous Infrastructure

Jun He, Deying Yu

The paper introduces Post-Deterministic Distributed Systems (PDDS) as a new model to coordinate autonomous infrastructure where participants, including stochastic agents, produce divergent reasoning p…

View →
cs.CRRecentApr 25, 2026

Core Logic and Algorithmic Performance Enhancements for a System Vulnerability Analysis Technique for Complex Mission Critical Systems Implementation

Matthew Tassava, Cameron Kolodjski, Jordan Milbrath, Jeremy Straub

The paper details significant enhancements to the SONARR system's core logic, replacing restrictive Boolean logic with generic data type support and adding multi-compute capabilities to improve vulner…

View →
cs.LGcs.AIcs.CLRecentJun 3, 2026

Failed Reasoning Traces Tell You What Is Fixable (But Not by Reading Them)

Nizar Islah, Istabrak Abbes, Irina Rish, Sarath Chandar +1 more

This paper proposes a method to recover recoverability structure from failed traces of post-trained language models, enabling test-time routing and post-training analysis.

View →
cs.SEcs.AIEmpiricalRecentJul 22, 2026

Beyond Fail-to-Pass: Iterative Hardening of Co-Generated Bug Reproduction Tests and Fixes

Yuhao Tan, Zhibang Yang, Fangkai Yang, Yuan Yao +8 more

This paper proposes CoHarden, a co-generation framework for automated program repair that uses a lax signal as an in-loop convergence criterion to prevent lax regressions.

View →
cs.LOcs.PLEmpiricalRecentJun 26, 2026

KoAT: Automatic Complexity and Termination Analysis of Integer Programs

Nils Lommen, Éléanore Meyer, Jürgen Giesl

KoAT is a tool that automatically infers complexity bounds and proves termination of integer programs using an alternating modular analysis approach and a portfolio of techniques.

View →
cs.SEEmpiricalRecentJun 18, 2026

N-Version Programming with Coding Agents

Javier Ron, Benoit Baudry, Martin Monperrus

This paper revisits N-version programming with AI coding agents and finds substantial common-mode failures but also practical benefits.

View →
cs.AREmpiricalRecentJul 13, 2026

Reliable Associative Lookup in Content-Addressable Memory

Fan Li, Yanan Guo, Xin Xin

This paper introduces a new protection code design for Content Addressable Memory (CAM) to ensure reliability.

View →
cs.AIcs.CRcs.SERecentMay 24, 2026

Inverting the Shield: Systematically Generating Safety Tests from Policy Specifications

Xiaoyue Lu, Xianglin Yang, Haijun Liu, Jiahao Liu +3 more

The paper introduces POLARIS, a novel framework that systematically generates comprehensive and verifiable safety tests for LLMs by formalizing natural language policies into First-Order Logic and exp…

View →
cs.ARcs.AITheoreticalRecentJul 22, 2026

Formal Foundations for Known Good Reliable Die Screening in Chiplet-Based AI Systems-on-Chip

Prashanthi Metku, Chandra Gandu

This paper presents a methodology for transitioning from Known Good Die (KGD) to Known Good Reliable Die (KGRD) screening in chiplet-based artificial intelligence systems-on-chips (SoCs) as a constrai…

View →
cs.SEEmpiricalRecentJul 22, 2026

SequenceFI: Non-intrusive Temporal Fault Injection for Microservice Systems

Yuzhen Tan, Jian Wang, Bing Li, Shaolin Tan

This paper introduces SequenceFI, a non-intrusive framework for temporal fault injection in microservice systems, which observes message-level events, synthesizes temporal guards from traces, and achi…

View →