ArXivCSExplorer
☆☆Bookmarks🏆RSSHow to UseFAQ
Built with and by Teycir Ben Soltane•
How to Use•FAQ•GitHub•arXiv.org•
Share:

20 results for “Quoridor”

CS papers only

Hybrid search: Keyword + semantic, ranked by combined score.ⓘ

Want pure semantic search? Try claim verification →

cs.CCTheoreticalRecentJun 27, 2026

One Hex reduction to rule them all: Quoridor, Maze Attack, Pinko Pallino and Blockade are PSPACE-complete

Francesco Carboni, Daniele Muscillo

The paper settles the computational complexity of Quoridor and related games by reducing them to Reisch's planar graph-Hex.

View →
cs.LGcs.DCcs.MATheoreticalRecentJul 17, 2026

The Honest Quorum Problem: Epistemic Byzantine Fault Tolerance for Agentic Infrastructure

Jun He, Deying Yu

This paper introduces Epistemic Byzantine Fault Tolerance (EBFT), a fault-tolerance model for agentic infrastructure and post-deterministic distributed systems, addressing the Honest Quorum Problem an…

View →
cs.AIRecentMay 27, 2026

VeriTrip: A Verifiable Benchmark for Travel Planning Agents over Unstructured Web Corpora

Yuting Xu, Jiayi Tian, Jian Liang, Xin Xiong +3 more

The paper introduces VeriTrip, a new verifiable benchmark that evaluates travel planning agents' ability to perform evidence-grounded reasoning over complex, unstructured, and multimodal web data, rev…

View →
cs.CLcs.AIRecentJun 2, 2026

QUBRIC: Co-Designing Queries and Rubrics for RL Beyond Verifiable Rewards

Rongzhi Zhang, Rui Feng, Zhihan Zhang, Jingfeng Yang +7 more

QUBRIC introduces a co-design framework that simultaneously optimizes queries and rubrics, overcoming the bottleneck of vague rubrics derived from open-ended questions, leading to significant gains in…

View →
cs.AIEmpiricalRecentJun 30, 2026

FARS: A Fully Automated Research System Deployed at Scale

Qiong Tang, Xiangkun Hu, Xiangyang Liu, Yiran Chen +1 more

FARS is a fully automated AI-for-AI research system that generated and advanced 166 complete research papers across 67 topics in a large-scale public deployment, with evaluations from 282 reviews.

View →
cs.CLRecentMay 31, 2026

DrugClaw and DrugAudit: A Primary-Source-Grounded Agent and Authority-Aware Benchmark for Drug-Information Question Answering

Qing Wang, Bo Li, Jialu Liang, Daling Shi +2 more

The paper introduces DrugClaw, a multi-agent system, and DrugAudit, a new benchmark, demonstrating that DrugClaw excels at answering drug-related questions by grounding answers in primary regulatory s…

View →
cs.AIcs.HCcs.SEEmpiricalRecentJun 29, 2026

Rehearsed Multi-Agent Live Product Demonstrations with Real-Time Voice Question Answering

Rahul Khedar, Mayank Malhotra, Avinash Karn, Mouli V +1 more

This paper proposes Rhetor, a multi-agent system that generates rehearsed live demonstrations with segment-synchronized narration and real-time voice question answering for web applications.

View →
cs.AIcs.DCcs.MARecentMay 27, 2026

SwarmHarness: Skill-Based Task Routing via Decentralized Incentive-Aligned AI Agent Networks

Edwin Jose

SwarmHarness introduces a decentralized, incentive-aligned protocol enabling self-organizing compute swarms for AI tasks, eliminating the need for central coordinators or heavy blockchain infrastructu…

View →
cs.IRcs.CLcs.MAEmpiricalRecentJul 20, 2026

FinSAgent: Corpus-Aligned Multi-Agent RAG Framework for Evidence-Grounded SEC Filing Question Answering

Jijun Chi, Zhenghan Tai, Hanwei Wu, Tung Sum Thomas Kwok +19 more

This paper proposes FinSAgent, an evidence-grounded multi-agent framework for financial question answering over SEC filings, which improves retrieval coverage and answer correctness through corpus-sid…

View →
cs.AIcs.CRcs.LGRecentMay 17, 2026

ADR: An Agentic Detection System for Enterprise Agentic AI Security

Chenning Li, Pan Hu, Justin Xu, Baris Ozbas +8 more

The paper introduces ADR, a novel, production-proven detection system that provides high-fidelity security monitoring for AI agents operating via the Model Context Protocol, significantly outperformin…

View →
cs.IRcs.CYcs.LGEmpiricalRecentJun 26, 2026

Reproducing FACTER: Fairness via Conformal Thresholding and Prompt Repair

Oscar Miró López-Feliu, Daimy van Loo, Xanthos Kekkos, Mikel Blom +1 more

The paper conducts a reproducibility study on FACTER, a model-agnostic framework for fairness and statistical coverage in LLM-based recommendation, and evaluates its consistency and contribution.

View →
cs.LGcs.AIcs.CRRecentJun 2, 2026

RUBAS: Rubric-Based Reinforcement Learning for Agent Safety

Xian Qi Loye, Qinglin Su, Zhexin Zhang, Shiyao Cui +4 more

The paper introduces RUBAS, a rubric-based reinforcement learning framework that improves agent safety by providing fine-grained, multi-dimensional rewards for complex tool-use scenarios.

View →
cs.CLEmpiricalRecentJul 2, 2026

CheckRLM: Effective Knowledge-Thought Coherence Checking in Retrieval-Augmented Reasoning

Dingling Xu, Ruobing Wang, Qingfei Zhao, Yukun Yan +7 more

The paper proposes CheckRLM, a framework that improves the reliability of Reasoning Language Models by identifying and correcting factual errors using Retrieval-Augmented Generation.

View →
cs.IREmpiricalRecentJul 22, 2026

CIR at iKAT SCAI 2026: Exploring Clarification Need Prediction in Agentic Conversational Search

Nolwenn Bernard, Jüri Keller, Philipp Schaer

The Cologne Information Retrieval group participated in iKAT SCAI 2026 shared task using an agentic conversational search system with query rewriting, retrieval, reranking, answer generation, and clar…

View →
cs.CRcs.AIcs.CYRecentApr 13, 2026

Hardening x402: PII-Safe Agentic Payments via Pre-Execution Metadata Filtering

Vladimir Stantchev

The paper introduces presidio-hardened-x402, an open-source middleware that intercepts x402 payment requests to detect and redact PII and enforce spending policies before on-chain settlement.

View →
cs.SEcs.CRRecentMar 18, 2026

Who Tests the Testers? Systematic Enumeration and Coverage Audit of LLM Agent Tool Call Safety

Xuan Chen, Lu Yan, Ruqi Zhang, Xiangyu Zhang

The paper introduces SafeAudit, a meta-audit framework that systematically enumerates test cases and uses a quantitative metric to uncover significant residual unsafe behaviors in LLM agents that exis…

View →
cs.AIcs.CLEmpiricalRecentJul 23, 2026

OpenForgeRL: Train Harness-native Agents in Any Environment

Xiao Yu, Baolin Peng, Ruize Xu, Hao Zou +6 more

OpenForgeRL is an open-source framework for training harness-based AI agents end-to-end in various environments using a lightweight proxy and Kubernetes orchestrator.

View →
cs.CLeess.ASEmpiricalRecentJul 19, 2026

Robust Summarization of Doctor-Patient Conversations: TalTech Systems for the Beyond Transcription Challenge

Aivo Olev, Tanel Alumäe

TalTech submitted top-ranking systems to the Beyond Transcription Challenge using fine-tuned Voxtral models and reinforcement learning against Open Medical Concept F1.

View →
cs.CLcs.AIRecentMay 27, 2026

SafeRx-Agent: A Knowledge-Grounded Multi-Agent Framework for Safe and Explainable Medication Recommendation

Xinyu Wang, Hanwei Wu, Zhenghan Tai, Sicheng Lyu +6 more

The paper introduces SafeRx-Agent, a knowledge-grounded multi-agent framework that improves medication recommendation accuracy and safety by incorporating fine-grained ATC codes and rigorous safety ve…

View →