ArXivCSExplorer
☆☆Bookmarks🏆RSSHow to UseFAQ
Built with and by Teycir Ben Soltane•
How to Use•FAQ•GitHub•arXiv.org•
Share:

20 results for “Familiarity with AI systems”

CS papers only

Hybrid search: Keyword + semantic, ranked by combined score.ⓘ

Want pure semantic search? Try claim verification →

cs.AIcs.OSTheoreticalRecentJul 22, 2026

Defining AI-Native Systems: Autonomy as Revision Authority

Cheng Tan

This paper proposes a definition for 'AI-nativeness' in systems, based on an AI agent's authority over system decisions.

View →
cs.AIRecentMay 27, 2026

Benchmarking AI for low-resource contexts: Thinking beyond leaderboards

Aakash Pant, Kavya Shah, Apoorv Agnihotri, Sneha Nikam +2 more

The paper critiques current AI benchmarking practices for low-resource settings, arguing that evaluation must shift focus from isolated model performance to the holistic performance of the deployed sy…

View →
cs.CLcs.AIcs.CYRecentMay 29, 2026

If LLMs Have Human-Like Attributes, Then So Does Age of Empires II

Adrian de Wynter

The paper argues that purported anthropomorphic attributes of LLMs are not unique to language models but are substrate-dependent, demonstrating this by training a neural network on the game Age of Emp…

View →
cs.CLcs.AIEmpiricalRecentJul 17, 2026

Frontier AI performance across the business disciplines: a case-grounded benchmark of knowledge work and analytical reasoning

Ajay Patel, Kartik Hosanagar, Ramayya Krishnan, Chris Callison-Burch +2 more

The paper introduces BusinessCaseBench, a benchmark for measuring AI performance on analytical knowledge work using business cases and grading rubrics.

View →
cs.CYcs.HCEmpiricalRecentJul 15, 2026

Persona Migration and Expectation Recalibration in Generative AI Adoption: A Longitudinal Study at a State Department of Transportation

Omidreza Shoghli, Fatemeh Banani Ardecani, Amin Mohamadi Hezaveh

This study examines the impact of Microsoft 365 Copilot adoption on employee acceptance in a state Department of Transportation using a matched two-wave survey.

View →
cs.AIcs.CLcs.HCRecentMay 27, 2026

AI, Take the Wheel: What Drives Delegation and Trust in Human-Computer Cooperative Question Answering?

Maharshi Gor, Yoo Yeon Sung, Yu Hou, Eve Fleisig +3 more

This study investigates human-AI collaboration in question answering, finding that while collaboration is beneficial, humans make suboptimal decisions by both under-relying on correct AI suggestions a…

View →
cs.AIRecentMay 27, 2026

Towards Faithful Agentic XAI: A Verification Method and an Open-World Benchmark for Better Model Faithfulness

Jaechang Kim, Sunung Mun, Seungjoon Lee, Jaewoong Cho +1 more

The paper proposes Faithful Agentic XAI (FAX), a verification framework that explicitly checks LLM-generated explanations against model behavior, significantly improving explanation faithfulness on a…

View →
cs.CYcs.AIcs.HCEmpiricalRecentJul 15, 2026

Learning Engagement Assistant (LEA): Cross-Course Scalability and Classroom Evaluation of an Agentic AI Tutoring System

Teri Rumble, Javad Zarrin, P. George Lovell, Ruth Falconer

This paper reports the first classroom deployment of LEA, an adaptive AI tutoring agent, with real students across three courses and evaluates its cross-course scalability.

View →
cs.AIcs.DBRecentMay 27, 2026

A Query Engine for the Agents

Kenny Daniel

The paper introduces Hyperparam, a set of lightweight JavaScript libraries designed to enable direct, model-aware querying of unstructured data (like agent traces) within client-side AI applications.

View →
cs.HCEmpiricalRecentJul 17, 2026

A Human-Centric Evaluation of a Retrieval-Augmented Generation System for Explaining Quebec Insurance Contracts

David Beauchemin, Richard Khoury

A study evaluating a Retrieval-Augmented Generation system for making Quebec automobile insurance contracts more understandable, showing it improves satisfaction, trust, and clarity, especially for in…

View →
cs.AIcs.CRRecentApr 25, 2026

AI Identity: Standards, Gaps, and Research Directions for AI Agents

Takumi Otsuka, Kentaroh Toyoda, Alex Leung

The paper defines AI Identity as the correspondence between an agent's declared state and its observed behavior, concluding that current infrastructure and standards are fundamentally inadequate for g…

View →
cs.CLRecentMay 29, 2026

RealityTest: How People Probe AI Identity and Whether Models Disclose It

Anna Gausen, Sarenne Wallbridge, Bessie O'Dell, Christopher Summerfield +1 more

RealityTest introduces a large-scale, multimodal, and multilingual benchmark using real-world human data to test how AI systems disclose their identity, finding that context and phrasing are more crit…

View →
cs.CLcs.AIEmpiricalRecentJun 12, 2026

Fodor and Pylyshyn's Systematicity Challenge Still Stands

Michael Goodale, Salvador Mascarenhas

This paper challenges the claim that neural networks have met the challenge of systematicity in language and thought as proposed by Fodor and Pylyshyn, demonstrating limitations in a recent neural net…

View →
cs.AIcs.CRRecentMay 11, 2026

MATRA: Modeling the Attack Surface of Agentic AI Systems -- OpenClaw Case Study

Tim Van hamme, Thomas Vissers, Javier Carnerero-Cano, Mario Fritz +3 more

The paper introduces MATRA, a systematic threat modeling framework, to assess how known LLM threats translate into concrete, deployment-specific risks within autonomous agentic AI systems.

View →
cs.HCcs.AIcs.SEEmpiricalRecentJun 29, 2026

To Tab or Not to Tab: Measuring Critical Engagement in AI Code Completion Tools Using Behavioral Signals and Attention Checks

Jessica Hutchison, Ian Tyler Applebaum, Kenneth Angelikas, Kush Rakesh Patel +5 more

The paper introduces Clover, a code completion tool that logs students' interactions and offers attention checks to promote reflective engagement during programming tasks.

View →
cs.SDcs.AIEmpiricalRecentJul 16, 2026

RW-Voice-EQ Bench: A Real World Benchmark for Evaluating Voice AI Systems

David Ayllon, Alice Baird, Jeffrey Brooks, Franc Camps-Febrer +10 more

The paper introduces the Real World Voice EQ Bench, a multidimensional benchmark for evaluating voice AI across text-to-speech, speech-to-speech, speech understanding, and automatic speech recognition…

View →
cs.SEcs.CLcs.HCEmpiricalRecentJun 17, 2026

Written by AI, Managed by AI: Semantic Space Control and Index Sickness Elimination Across 391 Consecutive Sessions

Hui Zhang, Shuren Song

This paper documents and analyzes the failure process of strategies used to address conceptual drift in long-horizon LLM collaboration and introduces the concept of 'Index Sickness' and the 'Pang Prin…

View →
cs.AIcs.CYcs.HCEmpiricalRecentJul 15, 2026

AI advice suppresses people's willingness to say "I don't know", even when the advice is wrong and accuracy is incentivized

Chiara Marcoccia, Walter Quattrociocchi, Valerio Capraro

Five experiments investigated how access to AI advice affects human judgment and willingness to suspend judgment, finding that it nearly eliminated participants' willingness to do so and led to less a…

View →