20 results for “Familiarity with AI systems”
CS papers onlyHybrid search: Keyword + semantic, ranked by combined score.ⓘ
Want pure semantic search? Try claim verification →
This paper proposes a definition for 'AI-nativeness' in systems, based on an AI agent's authority over system decisions.
Aakash Pant, Kavya Shah, Apoorv Agnihotri, Sneha Nikam +2 more
The paper critiques current AI benchmarking practices for low-resource settings, arguing that evaluation must shift focus from isolated model performance to the holistic performance of the deployed sy…
The paper argues that purported anthropomorphic attributes of LLMs are not unique to language models but are substrate-dependent, demonstrating this by training a neural network on the game Age of Emp…
The paper introduces BusinessCaseBench, a benchmark for measuring AI performance on analytical knowledge work using business cases and grading rubrics.
This study examines the impact of Microsoft 365 Copilot adoption on employee acceptance in a state Department of Transportation using a matched two-wave survey.
Maharshi Gor, Yoo Yeon Sung, Yu Hou, Eve Fleisig +3 more
This study investigates human-AI collaboration in question answering, finding that while collaboration is beneficial, humans make suboptimal decisions by both under-relying on correct AI suggestions a…
Jaechang Kim, Sunung Mun, Seungjoon Lee, Jaewoong Cho +1 more
The paper proposes Faithful Agentic XAI (FAX), a verification framework that explicitly checks LLM-generated explanations against model behavior, significantly improving explanation faithfulness on a…
This paper reports the first classroom deployment of LEA, an adaptive AI tutoring agent, with real students across three courses and evaluates its cross-course scalability.
The paper introduces Hyperparam, a set of lightweight JavaScript libraries designed to enable direct, model-aware querying of unstructured data (like agent traces) within client-side AI applications.
A study evaluating a Retrieval-Augmented Generation system for making Quebec automobile insurance contracts more understandable, showing it improves satisfaction, trust, and clarity, especially for in…
The paper defines AI Identity as the correspondence between an agent's declared state and its observed behavior, concluding that current infrastructure and standards are fundamentally inadequate for g…
RealityTest introduces a large-scale, multimodal, and multilingual benchmark using real-world human data to test how AI systems disclose their identity, finding that context and phrasing are more crit…
This paper challenges the claim that neural networks have met the challenge of systematicity in language and thought as proposed by Fodor and Pylyshyn, demonstrating limitations in a recent neural net…
The paper introduces MATRA, a systematic threat modeling framework, to assess how known LLM threats translate into concrete, deployment-specific risks within autonomous agentic AI systems.
The paper introduces Clover, a code completion tool that logs students' interactions and offers attention checks to promote reflective engagement during programming tasks.
David Ayllon, Alice Baird, Jeffrey Brooks, Franc Camps-Febrer +10 more
The paper introduces the Real World Voice EQ Bench, a multidimensional benchmark for evaluating voice AI across text-to-speech, speech-to-speech, speech understanding, and automatic speech recognition…
This paper documents and analyzes the failure process of strategies used to address conceptual drift in long-horizon LLM collaboration and introduces the concept of 'Index Sickness' and the 'Pang Prin…
Five experiments investigated how access to AI advice affects human judgment and willingness to suspend judgment, finding that it nearly eliminated participants' willingness to do so and led to less a…