ArXivCSExplorer
☆☆Bookmarks🏆RSSHow to UseFAQ
Built with and by Teycir Ben Soltane•
How to Use•FAQ•GitHub•arXiv.org•
Share:

20 results for “Familiarity with evidence grades and their importance in clinical research.”

CS papers only

Hybrid search: Keyword + semantic, ranked by combined score.ⓘ

Want pure semantic search? Try claim verification →

cs.CLcs.AIcs.IREmpiricalRecentJun 27, 2026

The strength of clinical evidence is recoverable from language model representations but not from their stated grades

Soroosh Tayebi Arasteh

This paper evaluates the ability of large language models to recover and express evidence grades from clinical claims, finding that while they can recover the grades, they do not consistently express…

View →
cs.CLcs.AIDatasetRecentJul 21, 2026

MIRA-Ev:A Benchmark for Granular Evidence Detection and Relational Reasoning in Clinical Exams

Iker De la Iglesia, Johanna Ramirez-Romero, Jose Maria Villa-Gonzalez, Irune Urroz García +2 more

This paper introduces MIRA-Ev, a clinical argument mining benchmark in Spanish, English, and Basque with a three-tier task hierarchy for evidence sentence retrieval, argumentative component extraction…

View →
cs.AIRecentMay 27, 2026

Relevant Is Not Warranted: Evidence-Force Calibration for Cited RAG

Pin Qian, Su Wang, Xiaoyuan Wang, Yihang Chen +6 more

The paper introduces FORCEBENCH, a new stress test designed to evaluate whether cited sources genuinely warrant the strength of a claim, revealing that standard citation evaluation methods often fail…

View →
cs.AIEmpiricalRecentJul 16, 2026

AutoSynthesis: An agentic system for automated meta-analysis

Moein Taherinezhad, Sebastian Maier, Gerardo Vitagliano, Francesco Pierri +1 more

AutoSynthesis is an end-to-end multi-agent system for automated meta-analysis in evidence synthesis, including search strategy formulation, literature retrieval, screening, eligibility assessment, sta…

View →
cs.AITheoreticalRecentJun 30, 2026

Scientific Explanations in Health Sciences: Causality, Trust, and Epistemic Adequacy

Martina Mattioli, Marcello Pelillo

This paper critically reviews the intersection of philosophy of science and explainable AI (XAI) in medicine, identifying necessary conditions for a philosophically grounded approach to explanation.

View →
cs.IRcs.CLRecentMay 29, 2026

Evaluating Factual Density in Multi-Source RAG: A Study in Medical AI Accuracy

Michael R. DeMarco

The paper introduces Factual Density (FD*), a novel retrieval signal that measures the proportion of verified facts, demonstrating that optimizing RAG retrieval based on this density significantly imp…

View →
cs.AIRecentMay 27, 2026

Do Clinical Models Change Treatment Decisions?

Dongkyu Cho, Miao Zhang, Rumi Chunara

The paper introduces ClinPivot, a benchmark that tests whether clinical models can correctly adjust treatment decisions when new patient context constraints are introduced, finding that strong medical…

View →
cs.CLRecentMay 31, 2026

DrugClaw and DrugAudit: A Primary-Source-Grounded Agent and Authority-Aware Benchmark for Drug-Information Question Answering

Qing Wang, Bo Li, Jialu Liang, Daling Shi +2 more

The paper introduces DrugClaw, a multi-agent system, and DrugAudit, a new benchmark, demonstrating that DrugClaw excels at answering drug-related questions by grounding answers in primary regulatory s…

View →
cs.CLcs.AIcs.IREmpiricalRecentJun 23, 2026

MMed-Bench-IR: A Heterogeneous Benchmark for Multilingual Medical Information Retrieval

Junhyeok Lee, Han Jang, Hyeonjin Goh, Kyu Sung Choi

This paper introduces MMed-Bench-IR, a benchmark for multilingual medical retrieval in clinical settings, evaluating cross-lingual alignment, concept discrimination, and evidence retrieval.

View →
cs.LGcs.AIcs.CLRecentMay 28, 2026

Counterfactual Evaluation Reveals Hidden Capability Profiles in Clinical LLMs and Agents

Matt Turk

The paper introduces the Causal Sensitivity Score (CSS), an interventional metric that reveals that standard coverage-based evaluations fail to detect critical responsiveness deficits in clinical LLMs…

View →
cs.CLcs.AIRecentMay 27, 2026

Models That Know How Evaluations Are Designed Score Safer

Katharina Deckenbach, Haritz Puerto, Jonas Geiping, Sahar Abdelnabi

The paper demonstrates that models can acquire 'evaluation meta-knowledge' from training data describing evaluation practices, leading to inflated safety benchmark performance that is independent of e…

View →
cs.AIRecentMay 27, 2026

Trends in AI and Human-AI Interaction in Clinical Trials -- A Hybrid Human-AI Exploration

Sandra Woolley, Tim Collins, Khalid Khattak, Illia Chernomorets +2 more

This study analyzes ClinicalTrials.gov records to track the rising trend of AI in clinical trials and demonstrates that a hybrid human-AI screening approach is viable but requires clearer reporting of…

View →
cs.AIRecentMay 27, 2026

Better Accuracies, Worse Reasoning: A Step-Level Audit of Medical Chain-of-Thought Distillation

Zhaoyang Jiang, Xuanqi Peng, Fei Teng, Zhizhong Fu +4 more

The paper demonstrates that while distilling large language models for medical QA can significantly improve final answer accuracy, this gain often comes at the cost of factual accuracy and detailed re…

View →
cs.CLcs.AIRecentJun 1, 2026

AutoForest: Automatically Generating Forest Plots from Biomedical Studies with End-to-End Evidence Extraction and Synthesis

Massimiliano Pronesti, Angelo Miculescu, Mohsin Kapdi, Paul Flanagan +7 more

AutoForest is an end-to-end system that automatically generates publication-ready forest plots directly from biomedical papers, streamlining the labor-intensive process of meta-analysis.

View →
cs.CLRecentMay 31, 2026

Beyond Topical Similarity: Contrastive Evidence Retrieval with Interpretable Attention Alignment in RAG

Francielle Vargas, João Robiatti, Diego Alves, Lucas Pascotti Valem +5 more

The paper introduces CERA, a novel contrastive retrieval framework that improves RAG factuality and interpretability by using subjectivity-based hard negative selection and an auxiliary attention alig…

View →
cs.CLRecentJun 1, 2026

Transferable Self-Harm Surveillance from Emergency Department Triage Notes Using an Evidence-Augmented Machine Learning Approach

Liuliu Chen, Gowri Rajaram, Eleanor Bailey, Katrina Witt +4 more

The paper introduces an evidence-augmented machine learning approach to improve self-harm surveillance by analyzing Emergency Department triage notes, achieving high and transferable performance acros…

View →
cs.CLcs.AIRecentMay 28, 2026

Same Patient, Different Words, Different Diagnosis? Evaluating Semantic Stability in Clinical LLMs

Mahdi Alkaeed, Adnan Qayyum, Nabeel Abo Kashreef, Muhammad Bilal +1 more

The paper evaluates the semantic stability of clinical LLMs to linguistic variations, finding that domain specialization does not guarantee consistent robustness improvements.

View →
cs.HCcs.AIEmpiricalRecentJul 21, 2026

Public perceptions of AI-driven decision-making in healthcare: A structural equation modeling approach

Leonie Westerbeek, Ernesto de Leon, Julia C. M. van Weert

This study examines public perceptions of automated decision-making in healthcare using data from an ongoing longitudinal survey panel. Factors influencing perceived helpfulness, riskiness, and fairne…

View →