20 results for “Familiarity with evidence grades and their importance in clinical research.”
CS papers onlyHybrid search: Keyword + semantic, ranked by combined score.ⓘ
Want pure semantic search? Try claim verification →
This paper evaluates the ability of large language models to recover and express evidence grades from clinical claims, finding that while they can recover the grades, they do not consistently express…
This paper introduces MIRA-Ev, a clinical argument mining benchmark in Spanish, English, and Basque with a three-tier task hierarchy for evidence sentence retrieval, argumentative component extraction…
Pin Qian, Su Wang, Xiaoyuan Wang, Yihang Chen +6 more
The paper introduces FORCEBENCH, a new stress test designed to evaluate whether cited sources genuinely warrant the strength of a claim, revealing that standard citation evaluation methods often fail…
AutoSynthesis is an end-to-end multi-agent system for automated meta-analysis in evidence synthesis, including search strategy formulation, literature retrieval, screening, eligibility assessment, sta…
This paper critically reviews the intersection of philosophy of science and explainable AI (XAI) in medicine, identifying necessary conditions for a philosophically grounded approach to explanation.
The paper introduces Factual Density (FD*), a novel retrieval signal that measures the proportion of verified facts, demonstrating that optimizing RAG retrieval based on this density significantly imp…
The paper introduces ClinPivot, a benchmark that tests whether clinical models can correctly adjust treatment decisions when new patient context constraints are introduced, finding that strong medical…
Qing Wang, Bo Li, Jialu Liang, Daling Shi +2 more
The paper introduces DrugClaw, a multi-agent system, and DrugAudit, a new benchmark, demonstrating that DrugClaw excels at answering drug-related questions by grounding answers in primary regulatory s…
This paper introduces MMed-Bench-IR, a benchmark for multilingual medical retrieval in clinical settings, evaluating cross-lingual alignment, concept discrimination, and evidence retrieval.
The paper introduces the Causal Sensitivity Score (CSS), an interventional metric that reveals that standard coverage-based evaluations fail to detect critical responsiveness deficits in clinical LLMs…
The paper demonstrates that models can acquire 'evaluation meta-knowledge' from training data describing evaluation practices, leading to inflated safety benchmark performance that is independent of e…
This study analyzes ClinicalTrials.gov records to track the rising trend of AI in clinical trials and demonstrates that a hybrid human-AI screening approach is viable but requires clearer reporting of…
Zhaoyang Jiang, Xuanqi Peng, Fei Teng, Zhizhong Fu +4 more
The paper demonstrates that while distilling large language models for medical QA can significantly improve final answer accuracy, this gain often comes at the cost of factual accuracy and detailed re…
AutoForest is an end-to-end system that automatically generates publication-ready forest plots directly from biomedical papers, streamlining the labor-intensive process of meta-analysis.
The paper introduces CERA, a novel contrastive retrieval framework that improves RAG factuality and interpretability by using subjectivity-based hard negative selection and an auxiliary attention alig…
Liuliu Chen, Gowri Rajaram, Eleanor Bailey, Katrina Witt +4 more
The paper introduces an evidence-augmented machine learning approach to improve self-harm surveillance by analyzing Emergency Department triage notes, achieving high and transferable performance acros…
The paper evaluates the semantic stability of clinical LLMs to linguistic variations, finding that domain specialization does not guarantee consistent robustness improvements.
This study examines public perceptions of automated decision-making in healthcare using data from an ongoing longitudinal survey panel. Factors influencing perceived helpfulness, riskiness, and fairne…