20 results for “Credit annotations”
CS papers onlyHybrid search: Keyword + semantic, ranked by combined score.ⓘ
Want pure semantic search? Try claim verification →
The paper introduces FOSSIL, a new multilingual dataset and specialized workflow designed to significantly improve the extraction of citations embedded within complex footnotes common in law and human…
The paper introduces Score Broadcast and Decorrelation (SBD), a general theoretical framework that unifies broadcast-based credit assignment across various differentiable loss functions by leveraging…
This paper conducts a large-scale audit of human annotation reporting in NLP, finding that while reporting has improved, critical details needed to assess annotation validity, such as training and agr…
This paper proposes a framework for reward allocation in AI cooperatives using value-conditioned gradient filtering, online marginal contribution signals, and cumulative revenue settlement within a tr…
SCAFDS introduces a novel, seven-stage graph attention system that models fraud propagation using co-occurrence edge features and generates forensically traceable SAR narratives, significantly improvi…
The paper introduces ARCA, a novel credit assignment method that measures token salience directly from the adapter's residual hidden state, addressing the degeneracy of standard intrinsic signals when…
The paper introduces a typed claim network that models cross-document references by explicitly labeling the stance (e.g., agreement, disagreement) of a citation, significantly improving downstream tas…
A prototype annotation tool is presented that allows humans to identify important information and LLMs to match it to system outputs, creating a division of labor for reliable AI evaluation.
Rocio Jimenez-Villen, Ziwei Xu, Ying Chen, Oscar Araque +1 more
The paper presents FinKG-News, a framework that constructs factual, company-centric knowledge graphs from news events to improve credit risk report generation.
Wenwu Li, Yuran Song, Mingze Zhao, Bo Jin +1 more
The paper proposes a novel temporal and structural credit assignment framework to efficiently optimize multi-agent LLM systems by decomposing the error signal and using targeted, discrete gradient upd…
The paper proposes a robust, multi-stage pipeline combining rule-based classification and machine learning to map noisy retail product names to standardized consumption categories, finding that simple…
Shuheng Cao, Ruiqi Chen, Renjie Cao, Zhenhao Zhang +2 more
The paper introduces BioConCal, a supervised scoring mechanism that evaluates biomedical NER candidates surfaced by multiple LLMs, significantly improving the quality of the candidate pool for human c…
Pin Qian, Su Wang, Xiaoyuan Wang, Yihang Chen +6 more
The paper introduces FORCEBENCH, a new stress test designed to evaluate whether cited sources genuinely warrant the strength of a claim, revealing that standard citation evaluation methods often fail…
Jungyeul Park, Kyungtae Lim, Wonjun Oh, Benjamin Nguyen +3 more
This paper refines word-based grammatical error annotation for L2 Korean by adapting existing resources to better reflect Korean morphology and error types, improving the evaluation of Korean Grammati…
This paper compares specialized supervised Extreme Multi-Label Classification (XMLC) methods with lexical matching baselines and LLM-based methods for subject indexing contemporary German scientific l…
This paper extends gradient boosting to functions of vector inputs using a simple algorithm with histogram-based decision trees.
The paper compares two families of methods for ranking case-law sentences by their usefulness for explaining statutory concepts using ModernBERT and decoder-only models. Decoder-only models achieve th…
The paper extends the User Experience Research (UXR) Points of View (PoV) framework into an AI-augmented methodology specifically designed for guiding the development and governance of high-stakes, hu…
This paper audits eight automatic scorers for attribution in LLM retrieval-augmented generation and finds that none of them transfer across datasets for generated-answer attribution.