ArXivCSExplorer
☆☆Bookmarks🏆RSSHow to UseFAQ
Built with and by Teycir Ben Soltane•
How to Use•FAQ•GitHub•arXiv.org•
Share:

20 results for “Extractive summarization”

CS papers only

Hybrid search: Keyword + semantic, ranked by combined score.ⓘ

Want pure semantic search? Try claim verification →

cs.IREmpiricalRecentJul 11, 2026

SVD-RAG: Efficient Tree-Organized Retrieval-Augmented Generation via Singular Value Decomposition

Zhihui Sun

This paper introduces SVD-RAG, a cost-efficient and content-adaptive summarization method for hierarchical Retrieval-Augmented Generation systems using Singular Value Decomposition on dense sentence e…

View →
cs.CLcs.AIRecentMay 31, 2026

Understanding LLM Behavior in Multi-Target Cross-Lingual Summarization

Sangwon Ryu, Yihong Liu, Mingyang Wang, Yunsu Kim +3 more

The paper introduces a new benchmark for multi-target cross-lingual summarization (MTXLS) and proposes an activation steering method that significantly improves LLM performance by guiding the generati…

View →
cs.CLcs.AIcs.LGRecentMay 27, 2026

Enhancing BiGRU with a KAN Block for Legal Document Classification and Summarization

Ahmed Faizul Haque Dhrubo, Souvik Pramanik, Most. Aysha Siddika Sumona, Shahnewaz Siddique +3 more

The paper proposes a novel KAN-enhanced BiGRU architecture to improve legal document classification and summarization in a low-resource, multilingual setting using Bengali and English legal texts.

View →
cs.CLRecentMay 31, 2026

Peacemaker at ATE-IT: Automatic term extraction from Italian text for waste management data using encoder model

Mahdi Bakhtiyarzadeh, Hadi Bayrami Asl Tekanlou, Jafar Razmara

The paper proposes a low-cost and interpretable fine-tuning extraction strategy for automatic term extraction, demonstrating consistent and balanced performance on the ATE Shared Task.

View →
cs.IRcs.CLDatasetRecentJun 9, 2026

A PubMed-Scale Dataset of Structured Biomedical Abstracts

Chia-Hsuan Chang, Haerin Song, Brian Ondov, Hua Xu

The authors introduce Structured PubMed, a comprehensive corpus of section-labeled biomedical abstracts compiled from the complete PubMed database.

View →
cs.CLcs.AIcs.MAEmpiricalRecentJul 16, 2026

Dialogue Summarization with Emotion Dynamics Using Topic- and Participant-Centric Decomposition

Linyun Xiang, Mark Neerincx, Stephanie Tan

This paper proposes a framework for summarizing dialogues, modeling semantic and emotion dynamics using multimodal inputs and an adapted hierarchical Chain-of-Agents approach.

View →
cs.DLcs.CLRecentMay 31, 2026

Digging Up Citations: FOSSIL, a Dataset and Workflow for Reference Extraction in Law and the Humanities

Luca Foppiano, Christian Boulanger

The paper introduces FOSSIL, a new multilingual dataset and specialized workflow designed to significantly improve the extraction of citations embedded within complex footnotes common in law and human…

View →
cs.IREmpiricalRecentJul 6, 2026

Prompting Beats Fine-Tuning: Generative Expected Value Scoring for Statutory Term Retrieval

Alvin Wang, Jaromir Savelka

The paper compares two families of methods for ranking case-law sentences by their usefulness for explaining statutory concepts using ModernBERT and decoder-only models. Decoder-only models achieve th…

View →
cs.AIRecentJun 1, 2026

An NLP-Driven Framework for Curriculum-Labor Market Alignment: Schema-Constrained LLM Extraction, ESCO-Anchored Semantic Matching, and Multi-Dimensional Gap Quantification

Sherzod Turaev, Mary John, Mamoun Awad, Nazar Zaki +1 more

The paper introduces a robust four-stage NLP framework that uses schema-constrained LLMs and ESCO vocabulary to accurately extract and align educational competencies with labor market demands, quantif…

View →
cs.CLRecentJun 1, 2026

Towards Multidisciplinary Summarization of Hospital Stays: Efficient Sentence-Level Clinical Provenance Categorization

Baris Karacan, Vaibhav Bhargava, Barbara Di Eugenio, Natalie Parde +20 more

The paper introduces a supervised fine-tuning pipeline using large language models to accurately categorize sentence-level clinical provenance across multi-disciplinary hospital notes, demonstrating t…

View →
cs.IRcs.AIcs.CLEmpiricalRecentJul 2, 2026

Evaluating Chunking Strategies for Retrieval-Augmented Generation on Academic Texts

Valentin J. J. Kreileder, Johannes Reisinger, Andreas Fischer

This paper evaluates the effectiveness of cluster-based semantic chunking compared to fixed-size and recursive chunking in Retrieval-Augmented Generation systems using the Retrieval Augmented Generati…

View →
cs.AIRecentMay 30, 2026

Ryze: Evidence-Enriched Data Synthesis from Biomedical Papers

Yeqi Huang, Yue Chen, Yanwei Ye, Guanhao Su +1 more

The paper introduces Ryze, an automated system that synthesizes evidence-enriched Question-Answering (QA) pairs from raw biomedical papers, resulting in a specialized VLM (BioVLM-8B) that significantl…

View →
cs.HCcs.IREmpiricalRecentJun 26, 2026

Context-Aware Explanations for Spatialized Document Layouts

Wei Liu, John Wenskovitch, Chris North, Rebecca Faust

The paper presents CAPE, a framework that generates natural-language explanations for spatially organized document layouts, using context-aware representations and LLM-based explanation generation.

View →
cs.IREmpiricalRecentJul 24, 2026

Legal Nugget Extraction for Granular Retrieval over Long Jurisprudential Texts

Lucas Pereira, Erick Brito, Roberto Lotufo, Jayr Pereira

This paper explores the use of legal nuggets for dense retrieval over Brazilian legal collections, achieving significant improvements in two datasets.

View →
cs.DBcs.AIcs.DCEmpiricalRecentJul 3, 2026

Scalable Maximal Frequent Episode Mining with Desbordante

Maxim Ivanov, Matvei Smirnov, Alisa Strazdina, George Chernishev

This paper presents an improved parallel version of MaxFEM algorithm for maximal frequent episode mining, achieving up to 8x speedup in C++ implementation and 35x improvement overall.

View →
cs.CLeess.ASEmpiricalRecentJul 19, 2026

Robust Summarization of Doctor-Patient Conversations: TalTech Systems for the Beyond Transcription Challenge

Aivo Olev, Tanel Alumäe

TalTech submitted top-ranking systems to the Beyond Transcription Challenge using fine-tuned Voxtral models and reinforcement learning against Open Medical Concept F1.

View →
cs.CLcs.IREmpiricalRecentJun 10, 2026

uva-irlab-conv at SemEval-2026 Task 8: Multi-Turn RAG with Learned Sparse Retrieval and Listwise Reranking

Simon Lupart, Kidist Amde Mekonnen, Zahra Abbasiantaeb, Mohammad Aliannejadi

This paper proposes a multi-turn retrieval-augmented generation pipeline for conversational systems across four domains.

View →
cs.CLEmpiricalRecentJul 3, 2026

The Classics at SemEval-2026 Task 3: Combining Transformer Models and LLM-Generated Annotations for Dimensional Aspect-Based Sentiment Analysis

Rafif Alshawi, Amit Raj, Aleksey Kudelya, Alexander Shirnin

This paper proposes an approach for fine-grained sentiment analysis using regression and extraction tasks, involving a weighted ensemble of transformer-based encoder models and a large language model…

View →
cs.IREmpiricalRecentJun 26, 2026

Listwise Explanation of Embedding-Based Rankings via Semantic Chunk Grouping

Hyunkyu Kim, Yeeun Yoo, Youngjun Kwak

The paper introduces ChunkGroupSHAP, a listwise Shapley method that clusters semantantly related chunks into shared cross-document features for dense semantic ranking.

View →