ArXivCSExplorer
☆☆Bookmarks🏆RSSHow to UseFAQ
Built with and by Teycir Ben Soltane•
How to Use•FAQ•GitHub•arXiv.org•
Share:

20 results for “prefix tuning”

CS papers only

Hybrid search: Keyword + semantic, ranked by combined score.ⓘ

Want pure semantic search? Try claim verification →

cs.CLcs.AIcs.SDEmpiricalRecentJul 6, 2026

REDDIT: Correcting Model-Generated Timestamp Drift in ASR without Forgetting via Replay-Based Distribution Editing

Cheng-Kang Chou, Ming-To Chuang, Ke-Han Lu, Chan-Jan Hsu +1 more

This paper investigates timestamp drift in modern autoregressive ASR systems and proposes REDDIT, a two-stage post-training framework to correct timestamps while avoiding forgetting.

View →
cs.LGcs.AIRecentMay 27, 2026

Learning the Error Patterns of Language Models

Jinwoo Kim, Taylor Berg-KirkPatrick, Loris D'Antoni

The paper introduces prefix filters and an algorithm (Palla) to systematically learn and apply specific error patterns in Large Language Models, significantly improving constrained generation tasks li…

View →
cs.LGcs.CLEmpiricalRecentJul 8, 2026

Max Out GRPO Signal: Adaptive Trace Prefix Control for Hard Reasoning Problems

Vladislav Beliaev

The paper introduces AdaPrefix-GRPO, a method that adjusts the amount of reference solution assistance during training to improve the success rate and accuracy of Group Relative Policy Optimization (G…

View →
cs.CLcs.AIRecentMay 27, 2026

PrunePath: Towards Highly Structured Sparse Language Models

Zhexuan Gu, Zixun Fu, Yancheng Yuan

PrunePath introduces a budget-adaptive structured sparsification framework that efficiently prunes Feed-forward networks in large language models, achieving hardware-friendly sparsity and measurable s…

View →
cs.AIRecentMay 28, 2026

Accelerating Constrained Decoding with Token Space Compression

Michael Sullivan, Alexander Koller

The paper introduces CFGzip, an offline token space compression technique that significantly reduces the computational overhead of constrained decoding, making complex grammar enforcement feasible at…

View →
cs.AIRecentMay 30, 2026

TAPS: Target-Aware Prefix Tree Selection for Diffusion-Drafted Speculative Decoding

Zhuoyu Wang, Junnan Huang, Xinyu Chen

TAPS introduces a target-aware prefix selection method that optimizes the trade-off between draft tree acceptance and verification cost, achieving significant speedups in speculative decoding.

View →
cs.CLRecentMay 29, 2026

TRACE: Discovering Task-Specific Parameter via Adaptation-Aware Probing for Continual Fine-Tuning

Xiaosong Han, Ke Chen, Xindi Dai, Di Liang +6 more

TRACE proposes a novel method to mitigate catastrophic forgetting in continual LLM fine-tuning by identifying and isolating a small, task-specific subset of essential parameters for each task.

View →
cs.AIRecentJun 1, 2026

Revisiting Ripple Effects in Knowledge Editing through Pressure-Aware Joint Neighborhood Optimization

Haoben Huang, Shuxin Liu, Ou Wu, Di Gao

The paper proposes Joint Neighborhood Optimization (JNO), a novel knowledge-editing framework that jointly addresses the coupled pressures of desirable knowledge propagation and unintended knowledge l…

View →
cs.CRcs.AIcs.CLRecentApr 16, 2026

Route to Rome Attack: Directing LLM Routers to Expensive Models via Adversarial Suffix Optimization

Haochun Tang, Yuliang Yan, Jiahua Lu, Huaxiao Liu +1 more

The paper introduces R$^2$A, an adversarial attack that uses suffix optimization to mislead black-box LLM routers into consistently selecting expensive, high-capability models.

View →
cs.SDeess.ASEmpiricalRecentJun 19, 2026

Gradient-Based Learning of Parametric Engine Sound Representations for Real-Time Resynthesis and Tuning on Embedded Systems

Robin Doerfler, Matthieu Kuntz, Clemens Zimmer

This paper proposes a neural network-based approach for combustion engine sound modeling, extending conventional methods with machine learning and stochastic components.

View →
cs.CLcs.AIRecentJun 1, 2026

From Layers to Submodules: Rethinking Granularity in Replacement-Based LLM Compression

Elia Cunegatti, Marcus Vukojevic, Erik Nielsen, Giovanni Iacca

The paper proposes SubFit, a novel compression technique that achieves superior LLM compression by replacing non-contiguous, submodule-level components (Attention and FeedForward) with lightweight res…

View →
cs.FLcs.DScs.LGEmpiricalRecentJul 19, 2026

Stringological sequence prediction II: Right-to-left automaticity and related complexity measures

Vanessa Kosoy

This paper presents an efficient algorithm for right-to-left sequence prediction based on a new complexity measure called arithmetic repetition complexity, and demonstrates its application to predicti…

View →
cs.CRcs.AIcs.LGEmpiricalRecentJul 22, 2026

HijackKV: New Threat in Position-Independent KV Cache Reuse

Yichi Zhang, Zhiqi Wang, Huan Zhang, Yuchen Yang

This paper introduces HIJACKKV, the first attack framework to exploit the new threat of KV Cache Hijacking in large language models, achieving an average success rate of 94%.

View →
cs.PLcs.LOEmpiricalRecentJul 28, 2026

Towards Bottom-Up Enumeration in miniKanren via Pruning and Memoization

Nikolai Kudasov

The paper introduces two library combinators, prune and defrel/bank, to bring observational deduplication and bottom-up enumeration into the relational setting using plain miniKanren.

View →
cs.DSTheoreticalRecentJun 26, 2026

Incremental Submodular Maximization: Better Than Greedy

Marcin Bienkowski, Joakim Blikstad, Jarosław Byrka, Martín Costa +2 more

The paper presents an adaptive scaling algorithm with a competitive ratio of 1.373 for incremental submodular maximization under increasing cardinality constraint, improving upon the previous best res…

View →
cs.AREmpiricalRecentJul 6, 2026

NEMESIS: NEtlist-Driven Modeling and Equation Synthesis with Inversion-Aware SPICE Anchoring

Subhadip Ghosh, Ramesh Harjani, Sachin S. Sapatnekar

NEMESIS is a framework that uses large language models to generate accurate performance equations for operational transconductance amplifiers (OTAs) with a balance between speed and accuracy.

View →
stat.MLcs.LGEmpiricalRecentJul 23, 2026

Automatic knot selection in smooth additive models

Nicolás Carrizosa, Vanesa Guerrero, María Durbán

A new method for selecting knots in Generalized Additive Models using an extension of adaptive splines and a customized Fellner-Schall scheme.

View →
cs.LGcs.AIRecentJun 1, 2026

FOAM: Frequency and Operator Error-Based Adaptive Damping Method for Reducing Staleness-Oriented Error for Shampoo

Kyunghun Nam, Sumyeong Ahn

The paper proposes FOAM, an adaptive damping method that stabilizes the Shampoo optimization algorithm by dynamically controlling damping and eigendecomposition frequency, thereby reducing staleness-i…

View →