20 results for “prefix tuning”
CS papers onlyHybrid search: Keyword + semantic, ranked by combined score.ⓘ
Want pure semantic search? Try claim verification →
Cheng-Kang Chou, Ming-To Chuang, Ke-Han Lu, Chan-Jan Hsu +1 more
This paper investigates timestamp drift in modern autoregressive ASR systems and proposes REDDIT, a two-stage post-training framework to correct timestamps while avoiding forgetting.
The paper introduces prefix filters and an algorithm (Palla) to systematically learn and apply specific error patterns in Large Language Models, significantly improving constrained generation tasks li…
The paper introduces AdaPrefix-GRPO, a method that adjusts the amount of reference solution assistance during training to improve the success rate and accuracy of Group Relative Policy Optimization (G…
PrunePath introduces a budget-adaptive structured sparsification framework that efficiently prunes Feed-forward networks in large language models, achieving hardware-friendly sparsity and measurable s…
The paper introduces CFGzip, an offline token space compression technique that significantly reduces the computational overhead of constrained decoding, making complex grammar enforcement feasible at…
TAPS introduces a target-aware prefix selection method that optimizes the trade-off between draft tree acceptance and verification cost, achieving significant speedups in speculative decoding.
Xiaosong Han, Ke Chen, Xindi Dai, Di Liang +6 more
TRACE proposes a novel method to mitigate catastrophic forgetting in continual LLM fine-tuning by identifying and isolating a small, task-specific subset of essential parameters for each task.
The paper proposes Joint Neighborhood Optimization (JNO), a novel knowledge-editing framework that jointly addresses the coupled pressures of desirable knowledge propagation and unintended knowledge l…
Haochun Tang, Yuliang Yan, Jiahua Lu, Huaxiao Liu +1 more
The paper introduces R$^2$A, an adversarial attack that uses suffix optimization to mislead black-box LLM routers into consistently selecting expensive, high-capability models.
This paper proposes a neural network-based approach for combustion engine sound modeling, extending conventional methods with machine learning and stochastic components.
The paper proposes SubFit, a novel compression technique that achieves superior LLM compression by replacing non-contiguous, submodule-level components (Attention and FeedForward) with lightweight res…
This paper presents an efficient algorithm for right-to-left sequence prediction based on a new complexity measure called arithmetic repetition complexity, and demonstrates its application to predicti…
This paper introduces HIJACKKV, the first attack framework to exploit the new threat of KV Cache Hijacking in large language models, achieving an average success rate of 94%.
The paper introduces two library combinators, prune and defrel/bank, to bring observational deduplication and bottom-up enumeration into the relational setting using plain miniKanren.
The paper presents an adaptive scaling algorithm with a competitive ratio of 1.373 for incremental submodular maximization under increasing cardinality constraint, improving upon the previous best res…
NEMESIS is a framework that uses large language models to generate accurate performance equations for operational transconductance amplifiers (OTAs) with a balance between speed and accuracy.
A new method for selecting knots in Generalized Additive Models using an extension of adaptive splines and a customized Fellner-Schall scheme.
The paper proposes FOAM, an adaptive damping method that stabilizes the Shampoo optimization algorithm by dynamically controlling damping and eigendecomposition frequency, thereby reducing staleness-i…