Frank Shyu
5 indexed papers
Publications per year
Top categories
Frequent co-authors
Research Timeline
This paper proposes Learning to Allocate (L2A), an end-to-end framework for resource-adaptive inference in Large Language Models (LLMs) using budget-conditioned and input-aware gating networks.
This paper introduces Bifocal dLLMs (R2LM), a new paradigm for discrete diffusion language models that combines causal and bidirectional attention for improved throughput and generation quality.
This paper proposes Diffusion-GR2, a method to convert an autoregressive reasoning re-ranker into a block-diffusion re-ranker while maintaining accuracy and increasing speed.
This paper proposes SCOReD, a framework for optimizing chain-of-thought (CoT) distillation in the recommendation domain by parsing teacher traces into typed segments, scoring their importance, and dynamically selecting edits based on the student's output distribution.
This paper adapts Large Language Models as semantic representation backbones in a two-tower retrieval architecture for high-throughput, large-scale recommendation systems.
Papers
The Case Against Generation for Retrieval: Discriminative Language Models as Effective Retrievers
Zhe Xu, Prachi Agrawal, Kavosh Asadi, Tianyi Chen +16 more
This paper adapts Large Language Models as semantic representation backbones in a two-tower retrieval architecture for high-throughput, large-scale recommendation systems.