Yi Li
50 indexed papers
Publications per year
Top categories
Frequent co-authors
Research Timeline
The paper introduces a distribution-free statistical framework that allows existing rewrite-based detectors to achieve finite-sample False Discovery Rate (FDR) guarantees for detecting LLM-generated text without requiring model retraining.
The paper proposes PrefixMem, a dedicated encoder for Semantic IDs (SIDs), demonstrating that structured, prefix-conditioned representations significantly improve the accuracy and recall of generative recommendation systems.
LongTraceRL addresses long-context reasoning challenges by generating highly challenging training data and introducing a fine-grained rubric reward, significantly improving evidence-grounded reasoning in LLMs.
The paper introduces PIGMENT, a physics-informed foundation model that enables reliable quantitative mapping of brain microstructure from extremely sparse or challenging diffusion MRI scans.
The paper proposes a novel, explicitly exploratory iterative Nash Learning from Human Feedback (NLHF) algorithm that achieves strong regret bounds for optimizing LLMs based on complex, non-scalar human preferences.
The paper proposes DAG-MoE, a novel sparse Mixture-of-Experts framework that replaces standard weighted-sum aggregation with structural aggregation to enhance model performance and enable multi-step reasoning.
The paper establishes that finding approximate Hylland-Zeckhauser equilibria (a type of market allocation) is computationally hard, specifically showing it is PPAD-hard under certain complexity assumptions.
OmniOPD introduces a logit-free, chunk-level distillation framework that improves on standard On-Policy Distillation by using semantic similarity and peak-entropy scheduling, achieving state-of-the-art performance even with black-box teachers.
The paper reframes Parameter-Efficient Fine-Tuning (PEFT) from a mere cost-saving alternative to a robust architecture for creating persistent, personalized models that layer specific behaviors onto large shared foundation models.
TrafficRAG is a multimodal retrieval-augmented framework that automates traffic accident liability determination by integrating visual evidence, structured legal knowledge, and advanced LLM reasoning.
The paper introduces SkillHarm, a comprehensive benchmark and automated framework for evaluating skill-based attacks across the entire agent skill-use lifecycle, demonstrating that current agents remain highly vulnerable to both fixed-payload and self-mutating poisoning attacks.
The paper introduces ChartWalker, a framework for generating challenging cross-modal analytical tasks using charts, with a hierarchical knowledge graph construction method and structure-aware sampling algorithm.
This paper develops a new algorithm for policy optimization in online episodic tabular Markov decision processes with unknown transition kernels, providing data-dependent regret bounds and best-of-both-worlds guarantees.
This paper presents an empirical study on how developers respond to agentic code reviews using CodeRabbit, revealing mixed reception and opportunities for improvement.
This paper proposes TrapHunter, a framework to identify deceptive trap tokens in standardized contracts using intent deviation analysis and dynamic validation.
This paper introduces UniRank, an open benchmark for comparing and studying unified ranking models that combine sequential modeling and feature interaction.
This paper introduces ElasticTTT, a framework to preserve the generative prior in Test-Time Tuning (TTT) of pretrained diffusion models for video editing, preventing Prior Collapse.
The authors propose a new solution for the content cold-start problem in industry-scale search and recommender systems, reducing bias, improving model prediction, and validating long-term impact.
The paper introduces LatentFlow, a system for analyzing latent spaces in molecular graph neural networks using clustering and visualization.
This paper presents DOPS, a hardware-aware framework for optimizing operator scheduling and weight layouts in Large Language Models, achieving significant speedups over prefill-decode disaggregation.
Papers
Beyond Prefill-Decode Disaggregation: Dissecting LLM Inference for Heterogeneous Platforms via Dynamic Operator Scheduling
Jiaqi Yang, Jiayi Li, Yihan Fu, Hongxiao Zhao +4 more
This paper presents DOPS, a hardware-aware framework for optimizing operator scheduling and weight layouts in Large Language Models, achieving significant speedups over prefill-decode disaggregation.