Xin Li
37 indexed papers
Publications per year
Top categories
Frequent co-authors
Research Timeline
The paper proposes the Morlet Spectral Transformer (MST), a novel architecture that effectively decodes cross-subject emotion from EEG by designing specialized spectral and spatial representations, outperforming existing large foundation models.
The paper introduces Moment-Video, a new benchmark that diagnoses the ability of video MLLMs to understand brief, critical visual events, revealing that current models struggle significantly with temporal fidelity.
InfoMerge is a novel, training-free method that significantly compresses visual tokens for Video-LLMs by estimating temporal redundancy and allocating tokens based on content richness, achieving high efficiency with minimal performance loss.
The paper introduces SPADE-Bench, a new benchmark designed to rigorously evaluate 'agent deception'—the divergence between an agent's reported plan and its actual executed actions—which is a critical safety issue for autonomous LLM agents.
The paper proposes Credit-Attenuated Privileged Feedback (CAPF), a training-time mechanism that uses verifier-side information to guide LLM search agents, significantly improving their performance on complex QA tasks.
The paper proposes Joint Neighborhood Optimization (JNO), a novel knowledge-editing framework that jointly addresses the coupled pressures of desirable knowledge propagation and unintended knowledge leakage during single-edit updates in LLMs.
The JAMEL framework addresses the challenge of effective exploration in open-ended environments by jointly training agent memory and exploration policies using natural, novelty-driven signals.
The paper analyzes information-sharing mechanisms in oligopolies, finding that privacy protection alone is insufficient to incentivize suppliers to share data; successful sharing requires combining privacy safeguards with a sufficiently informative external signal.
QUBRIC introduces a co-design framework that simultaneously optimizes queries and rubrics, overcoming the bottleneck of vague rubrics derived from open-ended questions, leading to significant gains in RL performance.
MLEvolve is a novel self-evolving multi-agent framework that enables LLM agents to discover and optimize machine learning algorithms for complex, long-horizon tasks.
The paper proposes OneReason, a framework that enhances the reasoning capability of generative recommendation models by focusing on improving item perception and structuring user behavior into coherent latent interests.
This paper introduces CORE-Bench, a comprehensive benchmark for code retrieval in agentic coding.
The paper introduces COSM, a cooperative scheduling framework to facilitate concurrent operation of Processing-in-Memory (PIM) and CPU tasks on mobile platforms, improving PIM throughput by up to 2.8x with less than 2.0% CPU performance loss.
This paper identifies the root cause of performance degradation in full-duplex Spoken Language Models (SLMs) due to modality interference and proposes Lychee-FD, a framework that decouples conflicting modalities in deep layers while preserving cross-modality coherence.
This paper proposes ReChannel, a method for dense prediction using a pretrained DiT model, which keeps the encoder but removes the decoder and adapts it with task LoRA. ReChannel maps each token to its corresponding pixel-space patch through a shared linear head.
This paper introduces MonoIR-RS, a large-scale infrared remote-sensing vision-language dataset and benchmark for understanding infrared imagery.
This paper proposes an efficient human intervention mechanism, Deep Interaction, for correcting reasoning errors in large language models, achieving over 25% improvement in correction success rate and reducing token usage by approximately 40%.
A cloud-scale gateway system for MCP services is presented, which breaks the direct-connect model and offloads legacy service integration, consolidates incompatible MCP variants, and reduces tool selection time and token usage.
This paper proposes an end-to-end Markov framework for auditory attention decoding using conditional random fields and an EEG--speech correlation backbone.
This paper conducts an empirical study on the effects of router-side injection in coding agents and evaluates the effectiveness of existing client-side safeguards.
Papers
Where Is the Cost of Third-Party API Routers in Agentic Software Development?
This paper conducts an empirical study on the effects of router-side injection in coding agents and evaluates the effectiveness of existing client-side safeguards.