Ting Gao
6 indexed papers
Publications per year
Top categories
Frequent co-authors
Research Timeline
The paper introduces FHE-DiCSNN, a novel framework that uses the TFHE scheme to enable secure and efficient computation on Spiking Neural Networks (SNNs), achieving high accuracy and fast inference times.
VCap introduces a novel Witness-Adjudicator reward mechanism that provides highly precise, factually grounded feedback for visual captioning, enabling state-of-the-art performance in RL-trained multimodal models.
ROVER is a lightweight, learnable plugin that efficiently routes and integrates object-centric visual evidence across multiple images and objects, significantly improving performance on grounded multi-image reasoning tasks.
The paper introduces COMET, a novel PLS-SVD framework, to analyze the audio-text modality gap in CLAP models, showing that shared concepts are captured by a small subset of axes, and proposes a spectral truncation method to mitigate this gap for improved zero-shot performance.
The paper proposes OneReason, a framework that enhances the reasoning capability of generative recommendation models by focusing on improving item perception and structuring user behavior into coherent latent interests.
This paper proposes a hybrid two-stage diffusion transformer architecture for instruction-guided audio editing, balancing performance and efficiency.
Papers
Hybrid Diffusion Transformer for Instruction-Guided Audio Editing via Rectified Flow
Liting Gao, Yonggang Zhu, Yaru Chen, Dongyu Wang +4 more
This paper proposes a hybrid two-stage diffusion transformer architecture for instruction-guided audio editing, balancing performance and efficiency.