Ao Zhang
50 indexed papers
Publications per year
Top categories
Frequent co-authors
Research Timeline
This paper introduces DyadEE, a dataset for emotional entrainment detection in conversational interactions, and TRACE, a window-level framework for modeling dyadic interaction using emotion fine-tuned Whisper representations. TRACE achieves the highest accuracy of 97.01% on DyadEE.
This paper proposes ShopX, a model-centric framework for intent-driven shopping experiences using a single foundation model for intent understanding, execution planning, and item-space operations.
The paper introduces PlanRAG, a framework for Retrieval-Augmented Generation (RAG) that models exploratory reasoning problems as logical query trees, addressing representation and optimization gaps between structured SQL and unstructured natural language.
The paper proposes Fuzzy-Function Programming and introduces Program-as-Weights (PAW), a compact, locally-executable neural artifact for everyday programming tasks.
The paper introduces SPEARBench, a benchmark for evaluating naturalness in speech-to-speech language models using a multidimensional protocol.
This paper proposes eBIM, a software-hardware collaborative paradigm for blockchain infrastructure management using RISC-V, and surveys related research and technologies.
The paper presents GIFT, a method for reducing communication volume in large language model pretraining by transforming gradients into a near-isotropic space before quantization.
The paper introduces WaspMOT, a new benchmark for long-term multi-object tracking, and evaluates five tracking-by-detection methods.
This paper introduces LingBot-VA 2.0, a video-action foundation model designed for embodiment, with semantic visual-action tokenization, causal pretraining, sparse MoE backbone, and enhanced asynchronous inference.
This paper proposes SolarChain-Eval, a physics-constrained benchmark for evaluating trustworthy economic agents in decentralized energy markets.
This paper proposes VOP-Nav, a novel navigation system for quadruped robots that combines the geometric safety of Velocity Obstacles with the agile adaptability of end-to-end learning.
JoyNexus is a multi-tenant service for VLA model supervised fine-tuning, reinforcement learning, and evaluation, which decouples services, introduces group batching, and improves training efficiency.
RecGPT-V3 is a stateful, hybrid-modal recommender system that uses a Memory Hub for user memory and a Hybrid-modal Foundation Model for joint reasoning over text tags and Semantic IDs, achieving consistent gains in user experience and commercial outcomes.
The paper proposes SALMONN-2, an ALLM built on a unified SSL encoder, and presents a multi-layer feature fusion adapter to better exploit hierarchical SSL encoder representations. It also explores multimodal in-context learning in ALLMs and shows that a general-purpose SSL encoder achieves comparable performance to specialized audio encoders.
This paper proposes a spatial semantic communication (SSC) system using fluid antenna-index modulation (FA-IM) technology, which synergizes residual quantization (RQ) and IM for efficient semantic transmission.
The paper introduces ReferTrack, a method for embodied visual tracking using a single forward-facing camera, achieving state-of-the-art performance on EVT-Bench.
This paper proposes RPPNet, a two-stage deep learning architecture for music generation with variable structural boundaries, which automatically derives grouping of Rhythm-Pitch Primitive sequences from acoustic cues and musical psychology.
This paper proposes RL-MACRO, a cybernetic closed-loop intelligence framework for autonomous robotic craniotomy, which includes a CNN-LSTM observer for temperature reconstruction, an offline Implicit Q-Learning policy, and a novel dual-head Actor for coordinating cutting parameters.
The paper introduces an agentic soundscape construction framework for controllable compositional audio generation, which makes explicit the scene planning, source selection, temporal layout, and rendering steps.
This paper presents SpecBox, a runtime system for LLM agents that uses speculative sandbox preallocation to improve resource utilization and reduce interactive tail latency.
Papers
SpecBox: Speculative Sandbox Scheduling for Efficient LLM Agent Serving
Yihui Zhang, Tianyu Wo, Jinghao Wang, Xiaoyang Sun +6 more
This paper presents SpecBox, a runtime system for LLM agents that uses speculative sandbox preallocation to improve resource utilization and reduce interactive tail latency.