Heng Zhang
39 indexed papers
Publications per year
Top categories
Frequent co-authors
Research Timeline
The paper introduces Harness-Bench, a diagnostic benchmark that measures how different system 'harnesses' affect LLM agent performance in realistic workflows, showing that agent capability must be reported at the model-harness configuration level.
LoSATok proposes a low-dimensional semantic-acoustic tokenizer that efficiently compresses high-dimensional audio features into a compact latent space, significantly improving the performance and efficiency of audio generation models.
The paper introduces AgentDoG 1.5, a lightweight and scalable alignment framework that significantly improves AI agent safety and security for complex, open-world agentic scenarios.
AliMark proposes a novel watermarking framework that treats sentence-level watermarking as a bit sequence alignment problem, significantly enhancing robustness against structural text perturbations like sentence splitting and merging.
The paper introduces DynSess, a novel session-level framework that evaluates and optimizes role-playing agents by assessing long-horizon conversational quality, significantly outperforming existing turn-level methods.
The paper introduces AgentDoG 1.5, a lightweight and scalable alignment framework that significantly improves AI agent safety and security for complex open-world agent deployments.
AliMark proposes a novel framework that enhances the robustness of sentence-level watermarking by reformulating the problem as a bit sequence encoding and alignment task, significantly improving resilience against structural text perturbations like sentence splitting and merging.
HunterAgent is a neuro-symbolic framework that reconstructs causal attack chains from fragmented, anti-forensics-corrupted logs, achieving high accuracy while drastically reducing hallucination.
The paper proposes AsyMoE, a novel Mixture of Experts architecture for Large Vision-Language Models that explicitly models the inherent asymmetry between visual and linguistic modalities, achieving significant performance gains and efficiency improvements.
This paper proposes two horizon-control strategies, Progressive OPD (POPD) and Truncated OPD (TOPD), demonstrating that full rollouts are often unnecessary for On-Policy Distillation, leading to significant improvements in training efficiency.
The paper introduces MineExplorer, a new benchmark in Minecraft, to evaluate the sustained open-world exploration capabilities of MLLM agents, finding that long-horizon coordination remains a significant challenge.
The paper proposes a sequence-alignment framework using Soft Dynamic Time Warping to evaluate audio-driven talking-head generation, demonstrating that this approach provides more robust and fair comparisons than traditional frame-wise metrics.
This paper synthesizes over 150 scattered studies and reports to provide the first comprehensive primer on post-training reasoning data, organizing the field around data objects, utility, construction, and scalability.
The paper introduces Tree-like Self-Play (TSP), a novel framework that treats secure code generation as a fine-grained decision process, significantly improving LLM security by forcing the model to self-correct localized vulnerabilities.
The paper introduces MedPMC, a framework that transforms permissively licensed literature into high-fidelity infrastructure for medical multimodal models, resulting in improved performance on various benchmarks.
The paper proposes WebSwarm, a multi-agent search system that dynamically instantiates agentic search nodes for task decomposition, recursive expansion, and agent collaboration.
The paper introduces FabriVLA, a lightweight Vision-Language-Action model that achieves strong performance on the Meta-World MT50 benchmark using a compact 1B scale VLM backbone and a flow-matching action head.
The paper presents BoxTwin, an interactive digital twin framework that learns the full dynamics of elastoplastic articulated objects from videos and accurately tracks joint trajectories and reproduces post contact plastic behavior.
This paper proposes a test-time scaling approach for Vision-Language Models in Unmanned Aerial Vehicle navigation, enabling self-correction and generation of more accurate and reliable flight plans.
This paper introduces a unified framework for generating high-quality full-length music from lyrics, text descriptions, and musical attributes, consisting of a semantic-aware tokenizer, hybird-LM, FullDiT, and a two-level melody module.
Papers
Pushing the Frontier of Full-Song Generation: Hierarchical Autoregressive Planning Meets Flow-Matching Rendering
Junyu Dai, Xinyue Fan, Weiqin Li, Xiangang Li +12 more
This paper introduces a unified framework for generating high-quality full-length music from lyrics, text descriptions, and musical attributes, consisting of a semantic-aware tokenizer, hybird-LM, Ful…