Hao Wang
50 indexed papers
Publications per year
Top categories
Frequent co-authors
Research Timeline
SafeSteer proposes a localized on-policy distillation method that restricts safety alignment to specific safety tokens, thereby achieving strong safety performance with minimal degradation to general capabilities and significantly reducing data requirements.
The paper introduces MCP-Persona, a novel benchmark designed to evaluate LLM agents' performance on real-world, personalized applications using the Model Context Protocol (MCP), revealing that current state-of-the-art agents struggle with such personalized tool use.
The paper proposes OneReason, a framework that enhances the reasoning capability of generative recommendation models by focusing on improving item perception and structuring user behavior into coherent latent interests.
This paper addresses the challenging problem of multi-objective submodular maximization under a cardinality constraint while ensuring differential privacy, proposing novel algorithms with approximation guarantees.
This paper introduces PASS (Process Advantage Signal Shaping), a method to address three pathologies in GRPO (Group Relative Policy Optimization) for process-supervised reinforcement learning of LLM reasoners.
This paper introduces Active Task Driving Memory (ATMem), an actively maintained execution state for mobile GUI agents, and STR-GRPO, an online reinforcement learning method that uses ATMem selectively.
This paper proposes ERA, a framework for efficient multimodal large language models using entropy-guided visual token pruning, rectified attention, and bias-aware token recycling.
This paper introduces Flex-Forcing, a framework for video generation that enables a model to operate under both bidirectional and autoregressive generation regimes, achieving better video quality and faster inference than existing methods.
This paper proposes LBR, a framework to mitigate length bias in large language model-based recommendation systems.
The paper introduces LingBot-World 2.0, an advanced version of a language model with unbounded interaction horizon, rapid response time, diverse interactive elements, and agentic harness integration.
This paper proposes Evolutionary Intelligence (EI) for scientific discovery, which links candidate refinement with experience retention across evolutionary cycles.
This paper introduces BadWAM, a framework for modeling and evaluating World-Action Drift Attacks, a new class of adversarial attacks that break the alignment between a World-Action Model's (WAM's) imagined future and its executed actions.
The paper introduces CRAFT, a method for converting evaluation datasets into model-specific diagnoses of weak capabilities, achieving stronger results than baselines on four open source models and two professional domains.
This paper proposes a recovery routing system for coding agents that uses a supervised router and Conformal Risk Control (CRC) layer to determine when to spend more compute or escalate to a stronger model.
A new framework called HOST enables robots to acquire new skills from a single human video in seconds while retaining previously mastered skills.
This paper introduces Byte-Prefix Marginalization (BPM), a method for consolidating open-weight language models through on-policy distillation while preserving teacher probability mass and maintaining a shared byte space.
This paper proposes Twins, a unified continuous token space for multimodal models using ViT and VAE features, and addresses optimization imbalance with a focal regression objective.
This paper presents SpecBox, a runtime system for LLM agents that uses speculative sandbox preallocation to improve resource utilization and reduce interactive tail latency.
The paper presents an algorithm for updating a Directed Minimum Spanning Tree using the weighted matroid intersection algorithm and a dynamic auxiliary graph.
This paper provides the first algorithmic separation between constant-depth and logarithmic-depth networks, identifying a class of Boolean functions that logarithmic-depth networks can learn efficiently and exhibiting a subclass for which constant-depth networks incur constant approximation error.
Papers
Optimization of the directed spanning trees using the weighted matroid intersection algorithm
The paper presents an algorithm for updating a Directed Minimum Spanning Tree using the weighted matroid intersection algorithm and a dynamic auxiliary graph.