Hao Lin
16 indexed papers
Publications per year
Top categories
Frequent co-authors
Research Timeline
The paper proposes SkillProbe, a multi-agent security auditing framework, demonstrating that high-popularity skills in LLM agent marketplaces are often insecure due to systemic combinatorial risks.
The paper analyzes that while multimodal large language models (MLLMs) offer superior semantic understanding for image generation, this enhanced capability significantly increases safety risks, particularly in generating unsafe content and creating harder-to-detect fake images compared to traditional diffusion models.
This paper introduces TwoHamsters, a new benchmark that rigorously tests Multi-Concept Compositional Unsafety (MCCU) in text-to-image models, demonstrating that current state-of-the-art models and safety defenses are highly vulnerable to subtle, compositionally unsafe prompts.
This survey establishes persistent, writable memory as an independent security problem for LLM agents, proposing a comprehensive framework for 'mnemonic sovereignty' to govern the entire memory lifecycle.
The paper introduces OR-Space, a novel full-lifecycle workspace benchmark designed to rigorously evaluate industrial optimization agents by simulating real-world, multi-stage OR workflows that go beyond simple model translation.
MACReD introduces a hierarchical multi-agent framework that achieves state-of-the-art performance in parsing complex chemical reaction diagrams by coordinating specialized agents for perception and global reasoning.
The paper introduces AgentDoG 1.5, a lightweight and scalable alignment framework that significantly improves AI agent safety and security for complex, open-world agentic scenarios.
The paper introduces AgentDoG 1.5, a lightweight and scalable alignment framework that significantly improves AI agent safety and security for complex open-world agent deployments.
DynaTree introduces a two-stage framework that pre-constructs a reusable retrieval tree offline using coordinated agents, allowing for efficient, structure-aware, and highly effective time-sensitive news retrieval online.
SeClaw is a new framework that synthesizes security tasks from structured risk specifications to evaluate autonomous LLM agents' behavior in stateful environments, focusing on the process of unsafe actions rather than just the final outcome.
SeClaw is a new framework that uses specification-driven task synthesis to create comprehensive and controllable security benchmarks for evaluating the unsafe behaviors of autonomous LLM agents.
This paper studies the problem of retrospectively reconstructing commit history from squashed patches, and evaluates different methods using a benchmark dataset.
This paper introduces DynaKRAG, a method for multi-hop retrieval-augmented generation that learns a shared policy for evidence operations, achieving state-of-the-art results on three benchmarks.
This paper introduces GraphBU, a graph-native MILP instance generator whose unit is a local subproblem plus its interface, promoting coupling and preserving feasibility.
Muon improves agentic RL performance in sparse-reward tasks under GiGPO, particularly with specific advantage estimators and learning rates.
This paper investigates the feasibility of large-scale HAPS deployment for Internet connectivity in islands and maritime zones, considering real-world shadowing effects.
Papers
HAPS-enabled Downlink Coverage Enhancement in Islands and Maritime Areas
This paper investigates the feasibility of large-scale HAPS deployment for Internet connectivity in islands and maritime zones, considering real-world shadowing effects.