Zhe Zhao
10 indexed papers
Publications per year
Top categories
Frequent co-authors
Research Timeline
MANA introduces an agentic multimodal reasoning framework that significantly improves the efficiency and accuracy of detecting mobile advertisements by integrating multiple types of signals into a guided UI navigation strategy.
FedMPT introduces a novel federated learning framework for Multi-Label Recognition (MLR) using Vision-Language Models (VLMs) by leveraging generalizable conditions to mitigate label overfitting and improve robustness.
The paper introduces Crafter, a multi-agent harness that significantly improves the generation of editable, publication-quality scientific figures from diverse inputs, addressing the limitations of existing single-purpose systems.
The paper introduces RouteGuard, a router-expert framework, to improve the robustness and generalization of safety guardrails by specializing threat detection across multiple distinct unsafe categories.
The paper introduces RouteGuard, a router-expert framework, to improve the robustness and generalization of safety guardrails by specializing threat detection across multiple unsafe categories.
The paper introduces Science Earth, a planet-scale scientific runtime that enables diverse, siloed AI capabilities to connect and collaborate dynamically, demonstrating that scientific discovery can become a distributed, self-correcting process.
The paper introduces Grounded Agentic Interaction Synthesis (GAIS), a framework that generates high-quality, diverse, and complex agentic training data by anchoring tasks to real-world protocols, significantly improving base model performance.
The paper introduces State-Adaptive Prompt Optimization (SAPO), a novel training strategy that treats prompts as dynamic variables to achieve robust fine-tuning, significantly mitigating catastrophic forgetting and improving generalization in LLMs.
This paper presents a comprehensive survey on reconfigurable antennas for next-generation mobile networks, focusing on their potential and applications.
Proposed G$^3$VLA, a camera-aware geometric module for pretrained vision-language-action models, injecting calibrated structure into visual tokens using intrinsic-conditioned ray embeddings, projective positional encoding, and bidirectional cross-view fusion.
Papers
G$^3$VLA: Geometric inductive bias for Vision-Language-Action Models
Yue Peng, Yongzhe Zhao, Artur Habuda, Khuyen Pham +4 more
Proposed G$^3$VLA, a camera-aware geometric module for pretrained vision-language-action models, injecting calibrated structure into visual tokens using intrinsic-conditioned ray embeddings, projectiv…