Youyong Kong
3 indexed papers
Publications per year
Top categories
Frequent co-authors
Research Timeline
The paper introduces ForeSci, a novel benchmark that evaluates LLM agents' ability to make forward-looking research judgments using only historical evidence, finding that explicit evidence organization improves performance but agents often decouple evidence from correct predictions.
This paper proposes BrainAgent, an agentic LLM framework for knowledge-enhanced brain network analysis, which improves performance and produces more comprehensive, multi-level, and verifiable explanations.
The paper introduces E-Bench, a synthetic benchmark for evaluating multi-step tool use in Large Language Models across three product domains.
Papers
E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios
Weihuang Zheng, Tianyuan Zou, Eileen Ye, Alphet Liu +4 more
The paper introduces E-Bench, a synthetic benchmark for evaluating multi-step tool use in Large Language Models across three product domains.