Qiang Hu
4 indexed papers
Publications per year
Top categories
Frequent co-authors
Research Timeline
E-MIA introduces a novel, stealthy black-box membership inference attack that converts verifiable hard evidence within a candidate document into an objective, multi-part exam score to determine if the document was ingested into a RAG knowledge base.
The paper addresses the gap in understanding real-world LLM-in-the-loop vulnerabilities by creating the LLMCVE dataset and demonstrating that these vulnerabilities are significantly harder to repair than conventional software flaws.
The paper introduces LL-Bench, a comprehensive benchmark for evaluating large-scale generative models on low-level vision tasks, and proposes LL-Score, an MLLM-based evaluator that better aligns quality assessment with human preferences.
This paper introduces AGENTMETER, a benchmark for evaluating model-CLI matching in local task-solving agents, and the AgentMeter Score (AMS) metric.
Papers
AgentMeter: Evaluating Model-CLI Matching for CLI-Based Local Task-Solving Agents
Han Chi, Jiaxin Qi, Yan Cui, Baisheng Lai +1 more
This paper introduces AGENTMETER, a benchmark for evaluating model-CLI matching in local task-solving agents, and the AgentMeter Score (AMS) metric.