Hiskias Dingeto
3 indexed papers
Publications per year
Top categories
Frequent co-authors
Research Timeline
The paper introduces AGENTREDBENCH, a dynamic redteaming benchmark that significantly measures indirect prompt injection threats in LLM agents using SaaS integrations, and releases AGENTREDGUARD, a superior defense model.
The paper introduces AGENTREDBENCH, a dynamic redteaming benchmark that significantly measures indirect prompt injection threats in LLM agents using third-party integrations, and releases AGENTREDGUARD, a superior defense model.
This paper examines the faithfulness of natural language explanations for hidden activations in neural networks using a reconstruction-based test, and finds that the test is not faithful and can be gamed.
Papers
Train the Model, Not the Reader: Decodability Supervision for Verifiable Activation Explanations
This paper examines the faithfulness of natural language explanations for hidden activations in neural networks using a reconstruction-based test, and finds that the test is not faithful and can be ga…