Built with and by Teycir Ben Soltane•
How to Use•FAQ•GitHub•arXiv.org•
Share:
ArXivCSExplorer
☆☆Bookmarks🏆RSSHow to UseFAQ
Home/Authors/Hiskias Dingeto

Hiskias Dingeto

3 indexed papers

Recent (6 mo)
3
With code
0
Influential cites
0
Benchmarked
0

Publications per year

3
26

Top categories

AI×3NLP×3Crypto×2Emerging Tech×2

Frequent co-authors

Will Leeney1×
William Leeney1×

Research Timeline

2026
AgentRedBench: Dynamic Redteaming and Integration-Aware Defense for LLM Agents over SaaS Integrations

The paper introduces AGENTREDBENCH, a dynamic redteaming benchmark that significantly measures indirect prompt injection threats in LLM agents using SaaS integrations, and releases AGENTREDGUARD, a superior defense model.

AgentRedBench: Dynamic Redteaming and Integration-Aware Defense for LLM Agents over SaaS Integrations

The paper introduces AGENTREDBENCH, a dynamic redteaming benchmark that significantly measures indirect prompt injection threats in LLM agents using third-party integrations, and releases AGENTREDGUARD, a superior defense model.

Train the Model, Not the Reader: Decodability Supervision for Verifiable Activation Explanations

This paper examines the faithfulness of natural language explanations for hidden activations in neural networks using a reconstruction-based test, and finds that the test is not faithful and can be gamed.

Highlighted terms show continued research focus across papers

Papers

cs.AIcs.CLEmpiricalRecentJul 22, 2026

Train the Model, Not the Reader: Decodability Supervision for Verifiable Activation Explanations

Hiskias Dingeto

This paper examines the faithfulness of natural language explanations for hidden activations in neural networks using a reconstruction-based test, and finds that the test is not faithful and can be ga…

View →
cs.CRcs.AIcs.CLRecent
Jun 1, 2026

AgentRedBench: Dynamic Redteaming and Integration-Aware Defense for LLM Agents over SaaS Integrations

Hiskias Dingeto, Will Leeney

The paper introduces AGENTREDBENCH, a dynamic redteaming benchmark that significantly measures indirect prompt injection threats in LLM agents using SaaS integrations, and releases AGENTREDGUARD, a su…

View →
cs.CRcs.AIcs.CLRecentJun 1, 2026

AgentRedBench: Dynamic Redteaming and Integration-Aware Defense for LLM Agents over SaaS Integrations

Hiskias Dingeto, William Leeney

The paper introduces AGENTREDBENCH, a dynamic redteaming benchmark that significantly measures indirect prompt injection threats in LLM agents using third-party integrations, and releases AGENTREDGUAR…

View →