Built with and by Teycir Ben Soltane•
How to Use•FAQ•GitHub•arXiv.org•
Share:
ArXivCSExplorer
☆☆Bookmarks🏆RSSHow to UseFAQ
Back to Paper
cs.CRcs.AIcs.RO

Local ID: 2604.27267v2

AI Summary: gemma4:e4b

From Prompt to Physical Actuation: Holistic Threat Modeling of LLM-Enabled Robotic Systems

By Neha Nagaraja, Hayretdin Bahsi, Carlo R. da Cunha

Revision History Timeline

v14/29/2026
4/29/2026

“Submitted to 23rd Annual International Conference on Privacy, Security, and Trust (PST2026)”

v25/4/2026
5/4/2026

“Submitted to 23rd Annual International Conference on Privacy, Security, and Trust (PST2026)”

★ Version indexed in Explorer

Comparing v1 vs v2

Green = Added • Red = Removed

Title Comparison

From Prompt to Physical Actuation: Holistic Threat Modeling of LLM-Enabled Robotic Systems

Authors Comparison

No author changes.

v1 Comment

“Submitted to 23rd Annual International Conference on Privacy, Security, and Trust (PST2026)”

v2 Comment

“Submitted to 23rd Annual International Conference on Privacy, Security, and Trust (PST2026)”

Abstract Word Diff

As large language models are integrated into autonomous robotic systems for task planning and control, compromised inputs or unsafe model outputs can propagate through the planning pipeline to physical-world consequences. Although prior work has studied robotic cybersecurity, adversarial perception attacks, and LLM safety independently, no existing study traces how these threat categories interact and propagate across trust boundaries in a unified architectural model. We address this gap by modeling an LLM-enabled autonomous robot in an edge-cloud architecture as a hierarchical Data Flow Diagram and applying STRIDE-per-interaction analysis across six boundary-crossing interaction points using a three-category taxonomy of Conventional Cyber Threats, Adversarial Threats, and Conversational Threats. The analysis reveals that these categories converge at the same boundary crossings, and we trace three cross-boundary attack chains from external entry points to unsafe physical actuation, each exposing a distinct architectural property: the absence of independent semantic validation between user input and actuator dispatch, cross-modal translation from visual perception to language-model instruction, and unmediated boundary crossing through provider-side tool use. To our knowledge, this is the first DFD-based threat analysis integrating all three threat categories across the full perception-planning-actuation pipeline of an LLM-enabled robotic system.
View Full Version History on arXiv