Abstract
As large language models (LLMs) become integrated into physical robots, they introduce semantic and situational jailbreaks that translate directly into real-world physical harm. Existing safety guardrails evaluate inputs in isolation, leaving embodied agents susceptible to multi-turn conversational social engineering attacks that incrementally smuggle malicious intent past static defences. We present a closed-loop adversarial evaluation and defence architecture for household robots in the AI2-THOR simulator. To expose these vulnerabilities we deploy an adaptive, LLM-driven adversarial agent that orchestrates multi-turn attack strategies such as role-play framing and crescendo escalation. To neutralise them we architect a constitutional multi-turn safety gate, a universal pre-planning interceptor that audits cumulative user intent before any action executes, monitored through a dashboard and an evaluator interface that automates adversarial campaigns. We ablate a memory-less single-turn constitution against a context-aware multi-turn constitution across four safety models, 144 attack episodes each. Single-turn defences leave an Attack Success Rate (ASR) as high as 8.33%. For Gemma-4-26B-A4B, the multi-turn constitution drops ASR from 3.47% to 0.69%. The same rules move Gemma-4-E2B the other way, from 6.25% to 9.03%. We report Wilson and bootstrap intervals for every cell, and because no single per-model change reaches significance at 144 episodes, we offer the split between the sparse mixture-of-experts model and the dense Per-Layer-Embedding edge models as a hypothesis rather than a scaling law. Every attack that survives our strongest configuration splits its payload across turns, prepositioning a physical hazard that the gate never reconciles with the action it is later asked to approve.
| Original language | English |
|---|---|
| Publication status | Published - 10 Aug 2026 |
| Event | 1st IJCAI Workshop on Safe Physical AI 2026 - Bremen, Germany Duration: 15 Aug 2026 → 16 Aug 2026 https://safephysicalai.github.io/ |
Conference
| Conference | 1st IJCAI Workshop on Safe Physical AI 2026 |
|---|---|
| Abbreviated title | IJCAI/ECAI 2026 |
| Country/Territory | Germany |
| City | Bremen |
| Period | 15/08/26 → 16/08/26 |
| Internet address |
Keywords
- AI-Driven Method
- Safety
- Embodied AI
- Large Language Models
- Ai Safety
- Jailbraek Attacks
- Adversarial Red Teaming
- Multi-Turn Interactions
- Constitutional AI
Fingerprint
Dive into the research topics of 'Defending Embodied AI Against Multi-Turn Adversarial Attacks via Constitutional Safety'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver