Skip to main navigation Skip to search Skip to main content

Defending Embodied AI Against Multi-Turn Adversarial Attacks via Constitutional Safety

Research output: Contribution to conferencePaperpeer-review

Abstract

As large language models (LLMs) become integrated into physical robots, they introduce semantic and situational jailbreaks that translate directly into real-world physical harm. Existing safety guardrails evaluate inputs in isolation, leaving embodied agents susceptible to multi-turn conversational social engineering attacks that incrementally smuggle malicious intent past static defences. We present a closed-loop adversarial evaluation and defence architecture for household robots in the AI2-THOR simulator. To expose these vulnerabilities we deploy an adaptive, LLM-driven adversarial agent that orchestrates multi-turn attack strategies such as role-play framing and crescendo escalation. To neutralise them we architect a constitutional multi-turn safety gate, a universal pre-planning interceptor that audits cumulative user intent before any action executes, monitored through a dashboard and an evaluator interface that automates adversarial campaigns. We ablate a memory-less single-turn constitution against a context-aware multi-turn constitution across four safety models, 144 attack episodes each. Single-turn defences leave an Attack Success Rate (ASR) as high as 8.33%. For Gemma-4-26B-A4B, the multi-turn constitution drops ASR from 3.47% to 0.69%. The same rules move Gemma-4-E2B the other way, from 6.25% to 9.03%. We report Wilson and bootstrap intervals for every cell, and because no single per-model change reaches significance at 144 episodes, we offer the split between the sparse mixture-of-experts model and the dense Per-Layer-Embedding edge models as a hypothesis rather than a scaling law. Every attack that survives our strongest configuration splits its payload across turns, prepositioning a physical hazard that the gate never reconciles with the action it is later asked to approve.
Original languageEnglish
Publication statusPublished - 10 Aug 2026
Event1st IJCAI Workshop on Safe Physical AI 2026 - Bremen, Germany
Duration: 15 Aug 202616 Aug 2026
https://safephysicalai.github.io/

Conference

Conference1st IJCAI Workshop on Safe Physical AI 2026
Abbreviated titleIJCAI/ECAI 2026
Country/TerritoryGermany
CityBremen
Period15/08/2616/08/26
Internet address

Keywords

  • AI-Driven Method
  • Safety
  • Embodied AI
  • Large Language Models
  • Ai Safety
  • Jailbraek Attacks
  • Adversarial Red Teaming
  • Multi-Turn Interactions
  • Constitutional AI

Fingerprint

Dive into the research topics of 'Defending Embodied AI Against Multi-Turn Adversarial Attacks via Constitutional Safety'. Together they form a unique fingerprint.

Cite this