AgentsAutonomy & tool use
Google Research workshop report proposes context-aware policy checks for AI agents, as a co-author's paper argues those checks have inherent limits
A Google-led workshop report argues that AI agents should check each action against the social norms of its context, using a runtime policy engine. The report is an agenda with no new results, and a co-author's separate paper argues such checks face inherent limits.

The report takes contextual integrity, a framework that judges information flows by sender, recipient, information type and governing norms, and applies it to the appropriateness of actions. It calls this 'contextual security.' In the proposed design, a supervisor layer holds a policy engine. The engine generates policies at runtime, including for tools the agent discovers during a task, and checks each flow before data leaves the user's workspace. Other proposed layers include sandboxing with agent identity, contextual user controls, guardrails against collusion between agents, and governance of shared norms. The report also calls for shared 'Agent Gym' sandboxes and long-running multi-agent benchmarks. It releases no code, benchmark or measurements. [1] [7] [8]
A May 2026 paper by two of the report's authors, Abdelnabi and Bagdasarian, treats prompt injections as violations of contextual integrity. It argues that an attacker can always fabricate a context in which a blocked data flow looks legitimate. That leaves defenders trading off between stopping attacks and blocking legitimate flows. In the same paper, classifiers detected context-based injections at only 0.43–0.59 AUROC, and red-teaming informed by contextual integrity reached a 96.7% attack success rate. [4]
Benchmarks show the problem the report targets. In the academic ConfAIde benchmark, GPT-4 and ChatGPT revealed private information in contexts where humans would not, 39% and 57% of the time respectively. In the PiSAs benchmark of multi-user agents, which a report co-lead co-authored, multi-agent designs leaked less than single agents. Even so, appropriateness violations ranged from more than 25% to 77% depending on configuration, and leaks spread into agents' memory stores. [9] [5]
For practitioners, the reviewed sources point toward enforcement at the execution layer rather than prompt-level or classifier defences alone. Earlier systems followed this pattern. AirGapAgent limits an agent to the user data relevant to its task. CaMeL wraps the model in a capability-based layer and solved 67% of AgentDojo tasks with provable security. A separate 33-expert agenda paper shares the report's worry that frequent confirmation prompts cause fatigue rather than real oversight, and recommends continuous auditing over one-time certification. [2] [3] [6]
benchmarks show LLMs and agents leaking private data where it is contextually inappropriate.
Read the full assessment
Implication: context-aware enforcement at the system layer is a plausible direction, but the report's design is untested and may not withstand adversaries who fabricate context.
Executive brief
Google Research has published a workshop report arguing that AI agents should check each action against the social norms of its context before carrying it out. The most consequential detail is that one of the report's lead authors co-wrote a paper in May 2026 that sets limits on this approach: it argues an attacker can always construct a context in which a blocked data flow looks legitimate. The report came out of Google's CAPS workshop, held in New York in late 2025 with more than 50 participants.
Read the full section
Google Research has published a workshop report arguing that AI agents should check each action against the social norms of its context before carrying it out. In the report's design, a "contextual policy engine" generates those norms at runtime. The most consequential detail is that one of the report's lead authors co-wrote a paper in May 2026 that sets limits on this approach: it argues an attacker can always construct a context in which a blocked data flow looks legitimate. The report came out of Google's CAPS workshop, held in New York in late 2025 with more than 50 participants. It is a research agenda with no new results, and no independent coverage was found.
What changed and event timeline
AirGapAgent applies contextual integrity to agents
Bagdasarian, Gruteser and colleagues showed how to restrict an agent to task-relevant user data. They tested the defence against context-hijacking attacks using Gemini, GPT and Mistral models; the work later appeared at CCS 2024 ().
CaMeL makes the system-layer case
A Google and ETH team placed a capability-based layer around the LLM. It solved 67% of AgentDojo tasks with provable security ().
CAPS workshop in New York
More than 50 academic and industry researchers met to map open problems in agent privacy and security ().
An impossibility result for prompt injection
Abdelnabi and Bagdasarian treat injections as violations of contextual integrity. They argue defenders must trade blocking attacks against blocking legitimate flows ().
The CAPS report is published
Google released the workshop report alongside a blog post by Bagdasarian and Gruteser ().
Capabilities and access
- The report is a position and agenda paper, available as a free PDF. It releases no model, code or benchmark.
- The publication page names Bagdasarian, Tsai, Abdelnabi, Balashankar and more than 50 other contributors.
- The proposed "Agent Gym" environments are a call to build shared sandboxes, not an existing product (blog).
Technical analysis for researchers and developers
- Contextual integrity (CI) defines a norm by who sends and receives information, what type of information it is, and the rules governing the flow.
- A supervisor layer contains a policy engine.
- Proposed layers: dynamic sandboxing with agent identity, model-level reasoning to resolve vague prompts, contextual user controls, guardrails against collusion between agents, and governance of shared norms.
Read the full section
- Framing. Contextual integrity (CI) defines a norm by who sends and receives information, what type of information it is, and the rules governing the flow. The report extends this idea from information sharing to the appropriateness of actions, which it calls "contextual security" (blog; Nissenbaum).
- Architecture. A supervisor layer contains a policy engine. The engine generates policies dynamically, including for tools discovered at runtime, and checks each flow before any data leaves the user's workspace.
- Proposed layers: dynamic sandboxing with agent identity, model-level reasoning to resolve vague prompts, contextual user controls, guardrails against collusion between agents, and governance of shared norms.
- Evaluation. The report calls for dynamic multi-agent benchmarks run over long periods. It reports no empirical results of its own.
Claims and evidence
Independent corroboration of the report's own proposals: none found.
Read the full section
| Claim | Status |
| LLMs make machine-readable, context-dependent policy feasible | Report's position (Google-led), not tested in the report (blog) |
| GPT-4 and ChatGPT leak private information in contexts where humans would not, 39% and 57% of the time | Independent academic benchmark (ConfAIde, ICLR 2024) |
| Classifiers detect context-based injections at only 0.43–0.59 AUROC; CI-informed red-teaming reaches 96.7% attack success | Overlapping authors with the report (arXiv 2605.17634) |
| Multi-agent designs leak less than single agents, but violations stay high | Bagdasarian is a co-author (PiSAs) |
Independent corroboration of the report's own proposals: none found.
Context and prior work
- ConfAIde (Mireshghallah et al.) first tested LLMs' contextual privacy reasoning using CI.
- AirGapAgent and CaMeL both anticipate the report's "system layer around the model" design.
- The independent Grand Challenges paper shares the report's concerns about confirmation fatigue and one-shot consent.
Read the full section
- ConfAIde (Mireshghallah et al.) first tested LLMs' contextual privacy reasoning using CI.
- AirGapAgent and CaMeL both anticipate the report's "system layer around the model" design.
- The independent Grand Challenges paper shares the report's concerns about confirmation fatigue and one-shot consent. However, it judges existing CI frameworks insufficient without new conceptual work for agentic settings.
Limitations, safety and contested findings
- The report offers no implementation, measurements or threat-model evaluation of its policy engine.
- Its own co-lead's impossibility result implies that a norm-checking supervisor cannot be both complete and permissive.
- In PiSAs, appropriateness violations ranged from more than 25% to 77% depending on configuration. Leaks also moved into agents' memory stores.
Read the full section
- The report offers no implementation, measurements or threat-model evaluation of its policy engine.
- Its own co-lead's impossibility result implies that a norm-checking supervisor cannot be both complete and permissive. An adversary can fabricate a context that makes a blocked flow look legitimate.
- In PiSAs, appropriateness violations ranged from more than 25% to 77% depending on configuration. Leaks also moved into agents' memory stores.
- Much of the supporting evidence comes from authors connected to the report or to Google.
Business and practitioner implications
- Don't rely on prompt-level or classifier defences alone.
- Budget for continuous auditing rather than one-time certification, as the Grand Challenges paper recommends.
- Expect frequent per-action confirmation prompts to cause fatigue rather than provide real oversight.
Read the full section
- Don't rely on prompt-level or classifier defences alone. Add enforcement at the execution layer that limits data access per task, following the pattern of CaMeL and AirGapAgent.
- Budget for continuous auditing rather than one-time certification, as the Grand Challenges paper recommends.
- Expect frequent per-action confirmation prompts to cause fatigue rather than provide real oversight.
- For multi-user agents, test shared memory specifically, since PiSAs found violations moving there.
Sources
Read the full section
- Google Research blog: Open and Emergent Problems in Agentic Privacy and Security
- Technical report page · PDF
- AI Agents May Always Fall for Prompt Injections (arXiv 2605.17634)
- PiSAs (arXiv 2607.05318)
- Security and Privacy in Agentic AI: Grand Challenges (arXiv 2607.06608)
- AirGapAgent (arXiv 2405.05175)
- Defeating Prompt Injections by Design (CaMeL) · code
- Can LLMs Keep a Secret? (ConfAIde) · repo
- Nissenbaum, Contextual Integrity Up and Down