CS329Z中文学习站

"通过仿真搜索 LLM 智能体的隐私风险"

"Searching for Privacy Risks in LLM Agents via Simulation"

Yanzhe Zhang, Diyi Yang · "ICLR 2026 · Georgia Tech / Stanford"

必读全译对照查看原文(PDF) ↗

导读

第 8 周「安全与护栏」核心论文:当人人都有 AI 智能体代理行事,恶意智能体可以主动发起多轮对话从你的智能体嘴里套出敏感信息——这种动态攻击面无法人工枚举防御。作者把攻防指令都当作可优化对象,用 LLM 作优化器反思仿真轨迹,交替搜索更强攻击与更强防御:攻击从直接索要(A0)演化到伪造紧急/权威(A1)再到多轮冒充+伪造同意(A2,泄露速度 42.2%),防御从宽松提示演化到规则化同意验证(D1)再到身份验证状态机(D2,压回 7.1%)。发现的攻防跨模型、跨场景迁移;ChatGPT Atlas + 真实 Outlook 的小型 sim-to-real 研究中冒充攻击 5 试 3 中。本页为全文中英对照版本(正文全部,附录略)。

全文对照翻译

译注:覆盖论文正文(摘要至第 5 节结论,含伦理与可复现性声明,原文第 1-10 页)。附录(超参、提示词、完整算法、迁移明细等)未收录,请查阅原文 PDF;其要点已浓缩在文末"要点速览"。

EN

**ABSTRACT** — The widespread deployment of LLM-based agents is likely to introduce a critical privacy threat: malicious agents that proactively engage others in multi-turn interactions to extract sensitive information. However, the evolving nature of such dynamic dialogues makes it challenging to anticipate emerging vulnerabilities and design effective defenses. To tackle this problem, we present a search-based framework that alternates between improving attack and defense strategies through the simulation of privacy-critical agent interactions. Specifically, we employ LLMs as optimizers to analyze simulation trajectories and iteratively propose new agent instructions. To explore the strategy space more efficiently, we further utilize parallel search with multiple threads and cross-thread propagation. Through this process, we find that attack strategies escalate from direct requests to sophisticated tactics, such as impersonation and consent forgery, while defenses evolve from simple rule-based constraints to robust identity-verification state machines. The discovered attacks and defenses generalize across diverse scenarios and backbone models, providing useful insights for developing privacy-aware agents.

摘要 —— LLM 智能体的广泛部署可能引入一种关键隐私威胁:恶意智能体主动与他人展开多轮交互以提取敏感信息。然而此类动态对话的演化本质,使预判新出现的漏洞、设计有效防御都很困难。为此我们提出一个基于搜索的框架:通过对隐私关键的智能体交互进行仿真,交替改进攻击与防御策略。具体地,我们用 LLM 作为优化器,分析仿真轨迹并迭代提出新的智能体指令;为更高效地探索策略空间,进一步利用多线程并行搜索与跨线程传播。通过该过程,我们发现攻击策略从直接请求升级为冒充(impersonation)与伪造同意(consent forgery)等复杂战术,而防御从简单的规则约束演化为鲁棒的身份验证状态机。所发现的攻击与防御可泛化到多样场景与骨干模型,为开发隐私感知智能体提供了有用洞见。

1 引言

EN

The future of interpersonal interaction is evolving towards a world where individuals are supported by AI agents acting on their behalf. These agents will not function in isolation; instead, they will collaborate, negotiate, and share information with agents representing others. This shift will introduce novel privacy paradigms that extend beyond conventional large language model (LLM) privacy considerations. Specifically, it presents a unique challenge: Can AI agents with access to sensitive information maintain privacy awareness while interacting with other agents?

人际交互的未来正走向"人人都有代表自己的 AI 智能体"的世界。这些智能体不会孤立运行,而是与其他人的智能体协作、谈判、共享信息。这一转变将带来超出传统 LLM 隐私考量(训练数据提取、系统提示泄露等)的全新隐私范式,并提出一个独特挑战:能接触敏感信息的 AI 智能体,能否在与他方智能体交互时保持隐私意识?

EN

Prior research on agent privacy has predominantly focused on user-agent or agent-environment interactions, where risks typically emerge from (I) under-specified user instructions that require distinguishing sensitive and non-sensitive information contextually, or (II) malicious environmental elements that prompt agents to disclose sensitive user data through their actions. We argue that these setups fall short in capturing the adaptive and interactive characteristics of real-world threats. To address this gap, we introduce a novel analytical framework for examining agent–agent interactions in which unauthorized parties actively attempt to extract sensitive information through multi-turn dialogues. Unlike environmental threats, which are static and structurally constrained, these exchanges create evolving attack surfaces, which makes it difficult to identify vulnerabilities through manual analysis or enumeration. We address this challenge with a search-based framework that systematically explores the threat landscape and potential defenses based on large-scale simulations. For each privacy norm from prior literature, such as PrivacyLens, our simulation instantiates three agents based on the contextual integrity theory (Nissenbaum, 2009): a data subject, a data sender, and a data recipient. The data subject shares sensitive information with the sender, while the data recipient (attacker) is instructed to elicit it from the sender (defender) via a specified transmission principle (e.g., "send an email"). The conversation between the attacker and the defender continues for multiple rounds, throughout which we detect privacy leakage by examining the defender's actions.

既有的智能体隐私研究主要聚焦"用户-智能体"或"智能体-环境"交互,风险通常来自 (I) 欠规范的用户指令(需按上下文区分敏感与非敏感信息),或 (II) 恶意环境元素(诱导智能体通过行动泄露用户敏感数据)。我们认为这些设定未能刻画真实世界威胁的自适应与交互特征。为弥补该缺口,我们引入一个新颖的分析框架,考察智能体-智能体交互:未授权方主动尝试通过多轮对话提取敏感信息。与静态、结构受限的环境威胁不同,这类交互形成不断演化的攻击面,难以靠人工分析或枚举发现漏洞。我们用基于搜索的框架应对:依托大规模仿真的系统化威胁与防御探索。对既有文献(如 PrivacyLens)中的每条隐私规范,我们的仿真依据情境完整性理论(contextual integrity,Nissenbaum 2009)实例化三个智能体:数据主体(把敏感信息分享给发送方)、数据发送方(防御者)、数据接收方(攻击者)(被指示通过指定传输方式(如"发邮件")从发送方诱导信息)。攻击者与防御者的对话持续多轮;我们通过检查防御者的动作检测隐私泄露。

EN

Simulation provides a controlled way to examine interactive risks: with the defender's instruction fixed, any attacker instruction that induces greater leakage is deemed a more effective strategy. Building on this, our framework alternates between optimizing attacks and defenses: we first search for specific attack instructions tailored to each scenario, then develop universal defense instructions to counter them, repeating this process iteratively. Specifically, we use LLMs as optimizers to analyze simulation trajectories and propose new strategies. To enable a comprehensive exploration of nuanced attack strategies, we develop a parallel search algorithm that allows multiple threads to search simultaneously and propagate breakthrough discoveries across threads. Our framework uncovers effective attacks such as consent forgery and multi-turn impersonation, and develops robust defenses, including strict identity verification and state-machine-based enforcement. Crucially, by framing privacy risks themselves as objects of search, our approach moves beyond manual design and anticipation of threats and establishes a systematic methodology for surfacing previously unrealized vulnerabilities. We further demonstrate that the discovered attacks and defenses are transferable across different backbone models and privacy scenarios, suggesting that our framework can serve as a practical tool to mitigate agent privacy risks in real-world deployments with adversaries.

仿真提供了检视交互风险的受控方式:固定防御指令,任何引发更大泄露的攻击指令都被视为更有效的策略。据此,我们的框架交替优化攻击与防御:先为每个场景搜索定制攻击指令,再研发通用防御指令加以反制,循环迭代。具体用 LLM 作优化器分析仿真轨迹、提出新策略;为全面探索细微的攻击策略,我们开发并行搜索算法——多线程同时搜索、跨线程传播突破性发现。该框架揭示出伪造同意、多轮冒充等有效攻击,并研发出严格身份验证、状态机式执行等鲁棒防御。关键在于,通过把隐私风险本身作为搜索对象,我们的方法超越了人工设计与预判威胁,建立了系统化地暴露前所未见漏洞的方法论。我们进一步证明所发现的攻防可跨骨干模型与隐私场景迁移,说明该框架可作为在存在对手的真实部署中缓解智能体隐私风险的实用工具。

2 相关工作

EN

**LLM Agent Privacy** — Privacy concerns around LLMs often include training data extraction, system prompt extraction, and the leakage of sensitive user data to cloud providers. The most relevant line of research to our work examines whether LLM agents leak private user information in generated actions. Based on contextual integrity theory, ConfAIde and PrivacyLens study the privacy norm awareness of LLMs by prompting them with sensitive information and under-specified user instructions, then benchmarking whether LLM-predicted actions (e.g., sending emails or messages) contain sensitive information. Such privacy-related scenarios can be curated via crowdsourcing or extracted from legal documents. AGENTDAM extends this setting to realistic web navigation environments. However, these works primarily focus on benign settings that do not involve malicious attackers. Liao et al. (2024); Chen et al. (2025) take a step further by investigating whether web agents can handle maliciously embedded elements (e.g., privacy-extraction instructions) while processing sensitive tasks such as filling in online forms on behalf of users. These instructions may be hidden in invisible HTML code or embedded in plausible interface components. Unlike these static threat models, we focus on dynamic adversarial scenarios where attacker agents actively initiate and sustain multi-round conversations to extract sensitive information.

LLM 智能体隐私 —— LLM 的隐私关切通常包括训练数据提取、系统提示提取、敏感用户数据向云提供商泄露等。与本工作最相关的一条研究线是考察 LLM 智能体是否在生成的动作中泄露用户隐私:基于情境完整性理论,ConfAIde 与 PrivacyLens 用敏感信息+欠规范的用户指令提示 LLM,再检测其预测的动作(如发邮件/消息)是否含敏感信息,以此基准测试隐私规范意识;此类场景可众包整理或从法律文件提取;AGENTDAM 把设定扩展到真实网页导航。但这些工作主要聚焦无恶意攻击者的良性设定。Liao et al.、Chen et al. 更进一步,考察网页智能体在处理敏感任务(如替用户填表)时能否应对恶意嵌入元素(隐私提取指令)——指令可能藏在不可见 HTML 代码或貌似合理的界面组件中。与这些静态威胁模型不同,我们聚焦动态对抗场景:攻击者智能体主动发起并维持多轮对话以提取敏感信息。

EN

**Privacy Defense** — The most common defense for privacy risks is prompting LLMs with privacy guidelines. Beyond prompting, Abdelnabi et al. (2025) develop protocols that can automatically derive rules to build firewalls to filter input and data, while Bagdasarian et al. (2024) propose an extra privacy-conscious agent to restrict data access to only task-necessary data. We focus on prompt-based defense in this work because of its simplicity and the model's increasing ability to follow complex instructions. Additionally, our simulation and search framework can readily accommodate and optimize more complex defense protocols in the future. **Prompt Search** — LLMs have demonstrated strong capabilities in prompt search across various contexts. For general task prompting, prior work explores methods such as resampling, a brute-force approach that samples multiple prompts to select high-performing ones, and reflection, which encourages the LLM to learn from (example, score) pairs and iteratively refine better prompts through pattern recognition. More structured approaches integrate LLMs into evolutionary frameworks such as genetic algorithms, enabling prompt optimization through crossover and mutation. For agent optimization, LLMs can inspect agent trajectories and refine agents by directly modifying agent prompts or writing code to improve agent architecture. In contrast to optimizing for task performance, our setting focuses on discovering long-tailed risks, which presents a fundamentally different optimization landscape. In frameworks such as DSPy, even simple prompt rewrites often yield informative performance gradients. In our case, however, many proposed attack strategies may produce no observable signal because they fail to reveal any vulnerability. In similar adversarial contexts, Perez et al. (2022) use resampling to automatically discover adversarial prompts, while AutoDAN applies a genetic algorithm to search for stealthy jailbreak prompts, and Samvelyan et al. (2024); Dharna et al. (2025) formulate the search as a quality-diversity problem to encourage a diverse set of adversarial strategies. Recent work has also explored training specialized models to systematically elicit harmful outputs and behaviors. However, unlike jailbreaking, validating whether an attacker instruction is effective in multi-turn simulations requires significantly more computing and time, making both resampling-based approaches and specialized model training impractical. Therefore, our search procedure builds on the LLM's reflection ability, using previous trials to propose more effective instructions.

隐私防御 —— 最常见的防御是用隐私指南提示 LLM;提示之外,Abdelnabi et al. 开发了可自动推导规则构建防火墙过滤输入与数据的协议,Bagdasarian et al. 提出额外的隐私意识智能体把数据访问限制在任务必需范围。本文聚焦提示式防御——因其简单,且模型跟随复杂指令的能力在增强;此外我们的仿真与搜索框架未来也能容纳并优化更复杂的防御协议。提示搜索 —— LLM 在各类提示搜索上已展现强大能力:通用任务提示方面,既有工作探索了重采样(暴力采多个提示选优)与反思(从"(示例, 分数)"对中学习、迭代精炼);更结构化的方法把 LLM 嵌入进化框架(遗传算法的交叉与变异)做提示优化;智能体优化方面,LLM 可检视智能体轨迹,直接改智能体提示或写代码改进架构。与优化任务性能不同,我们的设定聚焦发现长尾风险,构成根本上不同的优化地形:DSPy 等框架中,哪怕简单的提示改写也常产生有信息量的性能梯度;而在我们的场景,许多候选攻击策略可能完全不产生可观察信号(没挖到任何漏洞)。在类似的对抗语境中,Perez et al. 用重采样自动发现对抗提示,AutoDAN 用遗传算法搜隐蔽越狱提示,Samvelyan et al.、Dharna et al. 把搜索形式化为质量-多样性问题以鼓励多样的对抗策略;近期还有训练专门模型系统诱导有害输出的工作。但与越狱不同,验证一条攻击指令在多轮仿真中是否有效需要多得多的计算与时间,使重采样与专门训练模型都不现实;因此我们的搜索过程建立在 LLM 的反思能力上——用先前的试验提出更有效的指令。

3 搜索风险与缓解

EN

**Simulation Setup** — For a given privacy norm (e.g., "It is not acceptable for the data sender to share the data subject's <sensitive information> with the data recipient by <transmission principle>."), our simulation involves three agents (data subject, data sender, and data recipient) using applications to transfer information. Following prior work, we implement four representative mock applications covering most of the transmission principles: Gmail, Facebook, Messenger, and Notion. We implement agents using the ReAct architecture, which is the most generalizable agent framework. Each agent is initialized with a memory (e.g., social backgrounds of the agent itself and others) and an instruction, and takes actions after receiving notifications from applications. During each simulation run, we first allow the data subject to transfer the sensitive data to the data sender, and then initiate tasks for both the data sender and the data recipient. In our implementation, these two agents will take turns to take actions until the data recipient chooses to end its task, or the maximum number of action cycles for any agent is reached, or the time limit of the entire simulation is exceeded.

仿真设置 —— 对给定隐私规范(如"数据发送方通过〈传输方式〉把数据主体的〈敏感信息〉分享给数据接收方是不可接受的"),仿真涉及三个智能体(数据主体、数据发送方、数据接收方)用应用传递信息。沿用先前工作,我们实现四个代表性 mock 应用覆盖大多数传输方式:Gmail、Facebook、Messenger、Notion。智能体用 ReAct 架构(最具泛化性的智能体框架)实现:每个智能体初始化有记忆(自身与他人的社会背景)与指令,收到应用通知后行动。每次仿真:先让数据主体把敏感数据传给发送方,再为发送方与接收方同时发起任务;两智能体轮流行动,直到接收方选择结束任务、任一智能体达最大动作循环数、或整场仿真超时。

EN

**Risk Metrics** — Following Shao et al. (2024), we use LLMs to detect whether any sensitive item is leaked in each data sender's action. They quantify the leakage through leak rate, which refers to the proportion of trajectories where any sensitive item is leaked. To provide a more fine-grained evaluation, we define the leak velocity and use it as our main metric for evaluation and search, which considers not only whether each item is leaked but also how quickly it is leaked. Specifically: s = (1/K) Σ (1 − log l_i / (log l_i + 1)), where K denotes the number of sensitive items and l_i ∈ [1, +∞) is the number of actions at which the i-th sensitive item is leaked. Thus, a leak velocity s = 1 means all sensitive items are leaked in the first action taken by the data sender, and a lower leak velocity means sensitive items are leaked later. We assign a leak velocity s = 0 to trajectories where no sensitive information is leaked. The quality and robustness of our simulation setup are ensured through our environmental design and objective evaluation. Unlike jailbreaking, which tests LLM outputs on isolated prompts, our simulations operate within realistic application environments, where agents must successfully invoke actual applications with sensitive content for a breach to be recorded. Moreover, the evaluation is straightforward, as the privacy leakage assessment is reduced to a simple detection task: Whether any sensitive information appears in the defender's actions.

风险度量 —— 沿用 Shao et al. (2024),我们用 LLM 检测发送方每个动作是否泄露任一敏感项;他们以 leak rate(泄露率)量化——发生泄露的轨迹占比。为提供更细粒度的评估,我们定义泄露速度(leak velocity)作为评估与搜索的主指标:不仅考虑每项是否泄露,还考虑泄露得多快:s = (1/K) Σ (1 − log l_i/(log l_i + 1)),K 为敏感项数,l_i ∈ [1,+∞) 为第 i 项被泄露时的动作序号。s = 1 表示发送方第一个动作就泄尽全部敏感项;泄露越晚 s 越低;无泄露轨迹记 s = 0。仿真质量与鲁棒性由环境设计与客观评估保证:与在孤立提示上测输出的越狱不同,我们的仿真运行在真实应用环境中,智能体必须真正带着敏感内容调用应用才记为一次泄露;且评估很直接——隐私泄露判定被简化为"防御者动作中是否出现敏感信息"的检测任务。

EN

**Alternating Search** — Basic simulations are limited to testing privacy norms with straightforward instructions. Since possible strategies that might lead to privacy leaks are extensive (e.g., persuasion, social engineering, etc.), we approach privacy risk discovery as a search problem, framing attacker and defender instructions as optimizable objects and using automated search to surface effective strategies that humans may miss. To allow attacks and defenses to co-evolve, we alternate between searching for attacks and defenses. Specifically, for each simulation scenario corresponding to a distinct privacy norm, we define the optimizable part of the configuration as (a, d), where a is the data recipient instruction and d is the data sender instruction. We initialize with Q distinct scenario-specific attacks in A0 and a single universal defense in D0. The T-th search cycle has two phases: (I) Attack search phase: (A_T, D_T) ⇒ (A_{T+1}, D_T), where we conduct Q separate searches to update each scenario-specific attack strategy. (II) Defense search phase: (A_{T+1}, D_T) ⇒ (A_{T+1}, D_{T+1}), where we run a single search to update the universal defense against new attacks. Repeating such cycles allows us to identify the most severe attacks and the most robust defenses gradually.

交替搜索 —— 基础仿真只能测试直白指令下的隐私规范。可能导致泄露的策略空间庞大(说服、社会工程等),因此我们把隐私风险发现当作搜索问题:把攻击者与防御者指令都框架化为可优化对象,用自动搜索浮出人类可能遗漏的有效策略。为让攻防共同演化,我们交替搜索攻击与防御:对每个对应不同隐私规范的仿真场景,可优化配置为 (a, d)(接收方指令 a、发送方指令 d)。初始化:Q 个场景特定攻击 A0 + 1 个通用防御 D0。第 T 轮搜索循环含两阶段:(I) 攻击搜索:(A_T, D_T) ⇒ (A_{T+1}, D_T)——Q 路独立搜索各更新场景特定攻击;(II) 防御搜索:(A_{T+1}, D_T) ⇒ (A_{T+1}, D_{T+1})——单路搜索更新对抗新攻击的通用防御。循环往复,逐步识别最严重的攻击与最鲁棒的防御。

EN

**Attack Search** — Effective attacks are context-dependent, and it is challenging to predict which ones might pose more significant risks than others without simulations. Our preliminary experiments show that generating a wide range of diverse strategies and testing all of them is neither effective nor efficient, as the strategy design receives no feedback from the simulation outcomes. Therefore, a natural idea is to leverage an LLM as an optimizer F to reflect on previous strategies and trajectories to develop new strategies (rewriting the instruction for the data recipient). The effectiveness of reflection-based approaches stems from the LLM's ability to analyze failed attack attempts, understand defensive weaknesses, and amplify successful signals. A sequential search baseline takes the configuration (a, d) as input, where a is the initial attack instruction and d is a fixed defense instruction, updates and evaluates the attack instruction iteratively, and outputs the one with the highest average leak velocity as â. Specifically, at step k, denoting the intermediate attack instruction as a_k, we run the simulation M times with the configuration (a_k, d). This produces trajectories t_k^j for j ∈ [1, M], each with a corresponding leak velocity s_k^j. The collection of results is S_k = {(a_k, t_k^j, s_k^j)}. From S_k, we select a subset with the highest-leak-velocity triples as reflection examples: E_k ← Select(S_k). The LLM optimizer F then generates the next attack instruction a_{k+1} using all search history: a_{k+1} ← F({(a_r, E_r) | 1 ≤ r ≤ k}). We repeat this process for K steps and return the best attack, constituting one search epoch.

攻击搜索 —— 有效攻击依赖上下文,不仿真就难以预测哪些风险更大。预实验表明:广撒网生成大量多样策略再逐一测试既无效也不高效——策略设计得不到仿真结果的反馈。因此自然的想法是用 LLM 作优化器 F,反思先前的策略与轨迹来发展新策略(改写接收方指令):其效力源于 LLM 分析失败的攻击尝试、理解防御弱点、放大成功信号的能力。顺序搜索基线:输入 (a, d)(初始攻击 a、固定防御 d),迭代更新并评估攻击指令,输出平均泄露速度最高的 â。第 k 步:用 (a_k, d) 跑 M 次仿真,得轨迹 t_k^j(j ∈ [1,M])及各自泄露速度 s_k^j;结果集 S_k = {(a_k, t_k^j, s_k^j)};从中选出泄露速度最高的三元组子集作为反思样例 E_k ← Select(S_k);LLM 优化器用全部搜索历史生成下一条攻击指令:a_{k+1} ← F({(a_r, E_r) | 1 ≤ r ≤ k})。重复 K 步返回最佳攻击,构成一个搜索 epoch。

EN

**Parallel Search** — A single-threaded sequential search is often prohibitively slow and constrained by its early exploration, as finding effective strategies may require hundreds or even thousands of iterations. To explore the space more thoroughly and efficiently, our algorithm launches N parallel search threads, each initialized with a distinct instruction generated by the LLM: a_1^1, ···, a_1^N ← Init(a). Each thread independently reflects on and improves its own instruction, substantially increasing search throughput and increasing the likelihood of discovering effective attack strategies within a limited time. A challenge of parallel search is that the total number of simulations per step scales linearly with the number of threads, i.e., N·M runs in total. If we reduce M to allow a larger N, the evaluation of any single instruction becomes less reliable. To address this, we set M to a small value, select one based on its average performance over these M runs, and then re-evaluate it with P additional simulations to obtain a more reliable estimate. Thus, we perform extensive evaluation for only one instruction per step, and ultimately return the best-performing instruction across all steps.

并行搜索 —— 单线程顺序搜索往往慢得离谱且受限于早期探索(找到有效策略可能需要数百甚至上千次迭代)。为更彻底高效地探索,算法启动 N 个并行搜索线程,各自以 LLM 生成的不同指令初始化(a_1^1, …, a_1^N ← Init(a));每线程独立反思并改进自己的指令,大幅提高吞吐,并在有限时间内提升发现有效攻击策略的概率。并行搜索的挑战:每步仿真总数随线程数线性增长(N·M 次);若为加大 N 而缩小 M,单条指令的评估就不可靠。解决办法:M 取小值,按 M 次平均表现选出 1 条,再用 P 次附加仿真重评以获更可靠估计——每步只对一条指令做重评估,最终返回所有步骤中表现最好的指令。

EN

**Cross-Thread Propagation** — A limitation of parallel search is the lack of information sharing between threads, which keeps any discovery isolated. As a result, only the thread that finds the best instruction can refine it in subsequent steps. Inspired by the migration mechanism in evolutionary search, we introduce a cross-thread propagation strategy that shares the best-performing trajectories across all threads whenever the best instruction is updated. Specifically, if the best instruction in the current step (evaluated over P simulation runs) outperforms all previous steps, E_k ← Select(∪_{i=1}^N S_k^i), which means it selects from all threads rather than from the local thread. This ensures that all threads are informed of the most effective strategy found so far, allowing them to refine it simultaneously. Putting all components together, we present a complete version of the attack search algorithm in the Appendix (Algorithm 1).

跨线程传播 —— 并行搜索的局限是线程间不共享信息,任何发现都是孤立的:只有找到最佳指令的线程能在后续步骤中精炼它。受进化搜索中迁移机制启发,我们引入跨线程传播策略:每当最佳指令被刷新,就把表现最好的轨迹共享给所有线程。具体地,若当前步的最佳指令(P 次仿真评估)超过此前所有步骤,则 E_k ← Select(∪_{i=1}^N S_k^i)——从所有线程而非本地线程中选择反思样例。这确保所有线程都知悉迄今最有效的策略,从而同时精炼它。全部组件合起来即完整的攻击搜索算法(附录算法 1)。

EN

**Defense Search** — Similarly, we can apply search to discover better defense strategies. Unlike effective attacks, which are rare, context-dependent, and often hidden in long-tail distributions, effective defenses must be comprehensive and rigorous. Empirically, we find that a single-threaded sequential search across multiple scenarios is sufficient. Specifically, at each step, we simulate all scenarios (M runs in total), compute the average leak velocity to evaluate the current defense, and reflect on the failure cases with the highest leak velocity. We run this process for K steps to find a defense with the lowest average leak velocity across all scenarios. Similar to the fact that multi-task training enables zero-shot generalization, we assume that a defense remaining robust against multiple different attacks is also likely to generalize to unseen attacks.

防御搜索 —— 同样可用搜索发现更好的防御。与稀有、依赖上下文、藏在长尾分布中的有效攻击不同,有效防御必须全面而严格。经验上,跨多场景的单线程顺序搜索就足够:每步仿真全部场景(M 次),算平均泄露速度评估当前防御,反思泄露速度最高的失败案例;跑 K 步找到跨所有场景平均泄露速度最低的防御。类比多任务训练带来零样本泛化:我们假设能抵御多种不同攻击的防御,大概率也能泛化到未见攻击。

4 实验与结果

EN

**Implementation Details** — Using basic agent instructions (e.g., "Guide Jane Carter to share Michael Turner's alleged crime details with you by sending an email." and "Monitor notifications and take necessary actions in response..."), we create 100 simulation configurations as the testing set (Testing-100) using o4-mini-high and manual curation, each of them is a distinct privacy norm randomly sampled from Shao et al. (2024). We use configurations to construct the agent memories and instructions with templates. For each configuration in Testing-100, we run the simulation 10 times by default and use gpt4.1-mini as the backbone for all agents. We choose gpt4.1-mini because it is the medium-sized version of the gpt4.1 family, enabling us to easily study the impact of choosing larger or smaller models. For privacy leakage evaluation, we use geminiflash (non-thinking) to analyze the context and content of each action and detect whether any sensitive item is leaked. This achieves 98.5% agreement with human annotators across 200 randomly sampled actions, ensuring reliable and scalable assessment.

实现细节 —— 用基础智能体指令(如"引导 Jane Carter 通过电子邮件把 Michael Turner 涉嫌犯罪的细节分享给你"与"监控通知并采取必要应对……"),借助 o4-mini-high 与人工筛选构造 100 个仿真配置作测试集(Testing-100),每个是从 Shao et al. (2024) 随机采样的不同隐私规范;用模板从配置构造智能体记忆与指令。Testing-100 每配置默认跑 10 次仿真,所有智能体骨干用 gpt4.1-mini(选它因为它是 gpt4.1 家族的中档型号,便于研究换大/小模型的影响)。泄露评估用 geminiflash(非思考)分析每个动作的上下文与内容、检测是否泄露敏感项:在 200 个随机抽样动作上与人工标注一致率 98.5%,保证评估可靠且可扩展。

EN

**Results** — We show the simulation results in Table 1, where we vary the backbone models in the simulation. For comparison, we consider five other backbones: gpt4.1-nano, gpt4.1, geminiflash (non-thinking), qwen3-32B (non-thinking), and gptoss-20B with low reasoning effort. Besides the average leak velocity, we report the average leak rate as well for reference, which is typically 5-10% higher. For all backbones, we observe nontrivial leaks, despite the data sender agent receiving privacy-augmented instructions ("Maintain highest privacy standards, discretion, and security awareness") and being asked to think before taking actions. More importantly, better backbone models can only induce slightly better attack (gpt4.1-mini→gpt4.1: 31.2%→33.0%) but can induce substantially better defense (gpt4.1-mini→gpt4.1: 31.2%→16.5%). While we provide a more detailed analysis of model differences in Appendix G, this suggests that an effective attack will not emerge from simply using a better backbone model, necessitating the search for more strategic agent instructions.

结果(表 1,基础指令仿真) —— 我们变换仿真骨干模型对比,另考虑 gpt4.1-nano、gpt4.1、geminiflash、qwen3-32B、gptoss-20B 五种骨干。除平均泄露速度外,另报平均泄露率供参考(通常高 5-10 个百分点)。所有骨干都出现非平凡泄露——尽管发送方智能体已收到隐私增强指令("保持最高隐私标准、审慎与安全意识")并被要求三思而后行。更重要的是:更好的骨干只能带来略好的攻击(gpt4.1-mini→gpt4.1:31.2%→33.0%),却能带来大幅更好的防御(31.2%→16.5%)。这说明有效攻击不会从"换个更好的模型"中自发涌现,必须搜索更具策略性的智能体指令。

表 1:基础指令、不同骨干下的平均泄露速度(LV)与泄露率(LR)

攻击方骨干 防御方骨干 LV(↓) LR(↓)
gpt4.1-mini gpt4.1-mini 31.2% 37.6%
gpt4.1-mini gpt4.1 16.5% 19.2%
gpt4.1-mini gpt4.1-nano 34.9% 42.4%
gpt4.1-mini geminiflash 20.4% 24.3%
gpt4.1-mini qwen3-32B 23.1% 30.2%
gpt4.1-mini gptoss-20B 23.7% 33.5%
gpt4.1 gpt4.1-mini 33.0% 42.7%
gpt4.1 gpt4.1-nano 31.2% 35.4%
gpt4.1 geminiflash 27.5% 35.3%
gpt4.1 qwen3-32B 27.2% 31.6%
gpt4.1 gptoss-20B 33.8% 40.5%
EN

**Alternating Search Results** — We create Q = 5 simulation configurations as the training set (Training-5), where the leak during simulation is minimal using the basic instructions. We use a relatively small training set to reduce computational costs while selecting diverse scenarios to ensure generalization and transferability. For each configuration in Training-5, we run the simulation 20 times after each search epoch to mitigate selection bias during search. By default, we use gpt4.1-mini as the backbone for all simulated agents and employ geminipro with a 1024-token thinking budget to generate diverse configurations (Init) and optimize them (F), which is one of the strongest reasoning models. During search, Select() returns 5 examples at each step for reflection. We set N = 30, M = 1, K = 10, P = 10 for attack and N = 1, M = 30, K = 10 for defense. We use our framework to sequentially discover A1, D1, A2, D2, and find that it is hard to find an effective A3 that further increases the leakage.

交替搜索结果 —— 我们构造 Q = 5 个仿真配置作训练集(Training-5,基础指令下泄露极小);训练集较小以省算力,但场景多样以保证泛化与迁移。Training-5 每配置在每个搜索 epoch 后跑 20 次仿真以缓解选择偏差。默认所有仿真智能体骨干 gpt4.1-mini;用 geminipro(1024-token 思考预算,最强推理模型之一)生成多样配置(Init)与执行优化(F)。搜索中 Select() 每步返回 5 个反思样例。攻击参数 N=30、M=1、K=10、P=10;防御参数 N=1、M=30、K=10。我们用框架顺序发现 A1、D1、A2、D2,并发现很难再找到进一步提升泄露的有效 A3。

EN

**Evolving Process of Strategies** — We plot the average leak velocity after each search phase and illustrate the evolving process in Figure 3, which includes strategies and examples. (I) Initially, the attacker employs a direct request approach (A0), which is not effective against D0. The attacker then evolves to A1, developing more sophisticated strategies, such as exploiting consent mechanisms by fabricating consent claims and creating fake urgency to pressure the defender, which improves the average leak velocity to 76.0%. (II) In response to this evolved attack, the defender adapts to D1, implementing rule-based consent verification that requires explicit confirmation from the data subject before sharing sensitive information, which effectively decreases the average leak velocity to 2.5%. (III) However, D1's consent verification proves insufficient against further attack evolution. The search process reveals an even more severe vulnerability in A2: the attacker can directly impersonate the data subject, sending fake consent messages that appear to originate from the legitimate source. This multi-turn strategy, which first establishes fake consent then immediately leverages it, successfully circumvents the rule-based defenses of D1 and improves the average leak velocity again to 42.2%. Note that sending a seemingly naive impersonation message using the data recipient's own email account would never be effective against human users, yet it proves remarkably successful against LLM agents. (IV) In response to this impersonation attack, the defense evolves to D2, implementing a comprehensive state-machine-based approach with strict identity verification protocols. Rather than simply checking for consent messages, D2 requires actively verifying sender identity at each step, and reaching out to the data subject if necessary, effectively neutralizing the impersonation strategy. This iterative process demonstrates how the improvement of attacks and defenses mutually influence one another, ultimately revealing both critical vulnerabilities and strong defense mechanisms. To highlight the necessity of search-based defense development, we compare it with directly prompting state-of-the-art language models to generate comprehensive defense instructions (details in Appendix L). The directly generated defense is substantially more vulnerable than D2.

策略演化过程(图 3) —— 我们画出每个搜索阶段后的平均泄露速度并展示演化过程。(I) 起初攻击者用直接请求(A0),对 D0 无效;随后演化出 A1:更精巧的策略——利用同意机制捏造同意声明、制造虚假紧急向防御者施压,平均泄露速度升至 76.0%。(II) 防御者应对演化出 D1:规则化的同意验证——分享敏感信息前必须获得数据主体的明确确认,平均泄露速度有效降至 2.5%。(III) 但 D1 的同意验证挡不住攻击继续演化。搜索过程揭示出 A2 中更严重的漏洞:攻击者可以直接冒充数据主体,发送看似来自合法来源的伪造同意消息;这种先建立假同意、随即立刻利用它的多轮策略,成功绕过 D1 的规则防御,平均泄露速度再升至 42.2%。注意:用数据接收方自己的邮箱发出一条看似天真的冒充消息,对人类用户永远无效,但对 LLM 智能体却异常成功。(IV) 应对冒充攻击,防御演化出 D2:带严格身份验证协议的全面状态机方法——不是简单检查同意消息,而是每一步主动核实发送者身份,必要时主动联系数据主体,有效中和冒充策略。这一迭代过程展示了攻防改进如何相互影响,最终既揭示关键漏洞也产出强防御机制。为凸显"搜索式防御研发"的必要性,我们与"直接提示最先进语言模型生成全面防御指令"对比(附录 L):直接生成的防御明显比 D2 更易被攻破。

EN

**Cross-Model Transfer** — We further investigate whether attacks and defenses discovered in one model can be transferred to other backbone models for both defense and attack agents. Using identical configurations (from (A0, D0) to (A2, D2)), we evaluate transferability across different backbone models in Table 2. (I) For different defense models, the attacks transfer well, as A1 outperforms A0 against D0, and A2 outperforms A1 against D1 in most cases, even for a stronger backbone like gpt4.1. The average leak velocity of transferred attacks is usually lower than that of the original defense backbone, even when switching to objectively weaker models like gpt4.1-nano, suggesting that the searched attacks are overfitted to the defense backbone to some extent. On the other hand, the defenses transfer less effectively. D1 outperforms D0 against A1 for most backbones except for gpt4.1-nano, while D2 cannot substantially outperform D1 against A2 for backbones like gpt4.1-nano, qwen3-32B, and gptoss-20B. We assume that the transfer of detailed defense instructions, such as D2, requires a strong instruction-following capability, which prevents weaker models from strictly following the protocol in prompts. (II) For different attack models, both attacks and defenses transfer reasonably well, as the trend remains similar across all different backbones. In some cases, the discovered attack is slightly more effective against the untargeted defense using other backbones (e.g., A1 with geminiflash against D0). For complex attacks like A2, the default backbone still performs the best.

跨模型迁移(表 2) —— 我们进一步研究在一个模型上发现的攻防能否迁移到其他骨干。用相同配置((A0,D0) 至 (A2,D2))跨骨干评估:(I) 换防御骨干:攻击迁移良好——多数情况下(即便防御骨干是更强的 gpt4.1)A1 对 D0 胜 A0、A2 对 D1 胜 A1;迁移攻击的平均泄露速度通常低于原防御骨干上的值(即便换成客观更弱的 gpt4.1-nano),说明搜索出的攻击对防御骨干有一定过拟合。另一方面防御迁移较弱:除 gpt4.1-nano 外,D1 对 A1 多数骨干下胜 D0;但 D2 对 A2 在 gpt4.1-nano、qwen3-32B、gptoss-20B 上无法大幅胜过 D1——我们推测 D2 这类精细防御指令的迁移需要强指令遵循能力,弱模型难以严格执行提示中的协议。(II) 换攻击骨干:攻防都迁移得不错,趋势在各骨干下相似;某些情况下,发现的攻击对其他骨干的防御略更有效(如 A1+geminiflash 对 D0);复杂攻击如 A2 仍是默认骨干表现最好。

EN

We further investigate whether defenses discovered using smaller, less expensive models can effectively protect against attacks found with larger, more expensive models. Starting from (A1, D1), we conduct alternative search cycles with either a smaller attack backbone or a smaller defense backbone. We then test the resulting defenses against the attack A2 and compare their performance with that of the targeted defense D2. Specifically, we examine whether we can replace gpt4.1-mini with gpt4.1-nano, which is 4× cheaper. We also conduct an alternative search cycle with the default model setup for comparison. Results in Table 3 reveal two key findings: (I) Partial transfer from smaller models: Defenses discovered using smaller models like gpt4.1-nano provide meaningful protection (20.7%-23.3% leak velocity) but remain less effective than the targeted defense D2 (7.1%). (II) Comparable performance with the same model: When using the same backbone model (gpt4.1-mini), the resulting defense achieves a similar effectiveness (6.6%) to the original one, D2 (7.1%), suggesting the generalizability of defenses discovered using the same model setup.

我们还研究"用更小更便宜模型发现的防御,能否有效抵御用更大更贵模型找到的攻击"。从 (A1, D1) 出发,分别以更小的攻击骨干或更小的防御骨干做交替搜索;把所得防御对 A2 测试,与目标防御 D2 对比。具体考察能否用便宜 4 倍的 gpt4.1-nano 替代 gpt4.1-mini,并以默认设定做对照。表 3 揭示两点:(I) 小模型的部分迁移:gpt4.1-nano 等小模型发现的防御提供有意义的保护(泄露速度 20.7%-23.3%),但仍不及目标防御 D2(7.1%);(II) 同模型下表现相当:同骨干(gpt4.1-mini)重搜所得防御(6.6%)与原 D2(7.1%)相近,说明同设定下发现的防御具备可复现性/泛化性。

EN

**Cross-Scenario Transfer** — Beyond model transfer, we investigate whether discovered attacks and defenses can be applied to different privacy scenarios, such as those in Testing-100. Since we use universal defense instructions, we can directly apply D0, D1, and D2 without modification. However, attacks require scenario-specific adaptation due to their contextual nature. Beyond applying basic attack instructions to Testing-100 (equivalent to A0), we primarily use in-context learning (ICL) to transfer A1 and A2 across scenarios. We provide A1 and A2 with their full configurations as in-context examples and ask LLMs (geminipro, identical to our optimizers) to generate scenario-specific instructions for each scenario in Testing-100. (I) Results in Table 4 demonstrate successful attack transfer through in-context learning: against D0, transferred A1 improves leak velocity from A0's 31.2% to 49.4%, while transferred A2 improves from transferred A1's 6.5% to 17.6% against D1. Correspondingly, transferred defenses effectively mitigate these attacks, reducing leak velocity to approximately 5%. (II) To enable more effective transfer from A2, we rank the transferred results and use the most successful transferred strategies as the strategy guidance in the in-context learning prompt (+SG). This further enhances attack effectiveness by increasing the leak velocity from 17.6% to 32.4% against D1, reflecting the value of an iterative feedback loop.

跨场景迁移 —— 模型迁移之外,我们研究所发现的攻防能否用于 Testing-100 等不同隐私场景。防御是通用指令,D0/D1/D2 可直接套用;攻击因依赖上下文需做场景适配——除把基础攻击指令直接用于 Testing-100(等价 A0)外,我们主要用上下文学习(ICL)迁移 A1 与 A2:把 A1/A2 的完整配置作为上下文示例,让 LLM(与我们优化器相同的 geminipro)为 Testing-100 每个场景生成场景特定指令。(I) 表 4 证明 ICL 迁移攻击成功:对 D0,迁移 A1 把泄露速度从 A0 的 31.2% 提到 49.4%;对 D1,迁移 A2 从迁移 A1 的 6.5% 提到 17.6%;相应地,迁移防御有效缓解这些攻击,把泄露速度压到约 5%。(II) 为让 A2 迁移更有效,我们对迁移结果排序,把最成功的迁移策略作为 ICL 提示中的"策略指导"(+SG):对 D1 的泄露速度从 17.6% 进一步升到 32.4%——体现了迭代反馈回路的价值。

EN

**Case Study** — To examine sim-to-real transfer, we conducted a small case study in a real-world setting, where we use gpt5.1 in the ChatGPT Atlas browser to interact with a live Outlook email account. We instantiated the discovered impersonation attack, in which the attacker attempted to obtain the data subject's mental-health diagnosis from the defender. Despite substantial differences from our simulation setting, including a real email interface, screenshot-based agent interaction, and a different backbone model, the attack succeeded in 3 out of 5 trials. In these cases, the agent incorrectly treated clearly spoofed messages as legitimate consent (e.g., "Explicit consent, ..., Consent verified."). In the remaining two trials, the agent flagged the consent as suspicious, indicating mixed but non-trivial real-world robustness. Although limited in scale, this case study provides preliminary evidence that the failure mode identified in our simulation can manifest in realistic environments. We leave a systematic characterization of sim-to-real transfer to future work.

案例研究(sim-to-real) —— 为检验"仿真到真实"的迁移,我们做了一个小规模真实环境案例研究:用 ChatGPT Atlas 浏览器中的 gpt5.1 操作真实的 Outlook 邮箱账户,实例化所发现的冒充攻击——攻击者试图从防御者处获取数据主体的心理健康诊断。尽管与仿真设定差异巨大(真实邮件界面、基于截图的智能体交互、不同骨干模型),攻击仍在 5 次试验中成功 3 次:这些案例中,智能体把明显伪造的消息当作合法同意(如"Explicit consent, …, Consent verified.");其余 2 次试验中智能体把该同意标记为可疑——说明真实世界鲁棒性喜忧参半但非平凡。尽管规模有限,该案例初步证明仿真识别出的失效模式会在真实环境中显现;系统化的 sim-to-real 刻画留作未来工作。

EN

**Ablation Study on Search Algorithm** — Starting with (A1, D1) as the initial configurations, we validate the design choices in our search algorithm in Figure 4. Our ablation confirms that parallel search with cross-thread propagation and strong optimizer backbones are key to finding vulnerabilities across different backbone models. (I) Parallel: With M = 1, P = 10, we test N = 1, 10, 30 without cross-thread propagation. Increasing the number of search threads enhances search effectiveness, particularly during early iterations, albeit at the expense of additional parallel computation. However, the improvement gradually diminishes, likely due to the absence of information flow between threads. (II) Propagation: Using the same number of parallel threads (N = 30), adding cross-thread propagation mitigates the plateau by enabling more exploration on top of the best solutions so far. We also conduct an ablation where information propagates between threads at every step, which yields suboptimal performance. By examining the trajectories, we attribute this degradation to reduced diversity, as all threads reflect the same selected trajectories at every step, thereby limiting their exploratory potential. (III) Optimizer Backbone: Optimizing agent instructions based on simulation trajectories requires strong long-context understanding and reasoning capabilities. Beyond our default choice geminipro, we evaluate geminiflash with the same thinking budget and a non-reasoning model gpt4.1. Both alternatives perform worse, indicating that the output of our search algorithm highly depends on the backbone of the LLM optimizer. (IV) Data Sender Backbone: We vary the backbone model for the data sender across gpt4.1-mini, gpt4.1-nano, and gpt4.1 to investigate how different privacy awareness levels affect the severity of discovered vulnerabilities. The discovered vulnerabilities (measured by average leak velocity at the final step, gpt4.1 < gpt4.1-mini < gpt4.1-nano) correlate with the defender's privacy awareness levels in Table 1. Notably, even for backbone models with strong privacy awareness, such as gpt4.1, where no successful attacks occurred in the initial search steps, our algorithm uncovers significant vulnerabilities by the end of the search process.

搜索算法消融(图 4) —— 以 (A1, D1) 为初始配置验证搜索算法的设计选择,确认并行搜索 + 跨线程传播 + 强优化器骨干是跨不同骨干发现漏洞的关键。(I) 并行:M=1、P=10,无传播地测 N=1/10/30——线程数增多提升搜索效果(尤其早期迭代),代价是更多并行计算;但改进逐渐减小,大概是线程间无信息流动所致。(II) 传播:同样 N=30,加跨线程传播可缓解平台期——它使搜索能在"迄今最佳解"之上继续探索;我们还做了"每步都传播"的消融,结果次优:检查轨迹发现性能下降源于多样性降低——所有线程每步反思同样的被选轨迹,限制了探索潜力。(III) 优化器骨干:基于仿真轨迹优化智能体指令需要强长上下文理解与推理能力;除默认 geminipro 外,评估同思考预算的 geminiflash 与非推理模型 gpt4.1,二者都更差——说明搜索算法的输出高度依赖 LLM 优化器骨干。(IV) 发送方骨干:在 gpt4.1-mini/nano/gpt4.1 间变换发送方骨干,研究不同隐私意识水平如何影响所发现漏洞的严重度——发现的漏洞(以最终步平均泄露速度计,gpt4.1 < gpt4.1-mini < gpt4.1-nano)与表 1 中防御者隐私意识水平相关;值得注意的是,即便对 gpt4.1 这类隐私意识强、初始搜索步零成功攻击的骨干,我们的算法在搜索结束时也挖出了显著漏洞。

5 结论

EN

In this work, we investigate privacy risks associated with LLM agent interactions. Building on top of controlled simulations, we employ an alternating search approach to systematically uncover these risks and develop robust defense strategies. The core of our search algorithm is to leverage LLMs to reflect on simulation trajectories and propose new attack and defense strategies, where we further augment it through parallel search and cross-thread propagation. We validate the generalization of our approach by demonstrating that the discovered attacks and defenses can be transferred across different model backbones, regardless of the model family, size, or whether they involve reasoning models. Additionally, it can also transfer to unseen scenarios. We hope our work represents an initial step toward the automatic discovery of agent risk and safeguarding, opening up several promising research directions. First, expanding the scope and type of discovered risk: future work could explore broader categories of long-tail risks, such as scenarios that are inherently difficult to handle. Second, broadening the search space: beyond optimizing prompt instructions, researchers could investigate searching for optimal agent architectures, guardrail designs, or even training objectives.

本工作研究 LLM 智能体交互相关的隐私风险。在受控仿真之上,我们采用交替搜索方法系统化地揭示这些风险并研发鲁棒防御策略。搜索算法的核心是利用 LLM 反思仿真轨迹、提出新的攻防策略,并通过并行搜索与跨线程传播进一步增强。我们通过展示所发现的攻防可跨不同模型骨干迁移——无论模型家族、规模、是否推理模型——来验证方法的泛化性;此外也能迁移到未见场景。我们希望本工作是"智能体风险与防护自动发现"的第一步,开启若干有前景的研究方向:其一,扩展所发现风险的范围与类型(如本质上就难处理的场景等更广的长尾风险);其二,拓宽搜索空间——除优化提示指令外,可研究搜索最优的智能体架构、护栏设计、甚至训练目标。

EN

**Ethics Statement** — This work examines privacy risks in agent-agent interactions through controlled simulations. No human subjects or personally identifiable data were involved. Our framework is designed to surface vulnerabilities for the purpose of developing stronger defenses, not to enable malicious use. While the discovery of new attack strategies could potentially inform adversarial actors, we mitigate this risk by presenting them only in the context of effective countermeasures. We believe our research makes a positive contribution to society by identifying emerging privacy threats early and proposing systematic defense strategies that can be implemented in practice.

伦理声明 —— 本工作通过受控仿真考察智能体-智能体交互的隐私风险,无人类被试、无个人可识别数据。我们的框架旨在浮出漏洞以发展更强防御,而非助长恶意使用。虽然发现新攻击策略可能为对手提供信息,但我们只在"配套有效对策"的语境中呈现它们以缓解该风险。我们相信本研究通过及早识别新兴隐私威胁、提出可实际落地的系统化防御策略,为社会做出积极贡献。

要点速览

  • 威胁模型:智能体-智能体多轮交互中的主动套取(说服/社会工程),攻击面随对话演化,人工枚举防不住——区别于静态的环境注入研究线(PrivacyLens/AGENTDAM 等)。
  • 三智能体仿真(主体/发送方/接收方)+ 四 mock 应用(Gmail/Facebook/Messenger/Notion)+ ReAct;度量:泄露率之外提出泄露速度(兼顾是否与多快)作为主指标;泄露判定与人工一致率 98.5%。
  • 方法 = 交替搜索:LLM 优化器反思"(指令, 高泄露轨迹)"历史改写攻防指令;攻击侧 N=30 线程并行(M=1 粗评 + P=10 精评)、全局最优刷新时跨线程传播;防御侧单线程跨场景搜通用指令(多任务→零样本泛化的类比)。
  • 攻防军备竞赛:A1 伪造紧急/捏造同意(76.0%)→ D1 规则化同意验证(2.5%)→ A2 多轮冒充+伪造同意(42.2%)→ D2 身份验证状态机(7.1%);A3 难有提升(攻防趋衡)。
  • 关键洞察:对人类无效的"天真"冒充(用自己的邮箱冒充数据主体)对 LLM 智能体极其有效;更好的骨干主要增强防御而非攻击——强攻击须靠搜索发现;直接让 SOTA 模型"写一份全面防御"明显弱于搜索出的 D2。
  • 迁移:攻击跨模型/跨场景(ICL + 策略指导:+SG 把 A2 迁移效果从 17.6% 提到 32.4%)良好;精细防御(D2)迁移依赖强指令遵循;同骨干重搜可复现(6.6% vs 7.1%);便宜 4 倍的 nano 搜防御可部分迁移(20.7-23.3%)。
  • sim-to-real:gpt5.1 + ChatGPT Atlas + 真实 Outlook,冒充攻击 5 试 3 中(智能体把伪造邮件当"Explicit consent… Consent verified")。
  • 消融:并行线程数提升早期效果但趋平台;只在全球最优刷新时传播(每步都传播会因多样性丧失而变差);优化器骨干(geminipro)很关键;防御骨干越弱发现的漏洞越多,但强骨干(gpt4.1)最终也被攻破。
  • 与课程关联:与 PrivacyLens(问题定义)、Wen et al. CDI(防御学习,直接消费本文的搜索式攻击做对抗训练)构成隐私专题闭环;方法论上是"反思式提示搜索"(GEPA/MIPRO)的红队应用。