"ReAct:在语言模型中协同推理与行动"
"ReAct: Synergizing Reasoning and Acting in Language Models"
必读全译对照查看原文(PDF) ↗
导读
ReAct 是第 4 周「Agent 设计模式」的开篇必读,也是几乎所有 LLM 智能体(包括本课程后面每一讲)的祖师爷范式。它的出发点极其直觉:人类做菜时会在动作之间用语言推理("切完了,该烧水了""没盐了,用酱油代替吧"),为什么不让 LLM 也交错地生成推理轨迹(thought)与任务动作(action)?推理帮助模型制定、跟踪、调整行动计划并处理异常;动作让模型与外部环境(维基百科 API、网页、模拟家庭)交互获取证据。结果:在知识密集任务上,ReAct 比 CoT 更事实可靠、幻觉更少(错误模式分析:CoT 56% 的失败源于幻觉,ReAct 为 0%);在交互决策任务上,仅 1-2 个上下文示例就击败用 10³-10⁵ 条轨迹训练的模仿/强化学习方法(绝对提升 34%/10%)。本页为全文中英对照版本:英文原段与中文全译逐段交替,可用页面顶部按钮切换"仅看中文"。
全文对照翻译
译注:以下覆盖论文正文全部内容(摘要、第 1-6 节、致谢、可复现性与伦理声明,原文第 1-10 页)。附录 A-G(GPT-3 附加实验、微调细节、全部任务提示词、样例轨迹与人工分析)未收录,完整提示词请查阅原文 PDF;其要点已浓缩在文末"要点速览"。图 1 的示例轨迹在 PDF 提取中存在编码乱码,以下对照以文字描述补充。
**ABSTRACT** — While large language models (LLMs) have demonstrated impressive performance across tasks in language understanding and interactive decision making, their abilities for reasoning (e.g. chain-of-thought prompting) and acting (e.g. action plan generation) have primarily been studied as separate topics. In this paper, we explore the use of LLMs to generate both reasoning traces and task-specific actions in an interleaved manner, allowing for greater synergy between the two: reasoning traces help the model induce, track, and update action plans as well as handle exceptions, while actions allow it to interface with and gather additional information from external sources such as knowledge bases or environments. We apply our approach, named ReAct, to a diverse set of language and decision making tasks and demonstrate its effectiveness over state-of-the-art baselines in addition to improved human interpretability and trustworthiness. Concretely, on question answering (HotpotQA) and fact verification (Fever), ReAct overcomes prevalent issues of hallucination and error propagation in chain-of-thought reasoning by interacting with a simple Wikipedia API, and generating human-like task-solving trajectories that are more interpretable than baselines without reasoning traces. Furthermore, on two interactive decision making benchmarks (ALFWorld and WebShop), ReAct outperforms imitation and reinforcement learning methods by an absolute success rate of 34% and 10% respectively, while being prompted with only one or two in-context examples.
摘要 —— 大语言模型(LLM)在语言理解与交互决策的众多任务上展现了惊艳的表现,但其推理能力(如思维链提示)与行动能力(如行动计划生成)此前基本是两个被分开研究的课题。本文探索让 LLM 以交错方式同时生成推理轨迹与任务专用动作,使二者产生更强的协同:推理轨迹帮助模型诱导、跟踪与更新行动计划并处理异常;动作则让模型对接知识库或环境等外部来源并收集额外信息。我们把这一名为 ReAct 的方法应用于一组多样的语言与决策任务,证明其优于最强基线,同时改善了人类可解释性与可信度。具体而言,在问答(HotpotQA)与事实核查(Fever)上,ReAct 通过与简单的维基百科 API 交互,克服了思维链推理中普遍的幻觉与错误传播问题,并生成比无推理轨迹基线更可解释的、类人的任务求解轨迹。此外,在两个交互决策基准(ALFWorld 与 WebShop)上,ReAct 仅用一两个上下文示例提示,就分别以 34% 与 10% 的绝对成功率优势超越模仿学习与强化学习方法。
1 引言(Introduction)
A unique feature of human intelligence is the ability to seamlessly combine task-oriented actions with verbal reasoning (or inner speech, Alderson-Day & Fernyhough, 2015), which has been theorized to play an important role in human cognition for enabling self-regulation or strategization (Vygotsky, 1987; Luria, 1965; Fernyhough, 2010) and maintaining a working memory (Baddeley, 1992). Consider the example of cooking up a dish in the kitchen. Between any two specific actions, we may reason in language in order to track progress ("now that everything is cut, I should heat up the pot of water"), to handle exceptions or adjust the plan according to the situation ("I don't have salt, so let me use soy sauce and pepper instead"), and to realize when external information is needed ("how do I prepare dough? Let me search on the Internet"). We may also act (open a cookbook to read the recipe, open the fridge, check ingredients) to support the reasoning and to answer questions ("What dish can I make right now?"). This tight synergy between "acting" and "reasoning" allows humans to learn new tasks quickly and perform robust decision making or reasoning, even under previously unseen circumstances or facing information uncertainties.
人类智能的一个独特之处,是把面向任务的动作与言语推理(或"内心独白",Alderson-Day & Fernyhough, 2015)无缝结合的能力;理论认为它在人类认知中对自我调节与策略制定(Vygotsky, 1987; Luria, 1965; Fernyhough, 2010)以及维持工作记忆(Baddeley, 1992)起重要作用。以在厨房做一道菜为例:在任意两个具体动作之间,我们可能会用语言推理来跟踪进度("现在都切好了,我该烧一锅水了")、处理异常或按情况调整计划("没有盐了,那就用酱油和胡椒代替")、以及意识到何时需要外部信息("面团怎么和?上网搜一下")。我们也会行动(翻开菜谱、打开冰箱、检查食材)来支持推理并回答问题("我现在能做什么菜?")。这种"行动"与"推理"之间的紧密协同,让人类能快速学习新任务,即便在从未见过的情况下或面对信息不确定时,也能进行稳健的决策与推理。
Recent results have hinted at the possibility of combining verbal reasoning with interactive decision making in autonomous systems. On one hand, properly prompted large language models (LLMs) have demonstrated emergent capabilities to carry out several steps of reasoning traces to derive answers from questions in arithmetic, commonsense, and symbolic reasoning tasks (Wei et al., 2022). However, this "chain-of-thought" reasoning is a static black box, in that the model uses its own internal representations to generate thoughts and is not grounded in the external world, which limits its ability to reason reactively or update its knowledge. This can lead to issues like fact hallucination and error propagation over the reasoning process (Figure 1 (1b)). On the other hand, recent work has explored the use of pre-trained language models for planning and acting in interactive environments (Ahn et al., 2022; Nakano et al., 2021; Yao et al., 2020; Huang et al., 2022a), with a focus on predicting actions via language priors. These approaches usually convert multi-modal observations into text, use a language model to generate domain-specific actions or plans, and then use a controller to choose or execute them. However, they do not employ language models to reason abstractly about high-level goals or maintain a working memory to support acting, barring Huang et al. (2022b) who perform a limited form of verbal reasoning to reiterate spatial facts about the current state. Beyond such simple embodied tasks to interact with a few blocks, there have not been studies on how reasoning and acting can be combined in a synergistic manner for general task solving, and if such a combination can bring systematic benefits compared to reasoning or acting alone.
近期结果暗示了在自主系统中把言语推理与交互决策结合的可能性。一方面,经过恰当提示的大语言模型已展现出进行多步推理轨迹的能力,可从算术、常识与符号推理任务的问题中推导答案(Wei et al., 2022)。然而这种"思维链"推理是一个静态黑箱:模型用自身内部表示生成思考、并不接地(grounded)于外部世界,这限制了它反应式地推理或更新知识的能力,可能导致事实幻觉与推理过程中的错误传播(图 1(1b))。另一方面,近期工作探索了用预训练语言模型在交互环境中规划与行动(Ahn et al., 2022; Nakano et al., 2021 等),重点在通过语言先验预测动作:通常把多模态观测转为文本,让语言模型生成领域专用动作或计划,再由控制器选择或执行。但它们不用语言模型对高层目标做抽象推理,也不维护支持行动的工作记忆——除了 Huang et al. (2022b) 做了一种有限形式的言语推理(重述当前状态的空间事实)。在此类与几个积木交互的简单具身任务之外,尚未有研究探讨推理与行动如何以协同方式结合以求解一般任务,以及这种结合相比单独推理或单独行动能否带来系统性收益。
In this work, we present ReAct, a general paradigm to combine reasoning and acting with language models for solving diverse language reasoning and decision making tasks (Figure 1). ReAct prompts LLMs to generate both verbal reasoning traces and actions pertaining to a task in an interleaved manner, which allows the model to perform dynamic reasoning to create, maintain, and adjust high-level plans for acting (reason to act), while also interact with the external environments (e.g. Wikipedia) to incorporate additional information into reasoning (act to reason).
本工作中,我们提出 ReAct——一个用语言模型结合推理与行动的通用范式,用于求解多样的语言推理与决策任务(图 1)。ReAct 提示 LLM 以交错方式生成与任务相关的言语推理轨迹与动作,使模型能够进行动态推理,为行动创建、维护与调整高层计划(以推理指导行动,reason to act);同时与外部环境(如维基百科)交互,把额外信息纳入推理(以行动反哺推理,act to reason)。
We conduct empirical evaluations of ReAct and state-of-the-art baselines on four diverse benchmarks: question answering (HotPotQA, Yang et al., 2018), fact verification (Fever, Thorne et al., 2018), text-based game (ALFWorld, Shridhar et al., 2020b), and webpage navigation (WebShop, Yao et al., 2022). For HotPotQA and Fever, with access to a Wikipedia API that the model can interact with, ReAct outperforms vanilla action generation models while being competitive with chain-of-thought reasoning (CoT) (Wei et al., 2022). The best approach overall is a combination of ReAct and CoT that allows for the use of both internal knowledge and externally obtained information during reasoning. On ALFWorld and WebShop, two or even one-shot ReAct prompting is able to outperform imitation or reinforcement learning methods trained with 10³∼10⁵ task instances, with an absolute improvement of 34% and 10% in success rates respectively. We also demonstrate the importance of sparse, versatile reasoning in decision making by showing consistent advantages over controlled baselines with actions only. Besides general applicability and performance boost, the combination of reasoning and acting also contributes to model interpretability, trustworthiness, and diagnosability across all domains, as humans can readily distinguish information from model's internal knowledge versus external environments, as well as inspect reasoning traces to understand the decision basis of model actions.
我们在四个多样基准上对 ReAct 与最强基线做实证评测:问答(HotPotQA)、事实核查(Fever)、文字游戏(ALFWorld)与网页导航(WebShop)。对 HotPotQA 与 Fever,在可交互的维基百科 API 下,ReAct 胜过朴素动作生成模型,并与思维链推理(CoT)相当;总体最佳方案是 ReAct 与 CoT 的结合,使推理过程既能用内部知识也能用外部获取的信息。在 ALFWorld 与 WebShop 上,两例甚至一例的 ReAct 提示就能胜过用 10³∼10⁵ 条任务实例训练的模仿或强化学习方法,成功率绝对提升分别为 34% 与 10%。我们还通过与"仅动作"受控基线的一致优势,证明了稀疏而多样的推理在决策中的重要性。除了普适与性能提升,推理与行动的结合还改善了模型的可解释性、可信度与可诊断性:人类可以轻易区分信息来自模型内部知识还是外部环境,并检视推理轨迹以理解模型动作的决策依据。
To summarize, our key contributions are the following: (1) we introduce ReAct, a novel prompt-based paradigm to synergize reasoning and acting in language models for general task solving; (2) we perform extensive experiments across diverse benchmarks to showcase the advantage of ReAct in a few-shot learning setup over prior approaches that perform either reasoning or action generation in isolation; (3) we present systematic ablations and analysis to understand the importance of acting in reasoning tasks, and reasoning in interactive tasks; (4) we analyze the limitations of ReAct under the prompting setup (i.e. limited support of reasoning and acting behaviors), and perform initial finetuning experiments showing the potential of ReAct to improve with additional training data. Scaling up ReAct to train and operate on more tasks and combining it with complementary paradigms like reinforcement learning could further unlock the potential of large language models.
总结我们的主要贡献:(1) 提出 ReAct——在语言模型中协同推理与行动以求解一般任务的新颖提示范式;(2) 跨多样基准的大量实验,展示少样本设定下 ReAct 相对"孤立推理或孤立行动"既有方法的优越性;(3) 系统的消融与分析,理解行动之于推理任务、推理之于交互任务的重要性;(4) 分析提示设定下 ReAct 的局限(对推理与行动行为的支持有限),并做初步微调实验展示其随训练数据增长的潜力。把 ReAct 扩展到更多任务上训练与运行,并与强化学习等互补范式结合,有望进一步释放大语言模型的潜力。
2 ReAct:协同推理 + 行动
Consider a general setup of an agent interacting with an environment for task solving. At time step t, an agent receives an observation o_t ∈ O from the environment and takes an action a_t ∈ A following some policy π(a_t|c_t), where c_t = (o_1, a_1, ···, o_{t−1}, a_{t−1}, o_t) is the context to the agent. Learning a policy is challenging when the mapping c_t → a_t is highly implicit and requires extensive computation. For example, the agent shown in Figure 1(1c) is unable to generate the correct final action (Act 4) to finish the QA task as it requires complex reasoning over the trajectory context (Question, Act 1-3, Obs 1-3). Similarly, the agent shown in Figure 1(2a) fails to comprehend from the context that sinkbasin 1 does not contain peppershaker 1, thus keep producing hallucinating actions.
考虑智能体与环境交互求解任务的一般设定。在时间步 t,智能体从环境收到观测 o_t ∈ O,并按某个策略 π(a_t|c_t) 执行动作 a_t ∈ A,其中 c_t = (o_1, a_1, …, o_{t−1}, a_{t−1}, o_t) 是智能体的上下文。当映射 c_t → a_t 高度隐式且需要大量计算时,学习策略是困难的。例如图 1(1c) 的智能体无法生成正确的最终动作(Act 4)来完成问答任务,因为这需要对轨迹上下文(问题、Act 1-3、Obs 1-3)的复杂推理;类似地,图 1(2a) 的智能体无法从上下文中领会到 sinkbasin 1(洗手池)里并没有 peppershaker 1(胡椒罐),于是不断产生幻觉动作。
The idea of ReAct is simple: we augment the agent's action space to  = A ∪ L, where L is the space of language. An action â_t ∈ L in the language space, which we will refer to as a thought or a reasoning trace, does not affect the external environment, thus leading to no observation feedback. Instead, a thought â_t aims to compose useful information by reasoning over the current context c_t, and update the context c_{t+1} = (c_t, â_t) to support future reasoning or acting. As shown in Figure 1, there could be various types of useful thoughts, e.g. decomposing task goals and create action plans (2b, Act 1; 1d, Thought 1), injecting commonsense knowledge relevant to task solving (2b, Act 1), extracting important parts from observations (1d, Thought 2, 4), track progress and transit action plans (2b, Act 8), handle exceptions and adjust action plans (1d, Thought 3), and so on.
ReAct 的想法很简单:把智能体的动作空间增广为 Â = A ∪ L,其中 L 是语言空间。语言空间中的动作 â_t ∈ L——我们称之为思考(thought)或推理轨迹——不影响外部环境,因此不产生观测反馈;思考的目的是基于当前上下文 c_t 推理并组合出有用信息,更新上下文 c_{t+1} = (c_t, â_t) 以支持未来的推理或行动。如图 1 所示,有用的思考有多种类型:分解任务目标并创建行动计划(2b, Act 1;1d, Thought 1)、注入与任务求解相关的常识知识(2b, Act 1)、从观测中提取重要部分(1d, Thought 2、4)、跟踪进度并转换行动计划(2b, Act 8)、处理异常并调整计划(1d, Thought 3)等。
However, as the language space L is unlimited, learning in this augmented action space is difficult and requires strong language priors. In this paper, we mainly focus on the setup where a frozen large language model, PaLM-540B (Chowdhery et al., 2022), is prompted with few-shot in-context examples to generate both domain-specific actions and free-form language thoughts for task solving (Figure 1 (1d), (2b)). Each in-context example is a human trajectory of actions, thoughts, and environment observations to solve a task instance (see Appendix C). For the tasks where reasoning is of primary importance (Figure 1(1)), we alternate the generation of thoughts and actions so that the task-solving trajectory consists of multiple thought-action-observation steps. In contrast, for decision making tasks that potentially involve a large number of actions (Figure 1(2)), thoughts only need to appear sparsely in the most relevant positions of a trajectory, so we let the language model decide the asynchronous occurrence of thoughts and actions for itself.
但由于语言空间 L 无限大,在这个增广动作空间中学习是困难的,需要强大的语言先验。本文主要聚焦的设定是:一个冻结的大语言模型 PaLM-540B,以少样本上下文示例提示,同时生成领域专用动作与自由形式的语言思考来求解任务(图 1(1d)、(2b))。每个上下文示例是一条人工编写的"动作-思考-环境观测"轨迹(见附录 C)。对推理占主导的任务(图 1(1)),我们交替生成思考与动作,使求解轨迹由多个"思考-动作-观测"步组成;相反,对可能涉及大量动作的决策任务(图 1(2)),思考只需稀疏地出现在轨迹最相关的位置,因此我们让语言模型自行决定思考与动作的异步出现。
Since decision making and reasoning capabilities are integrated into a large language model, ReAct enjoys several unique features: A) Intuitive and easy to design: Designing ReAct prompts is straightforward as human annotators just type down their thoughts in language on top of their actions taken. No ad-hoc format choice, thought design, or example selection is used in this paper. We detail prompt design for each task in Sections 3 and 4. B) General and flexible: Due to the flexible thought space and thought-action occurrence format, ReAct works for diverse tasks with distinct action spaces and reasoning needs, including but not limited to QA, fact verification, text game, and web navigation. C) Performant and robust: ReAct shows strong generalization to new task instances while learning solely from one to six in-context examples, consistently outperforming baselines with only reasoning or acting across different domains. We also show in Section 3 additional benefits when finetuning is enabled, and in Section 4 how ReAct performance is robust to prompt selections. D) Human aligned and controllable: ReAct promises an interpretable sequential decision making and reasoning process where humans can easily inspect reasoning and factual correctness. Moreover, humans can also control or correct the agent behavior on the go by thought editing, as shown in Figure 5 in Section 4.
由于决策与推理能力被整合进同一个大语言模型,ReAct 具备若干独特优势:A) 直观易设计:标注者只需在自己的动作之上用语言写下思考即可,本文未做任何特别的格式选择、思考设计或示例挑选;B) 通用灵活:得益于灵活的思考空间与"思考-动作"出现格式,ReAct 适用于动作空间与推理需求迥异的多样任务(问答、事实核查、文字游戏、网页导航等);C) 高性能且稳健:仅从 1-6 个上下文示例学习,就对新任务实例表现出强泛化,在各领域一致超过仅推理或仅行动的基线;第 3 节展示微调的额外收益,第 4 节展示对提示选择的稳健性;D) 人类对齐且可控:ReAct 提供可解释的序贯决策与推理过程,人类可轻松检查推理与事实正确性,还能通过编辑思考实时控制或纠正智能体行为(见第 4 节图 5)。
3 知识密集型推理任务
We begin with knowledge-intensive reasoning tasks like multi-hop question answering and fact verification. As shown in Figure 1(1d), by interacting with a Wikipedia API, ReAct is able to retrieve information to support reasoning, while also use reasoning to target what to retrieve next, demonstrating a synergy of reasoning and acting.
我们先从多跳问答与事实核查等知识密集型推理任务开始。如图 1(1d) 所示,通过与维基百科 API 交互,ReAct 能检索信息支持推理,同时用推理确定下一步检索什么——展示推理与行动的协同。
**Domains** — We consider two datasets challenging knowledge retrieval and reasoning: (1) HotPotQA (Yang et al., 2018), a multi-hop question answering benchmark that requires reasoning over two or more Wikipedia passages, and (2) FEVER (Thorne et al., 2018), a fact verification benchmark where each claim is annotated SUPPORTS, REFUTES, or NOT ENOUGH INFO, based on if there exists a Wikipedia passage to verify the claim. In this work, we operate in a question-only setup for both tasks, where models only receive the question/claim as input without access to support paragraphs, and have to rely on their internal knowledge or retrieve knowledge via interacting with an external environment to support reasoning.
任务域 —— 我们考虑两个挑战知识检索与推理的数据集:(1) HotPotQA,需要基于两篇及以上维基百科段落进行推理的多跳问答基准;(2) FEVER,事实核查基准,每条断言依据是否存在可验证它的维基百科段落,被标注为 SUPPORTS(支持)、REFUTES(驳斥)或 NOT ENOUGH INFO(信息不足)。本文两任务都采用仅问题(question-only)设定:模型只收到问题/断言作为输入、拿不到支持段落,必须依靠内部知识,或通过与外部环境交互检索知识来支撑推理。
**Action Space** — We design a simple Wikipedia web API with three types of actions to support interactive information retrieval: (1) search[entity], which returns the first 5 sentences from the corresponding entity wiki page if it exists, or else suggests top-5 similar entities from the Wikipedia search engine, (2) lookup[string], which would return the next sentence in the page containing string, simulating Ctrl+F functionality on the browser. (3) finish[answer], which would finish the current task with answer. We note that this action space mostly can only retrieve a small part of a passage based on exact passage name, which is significantly weaker than state-of-the-art lexical or neural retrievers. The purpose is to simulate how humans would interact with Wikipedia, and force models to retrieve via explicit reasoning in language.
动作空间 —— 我们设计了一个简单的维基百科 Web API,含三类动作支持交互式信息检索:(1) search[实体]——若实体页面存在则返回该页前 5 句,否则从维基百科搜索引擎返回最相似的前 5 个实体;(2) lookup[字符串]——返回页面中下一个包含该字符串的句子,模拟浏览器的 Ctrl+F 功能;(3) finish[答案]——以该答案结束当前任务。需要指出,这个动作空间大多只能按精确页面名检索到段落的一小部分,显著弱于最强的词法或神经检索器;其目的是模拟人类与维基百科的交互方式,迫使模型通过显式的语言推理来检索。
**ReAct Prompting** — For HotpotQA and Fever, we randomly select 6 and 3 cases from the training set and manually compose ReAct-format trajectories to use as few-shot exemplars in the prompts. Similar to Figure 1(d), each trajectory consists of multiple thought-action-observation steps (i.e. dense thought), where free-form thoughts are used for various purposes. Specifically, we use a combination of thoughts that decompose questions ("I need to search x, find y, then find z"), extract information from Wikipedia observations ("x was started in 1844", "The paragraph does not tell x"), perform commonsense ("x is not y, so z must instead be...") or arithmetic reasoning ("1844 < 1989"), guide search reformulation ("maybe I can search/look up x instead"), and synthesize the final answer ("...so the answer is x"). See Appendix C for more details.
ReAct 提示 —— 对 HotpotQA 与 FEVER,我们分别从训练集随机选 6 与 3 个案例,人工编写 ReAct 格式轨迹作为提示中的少样本范例。与图 1(d) 类似,每条轨迹由多个"思考-动作-观测"步组成(即密集思考),自由形式的思考用于多种目的:分解问题("我需要搜 x,找 y,再找 z")、从维基百科观测中提取信息("x 始建于 1844 年"/"这段没有说 x")、进行常识("x 不是 y,所以 z 应该是……")或算术推理("1844 < 1989")、引导改写搜索("也许我可以改搜/改查 x")、以及综合出最终答案("……所以答案是 x")。详见附录 C。
**Baselines** — We systematically ablate ReAct trajectories to build prompts for multiple baselines (with formats as Figure 1(1a-1c)): (a) Standard prompting (Standard), which removes all thoughts, actions, observations in ReAct trajectories. (b) Chain-of-thought prompting (CoT) (Wei et al., 2022), which removes actions and observations and serve as a reasoning-only baseline. We also build a self-consistency baseline (CoT-SC) (Wang et al., 2022a;b) by sampling 21 CoT trajectories with decoding temperature 0.7 during inference and adopting the majority answer, which is found to consistently boost performance over CoT. (c) Acting-only prompt (Act), which removes thoughts in ReAct trajectories, loosely resembling how WebGPT (Nakano et al., 2021) interacts with the Internet to answer questions, though it operates on a different task and action space, and uses imitation and reinforcement learning instead of prompting.
基线 —— 我们通过系统性消融 ReAct 轨迹来构建多个基线提示(格式见图 1(1a-1c)):(a) 标准提示(Standard):去掉轨迹中全部思考、动作、观测;(b) 思维链提示(CoT):去掉动作与观测,作为"仅推理"基线;另构建自洽性基线(CoT-SC):推理时以解码温度 0.7 采样 21 条 CoT 轨迹、取多数答案(被发现能稳定提升 CoT);(c) 仅动作提示(Act):去掉轨迹中的思考,大致类似 WebGPT 与互联网交互答题的方式(但任务与动作空间不同,且其用模仿/强化学习而非提示)。
**Combining Internal and External Knowledge** — As will be detail in Section 3.3, we observe that the problem solving process demonstrated by ReAct is more factual and grounded, whereas CoT is more accurate in formulating reasoning structure but can easily suffer from hallucinated facts or thoughts. We therefore propose to incorporate ReAct and CoT-SC, and let the model decide when to switch to the other method based on the following heuristics: A) ReAct→CoT-SC: when ReAct fails to return an answer within given steps, back off to CoT-SC. We set 7 and 5 steps for HotpotQA and FEVER respectively as we find more steps will not improve ReAct performance. B) CoT-SC→ReAct: when the majority answer among n CoT-SC samples occurs less than n/2 times (i.e. internal knowledge might not support the task confidently), back off to ReAct.
结合内部与外部知识 —— 如 3.3 节所述,我们观察到 ReAct 展示的求解过程更事实、更接地,而 CoT 更擅长组织推理结构但容易遭受事实或思考的幻觉。因此我们提出把 ReAct 与 CoT-SC 结合,按以下启发式让模型决定何时切换:A) ReAct→CoT-SC:当 ReAct 在给定步数内未能给出答案,回退到 CoT-SC(HotpotQA/FEVER 分别设 7/5 步,更多步数不再提升);B) CoT-SC→ReAct:当 n 个 CoT-SC 样本中的多数答案出现次数少于 n/2(即内部知识可能不足以自信支持该任务),回退到 ReAct。
**Finetuning** — Due to the challenge of manually annotating reasoning traces and actions at scale, we consider a bootstraping approach similar to Zelikman et al. (2022), using 3,000 trajectories with correct answers generated by ReAct (also for other baselines) to finetune smaller language models (PaLM-8/62B) to decode trajectories (all thoughts, actions, observations) conditioned on input questions/claims. More details are in Appendix B.1.
微调 —— 由于大规模人工标注推理轨迹与动作很困难,我们采用类似 Zelikman et al. (2022) 的自举方法:用 ReAct(及其他基线)生成的 3,000 条答案正确的轨迹微调较小的语言模型(PaLM-8/62B),使其在输入问题/断言条件下解码完整轨迹(全部思考、动作、观测)。细节见附录 B.1。
**ReAct outperforms Act consistently** — Table 1 shows HotpotQA and Fever results using PaLM-540B as the base model with different prompting methods. We note that ReAct is better than Act on both tasks, demonstrating the value of reasoning to guide acting, especially for synthesizing the final answer, as shown in Figure 1 (1c-d). Fine-tuning results also confirm the benefit of reasoning traces for more informed acting.
ReAct 一致优于 Act —— 表 1 给出以 PaLM-540B 为基座、不同提示方法在 HotpotQA 与 FEVER 上的结果。ReAct 在两任务上都优于 Act,证明推理对引导行动的价值(尤其是综合出最终答案时,见图 1(1c-d));微调结果也证实推理轨迹有助于更有依据的行动。
表 1:PaLM-540B 在 HotpotQA 与 Fever 上的提示结果
| 提示方法 | HotpotQA(EM) | Fever(Acc) |
|---|---|---|
| Standard | 28.7 | 57.1 |
| CoT | 29.4 | 56.3 |
| CoT-SC | 33.4 | 60.4 |
| Act | 25.7 | 58.9 |
| ReAct | 27.4 | 60.9 |
| CoT-SC→ReAct | 34.2 | 64.6 |
| ReAct→CoT-SC | 35.1 | 62.0 |
| 监督式 SoTA | 67.5 | 89.5 |
**ReAct vs. CoT** — On the other hand, ReAct outperforms CoT on Fever (60.9 vs 56.3) and slightly lags behind CoT on HotpotQA (27.4 vs 29.4). Fever claims for SUPPORTS/REFUTES might only differ by a slight amount (see Appendix D.1), so acting to retrieve accurate and up-to-date knowledge is vital. To better understand the behavioral difference between ReAct and CoT on HotpotQA, we randomly sampled 50 trajectories with correct and incorrect answers (judged by EM) from ReAct and CoT respectively (thus 200 examples in total), and manually labeled their success and failure modes in Table 2. Some key observations are as follows:
ReAct vs. CoT —— 另一方面,ReAct 在 Fever 上胜过 CoT(60.9 对 56.3),在 HotpotQA 上略逊于 CoT(27.4 对 29.4)。FEVER 的 SUPPORTS/REFUTES 断言可能只有细微差别(见附录 D.1),因此行动去检索准确、时新的知识至关重要。为更好理解二者在 HotpotQA 上的行为差异,我们从 ReAct 与 CoT 各随机抽 50 条答案正确与错误的轨迹(共 200 例),人工标注其成功与失败模式(表 2)。关键观察:
表 2:ReAct 与 CoT 在 HotpotQA 上的成功/失败模式(人工分析)
| 类型 | 定义 | ReAct | CoT |
|---|---|---|---|
| 成功·真阳性 | 推理轨迹与事实均正确 | 94% | 86% |
| 成功·假阳性 | 推理轨迹或事实有幻觉 | 6% | 14% |
| 失败·推理错误 | 推理轨迹错误(含无法跳出重复步骤) | 47% | 16% |
| 失败·检索结果错误 | 检索返回空或无有用信息 | 23% | – |
| 失败·幻觉 | 推理轨迹或事实有幻觉 | 0% | 56% |
| 失败·标签歧义 | 预测正确但与标签不精确匹配 | 29% | 28% |
A) Hallucination is a serious problem for CoT, resulting in much higher false positive rate than ReAct (14% vs. 6%) in success mode, and make up its major failure mode (56%). In contrast, the problem solving trajectory of ReAct is more grounded, fact-driven, and trustworthy, thanks to the access of an external knowledge base. B) While interleaving reasoning, action and observation steps improves ReAct's groundedness and trustworthiness, such a structural constraint also reduces its flexibility in formulating reasoning steps, leading to more reasoning error rate than CoT. we note that there is one frequent error pattern specific to ReAct, in which the model repetitively generates the previous thoughts and actions, and we categorize it as part of "reasoning error" as the model fails to reason about the proper next action to take and jump out of the loop. C) For ReAct, successfully retrieving informative knowledge via search is critical. Non-informative search, which counts for 23% of the error cases, derails the model reasoning and gives it a hard time to recover and reformulate thoughts. This is perhaps an expected trade-off between factuality and flexibility, which motivates our proposed strategies of combining two methods.
A) 幻觉是 CoT 的严重问题:成功模式中的假阳性率远高于 ReAct(14% 对 6%),且构成其主要失败模式(56%)。相比之下,得益于外部知识库的访问,ReAct 的求解轨迹更接地、更事实驱动、更可信。B) 交错"推理-动作-观测"步虽然提升了 ReAct 的接地性与可信度,这种结构约束也降低了其组织推理步骤的灵活性,导致推理错误率高于 CoT;ReAct 特有的一种高频错误是重复生成先前的思考与动作——我们归入"推理错误",因为模型没能推理出正确的下一步动作、跳不出循环。C) 对 ReAct 而言,通过搜索成功检索到有信息量的知识至关重要:无信息量的检索占错误案例的 23%,会把模型推理带偏,使其难以恢复并重新组织思考。这可视为事实性与灵活性之间意料之中的权衡,也促使我们提出结合两法的策略。
**ReAct + CoT-SC perform best for prompting LLMs** — Also shown in Table 1, the best prompting method on HotpotQA and Fever are ReAct→CoT-SC and CoT-SC→ReAct respectively. Furthermore, Figure 2 shows how different methods perform with respect to the number of CoT-SC samples used. While two ReAct + CoT-SC methods are advantageous at one task each, they both significantly and consistently outperform CoT-SC across different number of samples, reaching CoT-SC performance with 21 samples using merely 3-5 samples. These results indicate the value of properly combining model internal knowledge and external knowledge for reasoning tasks.
ReAct + CoT-SC 是提示 LLM 的最佳组合 —— 表 1 还显示,HotpotQA 与 Fever 上的最佳提示方法分别是 ReAct→CoT-SC(35.1)与 CoT-SC→ReAct(64.6)。图 2 进一步展示不同方法随 CoT-SC 采样数的变化:两种结合法各自在一个任务上占优,但都显著且一致地胜过 CoT-SC——仅用 3-5 个样本即可达到 CoT-SC 用 21 个样本的性能。这说明恰当结合模型内部知识与外部知识对推理任务的价值。
**ReAct performs best for fine-tuning** — Figure 3 shows the scaling effect of prompting/finetuning four methods (Standard, CoT, Act, ReAct) on HotpotQA. With PaLM-8/62B, prompting ReAct performs worst among four methods due to the difficulty to learn both reasoning and acting from in-context examples. However, when finetuned with just 3,000 examples, ReAct becomes the best method among the four, with PaLM-8B finetuned ReAct outperforming all PaLM-62B prompting methods, and PaLM-62B finetuned ReAct outperforming all 540B prompting methods. In contrast, finetuning Standard or CoT is significantly worse than finetuning ReAct or Act for both PaLM-8/62B, as the former essentially teaches models to memorize (potentially hallucinated) knowledge facts, and the latter teaches models how to (reason and) act to access information from Wikipedia, a more generalizable skill for knowledge reasoning. As all prompting methods are still significantly far from domain-specific state-of-the-art approaches (Table 1), we believe finetuning with more human-written data might be a better way to unleash the power of ReAct.
ReAct 微调后表现最佳 —— 图 3 展示四种方法(Standard、CoT、Act、ReAct)在 HotpotQA 上提示/微调的规模效应。用 PaLM-8/62B 时,提示式 ReAct 在四法中最差——小模型难以从上下文示例中同时学会推理与行动;但仅用 3,000 个示例微调后,ReAct 跃居四法之首:PaLM-8B 微调 ReAct 超过全部 PaLM-62B 提示方法,PaLM-62B 微调 ReAct 超过全部 540B 提示方法。相比之下,微调 Standard 或 CoT 显著差于微调 ReAct 或 Act——前者本质上是教模型记忆(可能幻觉的)知识事实,后者教模型如何(推理并)行动去访问维基百科信息,这是知识推理中更可泛化的技能。由于所有提示方法与领域专用最优方法仍有明显差距(表 1),我们相信用更多人工数据微调才是释放 ReAct 潜力的更好途径。
4 决策任务
We also test ReAct on two language-based interactive decision-making tasks, ALFWorld and WebShop, both of which feature complex environments that require agents to act over long horizons with sparse rewards, warranting the need for reasoning to act and explore effectively.
我们还在两个基于语言的交互决策任务上测试 ReAct:ALFWorld 与 WebShop。二者都是复杂环境,要求智能体在长时程、稀疏奖励下行动,因而需要推理来有效地行动与探索。
**ALFWorld** — ALFWorld (Shridhar et al., 2020b) (Figure 1(2)) is a synthetic text-based game designed to align with the embodied ALFRED benchmark (Shridhar et al., 2020a). It includes 6 types of tasks in which an agent needs to achieve a high-level goal (e.g. examine paper under desklamp) by navigating and interacting with a simulated household via text actions (e.g. go to coffeetable 1, take paper 2, use desklamp 1). A task instance can have more than 50 locations and take an expert policy more than 50 steps to solve, thus challenging an agent to plan and track subgoals, as well as explore systematically (e.g. check all desks one by one for desklamp). In particular, one challenge built into ALFWorld is the need to determine likely locations for common household items (e.g. desklamps will likely be on desks, shelfs, or dressers), making this environment a good fit for LLMs to exploit their pretrained commonsense knowledge. To prompt ReAct, we randomly annotate three trajectories from the training set for each task type, where each trajectory includes sparse thoughts that (1) decompose the goal, (2) track subgoal completion, (3) determine the next subgoal, and (4) reason via commonsense where to find an object and what to do with it. We show prompts used for ALFWorld in Appendix C.4. Following Shridhar et al. (2020b), we evaluate on 134 unseen evaluation games in a task-specific setup. For robustness, we construct 6 prompts for each task type through each permutation of 2 annotated trajectories from the 3 we annotate. Act prompts are constructed using the same trajectories, but without thoughts — since task instances are randomly chosen from the training set, it favors neither ReAct nor Act and provides a fair and controlled comparison to test the importance of sparse thoughts. For baselines, we use BUTLER (Shridhar et al., 2020b), an imitation learning agent trained on 10⁵ expert trajectories for each task type.
ALFWorld —— ALFWorld(图 1(2))是与具身基准 ALFRED 对齐的合成文字游戏,含 6 类任务:智能体需通过文本动作(如 go to coffeetable 1、take paper 2、use desklamp 1)在模拟家庭中导航与交互以达成高层目标(如"在台灯下查看纸张")。一个任务实例可有 50+ 个地点、专家策略也需 50+ 步才能解决,因此考验智能体规划与跟踪子目标的能力以及系统性探索(如逐个检查所有桌子找台灯)的能力。ALFWorld 内置的一个挑战是判断常见家居物品的可能位置(台灯大概率在书桌、架子或衣柜上),这让该环境很适合 LLM 发挥其预训练常识。为提示 ReAct,我们对每类任务从训练集随机标注 3 条轨迹,每条含稀疏思考,用于:(1) 分解目标、(2) 跟踪子目标完成、(3) 确定下一子目标、(4) 用常识推理物品在哪、该怎么处理(提示词见附录 C.4)。沿用 Shridhar et al. (2020b),我们在任务特定设定下于 134 个未见评测游戏上评测;为稳健性,每类任务用 3 条标注轨迹中任取 2 条的排列构造 6 个提示。Act 提示用相同轨迹但去掉思考——由于任务实例从训练集随机选取,这对 ReAct 与 Act 都无偏,构成检验稀疏思考价值的公平受控对比。基线为 BUTLER——每类任务用 10⁵ 条专家轨迹训练的模仿学习智能体。
**WebShop** — Can ReAct also interact with noisy real-world language environments for practical applications? We investigate WebShop (Yao et al., 2022), a recently proposed online shopping website environment with 1.18M real-world products and 12k human instructions. Unlike ALFWorld, Webshop contains a high variety of structured and unstructured texts (e.g. product titles, descriptions, and options crawled from Amazon), and requires an agent to purchase a product based on a user instruction (e.g. "I am looking for a nightstand with drawers. It should have a nickel finish, and priced lower than $140") through web interactions (e.g. search "nightstand drawers", choose buttons such as "color: modern-nickel-white" or "back to search"). This task is evaluated by average score (percentage of desired attributes covered by the chosen product averaged across all episodes) and success rate (percentage of episodes where the chosen product satisfies all requirements) on 500 test instructions. We formulate Act prompts with actions to search, choose product, choose options, and buy, with ReAct prompts additionally reasoning to determine what to explore, when to buy, and what products options are relevant to the instruction. See Table 6 for an example prompt, and Table 10 for model predictions in the Appendix. We compare to an imitation learning (IL) method trained with 1,012 human annotated trajectories, and a imitation + reinforcement learning (IL + RL) method additionally trained with 10,587 training instructions.
WebShop —— ReAct 能否在嘈杂的真实语言环境中交互以服务实际应用?我们研究 WebShop——新提出的在线购物网站环境,含 118 万件真实商品与 1.2 万条人类指令。与 ALFWorld 不同,WebShop 包含高度多样的结构化与非结构化文本(从亚马逊爬取的商品标题、描述、选项),要求智能体依据用户指令(如"我想买个带抽屉的床头柜,镍色饰面,价格低于 140 美元")通过网页交互(如搜索 "nightstand drawers"、点击 "color: modern-nickel-white" 或 "back to search" 等按钮)购买商品。任务在 500 条测试指令上以平均分(所选商品覆盖期望属性的百分比,跨回合平均)与成功率(所选商品满足全部要求的回合占比)评估。Act 提示包含搜索、选商品、选选项、购买等动作;ReAct 提示额外推理:探索什么、何时买、哪些商品选项与指令相关(示例提示见表 6,模型预测见附录表 10)。对比基线:用 1,012 条人工标注轨迹训练的模仿学习(IL),以及额外用 10,587 条训练指令训练的模仿+强化学习(IL+RL)。
**Results** — ReAct outperforms Act on both ALFWorld (Table 3) and Webshop (Table 4). On ALFWorld, the best ReAct trial achieves an average success rate of 71%, significantly outperforming the best Act (45%) and BUTLER (37%) trials. In fact, even the worse ReAct trial (48%) beats the best trial of both methods. Moreover, the advantage of ReAct over Act is consistent across six controlled trials, with relative performance gain ranging from 33% to 90% and averaging 62%. Qualitatively, we saw that, without any thoughts at all, Act fails to correctly decompose goals into smaller subgoals, or loses track of the current state of the environment. Example trajectories comparing ReAct and Act can be found in Appendix D.2.1 and Appendix D.2.2.
结果 —— ReAct 在 ALFWorld(表 3)与 WebShop(表 4)上都胜过 Act。ALFWorld 上,最佳 ReAct 试验平均成功率 71%,显著超过最佳 Act(45%)与 BUTLER(37%);事实上最差的 ReAct 试验(48%)也赢过两法的最佳试验。且在六次受控试验中 ReAct 对 Act 的优势一致,相对增益 33%-90%、平均 62%。定性来看,完全没有思考的 Act 无法把目标正确分解为子目标,或跟丢环境的当前状态(对比轨迹见附录 D.2.1/D.2.2)。
表 3:ALFWorld 任务特定成功率(%)
| 方法 | Pick | Clean | Heat | Cool | Look | Pick 2 | 全部 |
|---|---|---|---|---|---|---|---|
| Act(6 选优) | 88 | 42 | 74 | 67 | 72 | 41 | 45 |
| ReAct(平均) | 65 | 39 | 83 | 76 | 55 | 24 | 57 |
| ReAct(6 选优) | 92 | 58 | 96 | 86 | 78 | 41 | 71 |
| ReAct-IM(平均) | 55 | 59 | 60 | 55 | 23 | 24 | 48 |
| ReAct-IM(6 选优) | 62 | 68 | 87 | 57 | 39 | 33 | 53 |
| BUTLER_g(8 选优) | 33 | 26 | 70 | 76 | 17 | 12 | 22 |
| BUTLER(8 选优) | 46 | 39 | 74 | 100 | 22 | 24 | 37 |
表 4:WebShop 上的得分与成功率(SR)
| 方法 | Score | SR |
|---|---|---|
| Act | 62.3 | 30.1 |
| ReAct | 66.6 | 40.0 |
| IL | 59.9 | 29.1 |
| IL+RL | 62.4 | 28.7 |
| 人类 | 82.1 | 59.6(专家) |
On Webshop, one-shot Act prompting already performs on par with IL and IL+RL methods. With additional sparse reasoning, ReAct achieves significantly better performance, with an absolute 10% improvement over the previous best success rate. By checking examples, we find that ReAct is more likely to identify instruction-relevant products and options by reasoning to bridge the gap between noisy observations and actions (e.g. "For 'space-saving ottoman bench for living room', the item has options '39x18x18inch' and 'blue' and seems good to buy."). However, existing methods are still far from the performance of expert humans (Table 4), who perform significantly more product explorations and query re-formulations that are still challenging for prompting-based methods.
WebShop 上,一例 Act 提示已与 IL、IL+RL 相当;加上稀疏推理后,ReAct 显著更强,成功率绝对提升 10% 超过此前最佳。查例发现,ReAct 更容易通过推理弥合嘈杂观测与动作之间的鸿沟,从而识别与指令相关的商品和选项(如"对于'客厅省空间软凳',该商品有 '39x18x18inch' 和 'blue' 选项,看起来值得买")。但现有方法距专家人类仍有差距(表 4)——人类会做多得多的商品探索与查询改写,这对提示式方法仍是挑战。
**On the value of internal reasoning vs. external feedback** — To our knowledge, ReAct is the first demonstration of combined reasoning and action using an LLM applied to an interactive environment within a closed-loop system. Perhaps the closest prior work is Inner Monologue (IM), from Huang et al. (2022b), in which actions from an embodied agent are motivated by an eponymous "inner monologue". However, IM's "inner monologue" is limited to observations of the environment state and what needs to be completed by the agent for the goal to be satisfied. In contrast, the reasoning traces in ReAct for decision making is flexible and sparse, allowing diverse reasoning types (see Section 2) to be induced for different tasks. To demonstrate the differences between ReAct and IM, and to highlight the importance of internal reasoning vs. simple reactions to external feedback, we ran an ablation experiment using a thought pattern composed of IM-like dense external feedback. As can be seen in Table 3, ReAct substantially outperforms IM-style prompting (ReAct-IM) (71 vs. 53 overall success rate), with consistent advantages on five out of six tasks. Qualitatively, we observed that ReAct-IM often made mistakes in identifying when subgoals were finished, or what the next subgoal should be, due to a lack of high-level goal decomposition. Additionally, many ReAct-IM trajectories struggled to determine where an item would likely be within the ALFWorld environment, due to a lack of commonsense reasoning. Both shortcomings can be addressed in the ReAct paradigm. More details about ReAct-IM is in Appendix B.2.
内部推理 vs. 外部反馈的价值 —— 据我们所知,ReAct 是首个用 LLM 把推理与行动结合、应用于交互环境的闭环系统演示。最接近的先前工作是 Inner Monologue(IM,Huang et al., 2022b):具身智能体的动作由同名的"内心独白"驱动。但 IM 的"内心独白"仅限于对环境状态的观察、以及为满足目标智能体还需完成什么。相比之下,ReAct 用于决策的推理轨迹灵活而稀疏,允许为不同任务诱导出多样的推理类型(见第 2 节)。为展示二者差异、凸显内部推理(而非简单响应外部反馈)的重要性,我们做了用 IM 式密集外部反馈作思考模式的消融(ReAct-IM)。表 3 显示 ReAct 大幅胜过 IM 式提示(总成功率 71 对 53),六类任务中五类一致占优。定性上,ReAct-IM 常因缺乏高层目标分解而在判断子目标何时完成、下一子目标是什么时出错;许多 ReAct-IM 轨迹也因缺乏常识推理而难以判断物品在 ALFWorld 环境中的可能位置。这两个短板都能在 ReAct 范式中解决(细节见附录 B.2)。
5 相关工作
**Language model for reasoning** — Perhaps the most well-known work of using LLMs for reasoning is Chain-of-Thought (CoT) (Wei et al., 2022), which reveals the ability of LLMs to formulate their own "thinking procedure" for problem solving. Several follow-up works have since been performed, including least-to-most prompting for solving complicated tasks (Zhou et al., 2022), zero-shot-CoT (Kojima et al., 2022), and reasoning with self-consistency (Wang et al., 2022a). Recently, (Madaan & Yazdanbakhsh, 2022) systematically studied the formulation and structure of CoT, and observed that the presence of symbols, patterns and texts is crucial to the effectiveness of CoT. Other work has also been extended to more sophisticated reasoning architecture beyond simple prompting. For example Selection-Inference (Creswell et al., 2022) divides the reasoning process into two steps of "selection" and "inference". STaR (Zelikman et al., 2022) bootstraps the reasoning process by finetuning the model on correct rationales generated by the model itself. Faithful reasoning (Creswell & Shanahan, 2022) decomposes multi-step reasoning into three steps, each performed by a dedicated LM respectively. Similar approaches like Scratchpad (Nye et al., 2021), which finetunes a LM on intermediate computation steps, also demonstrate improvement on multi-step computation problems. In contrast to these methods, ReAct performs more than just isolated, fixed reasoning, and integrates model actions and their corresponding observations into a coherent stream of inputs for the model to reason more accurately and tackle tasks beyond reasoning (e.g. interactive decision making).
用于推理的语言模型 —— 最著名的用 LLM 做推理的工作是思维链 CoT(Wei et al., 2022),揭示了 LLM 为解题构造自身"思考过程"的能力;后续包括最少到最多提示(Zhou et al., 2022)、零样本 CoT(Kojima et al., 2022)、自洽性推理(Wang et al., 2022a)等。近期 Madaan & Yazdanbakhsh (2022) 系统研究了 CoT 的表述与结构,发现符号、模式与文本的存在对 CoT 有效性至关重要。另一些工作扩展到简单提示之外的精巧推理架构:Selection-Inference 把推理分为"选择"与"推断"两步;STaR 用模型自生的正确理由微调模型来自举推理过程;Faithful reasoning 把多步推理拆成三步、各由专门 LM 执行;Scratchpad 在中间计算步骤上微调 LM,也在多步计算问题上带来提升。与这些方法相比,ReAct 不只做孤立、固定的推理,而是把模型动作及其观测整合成连贯的输入流,使模型推理更准确,并能处理推理之外的任务(如交互决策)。
**Language model for decision making** — The strong capability of LLMs has enabled them to perform tasks beyond language generation, and it is becoming more popular to take advantage of LLMs as a policy model for decision making, especially in interactive environments. WebGPT (Nakano et al., 2021) uses an LM to interact with web browsers, navigate through web pages, and infer answers to complicated questions from ELI5. In comparison to ReAct, WebGPT does not explicitly model the thinking and reasoning procedure, instead rely on expensive human feedback for reinforcement learning. In conversation modeling, chatbots like BlenderBot and Sparrow and task-oriented dialogue systems like SimpleTOD also train LMs to make decision about API calls. Unlike ReAct, they do not explicitly consider the reasoning procedure either, and also relies on expensive datasets and human feedback collections for policy learning. In contrast, ReAct learns a policy in a much cheaper way, since the decision making process only requires language description of the reasoning procedure.
用于决策的语言模型 —— LLM 的强大能力使其能胜任语言生成之外的任务,把 LLM 当作决策的策略模型(尤其在交互环境中)日益流行。WebGPT 用 LM 与浏览器交互、浏览网页、推断 ELI5 复杂问题的答案;与 ReAct 相比,WebGPT 不显式建模思考与推理过程,而是依赖昂贵的人类反馈做强化学习。对话建模中,BlenderBot、Sparrow 等聊天机器人与 SimpleTOD 等任务型对话系统也训练 LM 决策 API 调用;它们同样不显式考虑推理过程,且依赖昂贵的数据集与人类反馈做策略学习。相比之下,ReAct 以便宜得多的方式学习策略——决策过程只需要推理过程的语言描述。
LLMs have also been increasingly employed in interactive and embodied environments for planning and decision making. Perhaps most relevant to ReAct in this respect are SayCan (Ahn et al., 2022) and Inner Monologue (Huang et al., 2022b), which use LLMs for robotic action planning and decision making. In SayCan, LLMs were prompted to directly predict possible actions a robot can take, which is then reranked by an affordance model grounded on the visual environments for final prediction. Inner Monologue made further improvements by adding the eponymous "inner monologue", which is implemented as injected feedback from the environment. To our knowledge, Inner Monologue is the first work that demonstrates such a closed-loop system, which ReAct builds on. However, we argue that Inner Monologue does not truly comprise of inner thoughts — this is elaborated in Section 4. We also note that leveraging language as semantically-rich inputs in the process of interactive decision making has been shown to be successful under other settings. It is becoming more evident that with the help of LLMs, language as a fundamental cognitive mechanism will play a critical role in interaction and decision making. What is more, progress in LLMs has also inspired the development of versatile and generalist agents like Reed et al. (2022).
LLM 也越来越多地被用于交互与具身环境中的规划与决策。与此最相关的是 SayCan 与 Inner Monologue,都用 LLM 做机器人动作规划与决策:SayCan 提示 LLM 直接预测机器人可采取的动作,再由接地于视觉环境的可供性(affordance)模型重排序得出最终预测;Inner Monologue 进一步加入同名的"内心独白"——实现为从环境注入的反馈。据我们所知,Inner Monologue 是首个展示此类闭环系统的工作,ReAct 建立在其上;但我们认为 Inner Monologue 并不真正包含内在思考——第 4 节有详述。在交互决策中把语言用作语义丰富的输入,在其他设定下也被证明成功。越来越清楚的是:在 LLM 的助力下,语言作为一种基础认知机制将在交互与决策中扮演关键角色。LLM 的进步也催生了 Reed et al. (2022) 等多面手通用智能体。
6 结论
We have proposed ReAct – a simple yet effective method for synergizing reasoning and acting in large language models. Through a diverse set of experiments on multi-hop question-answering, fact checking, and interactive decision-making tasks, we show that ReAct leads to superior performance with interpretable decision traces. Despite the simplicity of our method, complex tasks with large action spaces require more demonstrations to learn well, which unfortunately can easily go beyond the input length limit of in-context learning. We explore the fine-tuning approach on HotpotQA with initial promising results, but learning from more high-quality human annotations will be the desiderata to further improve the performance. Scaling up ReAct with multi-task training and combining it with complementary paradigms like reinforcement learning could result in stronger agents that further unlock the potential of LLMs for more applications.
我们提出了 ReAct——在大语言模型中协同推理与行动的简单而有效的方法。通过在多跳问答、事实核查与交互决策任务上的多样实验,我们证明 ReAct 以可解释的决策轨迹带来更优性能。尽管方法简单,但动作空间大的复杂任务需要更多示范才能学好,而这很容易超出上下文学习的输入长度限制。我们在 HotpotQA 上探索了微调途径并取得初步可喜结果,但从更多高质量人工标注中学习是进一步提升性能的应然之需。通过多任务训练扩展 ReAct、并与强化学习等互补范式结合,有望造出更强大的智能体,进一步释放 LLM 在更多应用上的潜力。
**Ethics Statement** — ReAct prompts large language models to generate more human interpretable, diagnosable, and controllable task-solving trajectories than previous methods. However, hooking up a large language model with an action space to interact with external environments (e.g. the web, physical environments) has potential dangers, e.g. looking up inappropriate or private information, or taking harmful actions in an environment. Our experiments minimize such risks by limiting the interactions to specific websites (Wikipedia or WebShop) that are free of private information, without any dangerous actions in the action space design (i.e. models cannot really buy products on WebShop the research benchmark, or edit Wikipedia). We believe researchers should be aware of such risks before designing more extensive experiments in the future.
伦理声明 —— 与先前方法相比,ReAct 促使 LLM 生成更可解释、可诊断、可控的任务求解轨迹。但把大语言模型接上动作空间与外部环境(网络、物理环境)交互存在潜在危险,如查阅不当或隐私信息、在环境中采取有害动作。我们的实验把交互限制在不含隐私信息的特定网站(维基百科或 WebShop),且动作空间设计中不含危险动作(模型无法真正在研究基准 WebShop 上购买商品、也无法编辑维基百科),以此最小化此类风险。我们认为研究者在设计更大规模的实验前应意识到这些风险。
要点速览
- 核心思想:动作空间增广 Â = A ∪ L,"思考"是不影响环境的语言动作——交错生成思考与动作,推理引导行动(reason to act)、行动反哺推理(act to reason)。
- 知识任务(HotpotQA/FEVER):ReAct 胜 Act;与 CoT 互有胜负,但 ReAct 轨迹更接地——CoT 的失败 56% 源于幻觉,ReAct 为 0%(代价是推理错误率升高 47% vs 16%,含"重复循环"特有错误)。
- 最佳组合:ReAct↔CoT-SC 按启发式互为回退,用 3-5 个样本达到 CoT-SC 21 个样本的性能(HotpotQA 35.1 / FEVER 64.6)。
- 微调发现:3,000 条自举轨迹微调后,8B ReAct 胜过全部 62B 提示方法、62B ReAct 胜过全部 540B 提示方法——教"如何推理并行动获取信息"比教"记忆事实"更可泛化。
- 决策任务:ALFWorld 71% 最佳试验(Act 45%、BUTLER 模仿学习 37%),最差 ReAct 试验也赢两法最佳;WebShop 40% SR(绝对 +10% 胜 IL/IL+RL),仅 1-2 个示例。
- 稀疏思考 > 密集外部反馈:IM 式"复述环境状态"的 ReAct-IM 仅 53%,缺乏目标分解与常识推理。
- 设计四优点:直观易设计、通用灵活、高性能稳健(仅 1-6 示例)、人类对齐可控(可编辑思考实时纠偏)。
- 历史地位:几乎所有后续智能体(SWE-agent、OpenHands、Claude Code、AutoGen)都构建在"思考-动作-观测"循环之上;与 CoT 的互补也预示了第 5 周的"测试时计算"与"反思优化"路线。
- 附录提示:GPT-3 结果、全部任务提示词与样例轨迹在原文附录 A-G,可在本页顶部打开原文 PDF 查看。