CS329Z中文学习站

"LIMA:对齐中的「少即是多」"

"LIMA: Less Is More for Alignment"

"Chunting Zhou et al." · "NeurIPS 2023 · Meta AI"

必读全译对照查看原文(PDF) ↗

导读

在课程第 7 周「数据选择与质量」单元中,LIMA 是讨论「数据质量 vs 数量」绕不开的名作。2023 年,当业界普遍认为对齐(alignment)需要数百万条指令数据与人类反馈强化学习(RLHF)时,Meta AI 团队用一项极其简洁的实验提出异议:在 LLaMa 65B 上仅用 1,000 条精心筛选的提示-回复对做标准监督微调,不使用任何 RLHF 或偏好建模,得到的模型 LIMA 在人类评估中即可与 DaVinci003、Bard 抗衡,甚至在 43% 的场景下不输 GPT-4。

论文提出「表层对齐假说(Superficial Alignment Hypothesis)」:模型的知识与能力几乎全部在预训练阶段习得,对齐只是教模型在与用户交互时应使用哪种输出格式子分布。这一论断直接塑造了此后两年「以质取胜」的数据策展(data curation)思潮——质量与多样性带来的收益远超单纯堆量,这也是本讲将其与 SWE-smith、Data Flywheels 等数据工程文献并列的原因。对构建 AI Agent 的学习者而言,LIMA 提示:与其无差别扩大微调语料,不如把预算投入在高价值样本的筛选与撰写上。本页为全文中英对照版本:英文原段与中文全译逐段交替,图表与附录要点以译注形式覆盖。

全文对照翻译

译注:以下为全文中英对照,覆盖摘要、第 1-7 节正文与附录 A-E 的实质内容(原文第 1-15 页,依据 arXiv:2305.11206v1 提取)。References(参考文献)按惯例不收录。表 1 已转为 Markdown 表格(数据保留);图 1-13 为版式图形,以「[图 N: 英文图题] + 中文说明」呈现,关键数值均保留在说明中;图 4、8、10、13 中的完整样例文本较长,以内容概述替代,全文请查阅原 PDF。英文原段按论文原文誊录,仅修复 PDF 提取造成的空格缺失、连字(fi/fl)与希腊字母(τ、β)等显示问题,未改动任何措辞;个别因双栏版式被吞掉的短语按上下文补全。

EN

**LIMA: Less Is More for Alignment**
Chunting Zhou\*, Pengfei Liu\*, Puxin Xu, Srini Iyer, Jiao Sun, Yuning Mao, Xuezhe Ma, Avia Efrat, Ping Yu, Lili Yu, Susan Zhang, Gargi Ghosh, Mike Lewis, Luke Zettlemoyer, Omer Levy — Meta AI · Carnegie Mellon University · University of Southern California · Tel Aviv University.
*Preprint. Under review.* arXiv:2305.11206v1 [cs.CL] 18 May 2023.

标题与作者 —— 《LIMA:对齐中的「少即是多」》。作者:Chunting Zhou、Pengfei Liu(共同一作)、Puxin Xu、Srini Iyer、Jiao Sun、Yuning Mao、Xuezhe Ma、Avia Efrat、Ping Yu、Lili Yu、Susan Zhang、Gargi Ghosh、Mike Lewis、Luke Zettlemoyer、Omer Levy,来自 Meta AI、卡内基梅隆大学、南加州大学与特拉维夫大学。(预印本,审稿中。)

摘要(Abstract)

EN

Large language models are trained in two stages: (1) unsupervised pretraining from raw text, to learn general-purpose representations, and (2) large scale instruction tuning and reinforcement learning, to better align to end tasks and user preferences. We measure the relative importance of these two stages by training LIMA, a 65B parameter LLaMa language model fine-tuned with the standard supervised loss on only 1,000 carefully curated prompts and responses, without any reinforcement learning or human preference modeling. LIMA demonstrates remarkably strong performance, learning to follow specific response formats from only a handful of examples in the training data, including complex queries that range from planning trip itineraries to speculating about alternate history. Moreover, the model tends to generalize well to unseen tasks that did not appear in the training data. In a controlled human study, responses from LIMA are either equivalent or strictly preferred to GPT-4 in 43% of cases; this statistic is as high as 58% when compared to Bard and 65% versus DaVinci003, which was trained with human feedback. Taken together, these results strongly suggest that almost all knowledge in large language models is learned during pretraining, and only limited instruction tuning data is necessary to teach models to produce high quality output.

摘要 —— 大语言模型的训练分两个阶段:(1)从原始文本进行无监督预训练,学习通用表示;(2)大规模指令微调与强化学习,以更好地对齐到端任务与用户偏好。我们通过训练 LIMA 来衡量这两个阶段的相对重要性——LIMA 是一个 65B 参数的 LLaMa 语言模型,仅用 1,000 条精心策展(curated)的提示与回复,以标准监督损失微调而成,全程不使用任何强化学习或人类偏好建模。LIMA 展示了非常强劲的性能:仅凭训练数据中寥寥几个示例,它就学会了遵循特定的回复格式,应对从规划旅行行程到推测架空历史等各类复杂请求;而且模型往往能很好地泛化到训练数据中未出现的新任务。在一项受控人类研究中,LIMA 的回复在与 GPT-4 比较时有 43% 的情况「相等或严格更受偏好」;与 Bard 比较时该比例高达 58%,与采用人类反馈训练的 DaVinci003 比较时为 65%。综合来看,这些结果有力地表明:大语言模型中几乎所有知识都在预训练阶段习得,只需有限的指令微调数据即可教会模型产出高质量输出。

1 引言(Introduction)

EN

Language models are pretrained to predict the next token at an incredible scale, allowing them to learn general-purpose representations that can be transferred to nearly any language understanding or generation task. To enable this transfer, various methods for aligning language models have thus been proposed, primarily focusing on instruction tuning [Mishra et al., 2021, Wei et al., 2022a, Sanh et al., 2022] over large multi-million-example datasets [Chung et al., 2022, Beeching et al., 2023, Köpf et al., 2023], and more recently reinforcement learning from human feedback (RLHF) [Bai et al., 2022a, Ouyang et al., 2022], collected over millions of interactions with human annotators. Existing alignment methods require significant amounts of compute and specialized data to achieve ChatGPT-level performance. However, we demonstrate that, given a strong pretrained language model, remarkably strong performance can be achieved by simply fine-tuning on 1,000 carefully curated training examples.

语言模型以惊人的规模预训练于下一词元(token)预测任务,从而学到可迁移至几乎任何语言理解或生成任务的通用表示。为实现这种迁移,学界已提出多种对齐语言模型的方法:早期主要是指令微调,依赖数百万级样本的大型多任务数据集;近期则是人类反馈强化学习,依赖与人类标注者数百万次交互收集的数据。现有对齐方法都需要海量算力与专门数据才能达到 ChatGPT 级别的性能。然而我们证明:给定一个强大的预训练语言模型,仅仅在 1,000 条精心策展的训练示例上微调,就能取得非常强劲的性能。

EN

We hypothesize that alignment can be a simple process where the model learns the style or format for interacting with users, to expose the knowledge and capabilities that were already acquired during pretraining. To test this hypothesis, we curate 1,000 examples that approximate real user prompts and high-quality responses. We select 750 top questions and answers from community forums, such as Stack Exchange and wikiHow, sampling for quality and diversity. In addition, we manually write 250 examples of prompts and responses, while optimizing for task diversity and emphasizing a uniform response style in the spirit of an AI assistant. Finally, we train LIMA, a pretrained 65B-parameter LLaMa model [Touvron et al., 2023] fine-tuned on this set of 1,000 demonstrations.

我们假设,对齐可以是一个简单的过程:模型学习与用户交互的「风格或格式」,以暴露预训练期间已经获得的知识与能力。为检验这一假设,我们策展了 1,000 条接近真实用户提示及其高质量回复的示例:从 Stack Exchange、wikiHow 等社区论坛选出 750 条优质问答,采样时兼顾质量与多样性;另外手工撰写 250 条提示-回复示例,优化任务多样性,并强调统一为「AI 助手」精神的回复风格。最后,我们训练 LIMA——一个在这 1,000 条示范数据上微调的预训练 65B 参数 LLaMa 模型。

EN

We compare LIMA to state-of-the-art language models and products across 300 challenging test prompts. In a human preference study, we find that LIMA outperforms RLHF-trained DaVinci003 from OpenAI, which was trained with RLHF, as well as a 65B-parameter reproduction of Alpaca [Taori et al., 2023], which was trained on 52,000 examples. While humans typically prefer responses from GPT-4, Claude, and Bard over LIMA, this is not always the case; LIMA produces equal or preferable responses in 43%, 46%, and 58% of the cases, respectively. Repeating the human preference annotations with GPT-4 as the annotator corroborates our findings. Analyzing LIMA responses on an absolute scale reveals that 88% meet the prompt requirements, and 50% are considered excellent. Ablation experiments reveal vastly diminishing returns when scaling up data quantity without also scaling up prompt diversity, alongside major gains when optimizing data quality. In addition, despite having zero dialogue examples, we find that LIMA can conduct coherent multi-turn dialogue, and that this ability can be dramatically improved by adding only 30 hand-crafted dialogue chains to the training set. Overall, these remarkable findings demonstrate the power of pretraining and its relative importance over large-scale instruction tuning and reinforcement learning approaches.

我们在 300 条有挑战性的测试提示上,将 LIMA 与最先进的语言模型及产品比较。在人类偏好研究中,LIMA 胜过 OpenAI 用 RLHF 训练的 DaVinci003,也胜过在 52,000 条示例上训练的 Alpaca 65B 复现版。虽然人类总体上更偏好 GPT-4、Claude 与 Bard 的回复,但并非总是如此:LIMA 分别在 43%、46% 与 58% 的场合产出相等或更优的回复。用 GPT-4 替代人类重做偏好标注,结果印证了我们的发现。对 LIMA 回复的绝对量级分析显示:88% 满足提示要求,50% 被评为优秀。消融实验表明:在不同步扩大提示多样性的前提下扩大数据数量,收益急剧递减;而优化数据质量则带来重大收益。此外,尽管训练集中没有一条对话示例,LIMA 也能进行连贯的多轮对话,且只需向训练集追加 30 条手写对话链,该能力即可获得大幅改善。总体而言,这些非凡的发现展示了预训练的力量,及其相对于大规模指令微调与强化学习方法的重要性。

2 对齐数据(Alignment Data)

EN

We define the Superficial Alignment Hypothesis: A model's knowledge and capabilities are learnt almost entirely during pretraining, while alignment teaches it which subdistribution of formats should be used when interacting with users. If this hypothesis is correct, and alignment is largely about learning style, then a corollary of the Superficial Alignment Hypothesis is that one could sufficiently tune a pretrained language model with a rather small set of examples [Kirstain et al., 2021].

我们定义表层对齐假说(Superficial Alignment Hypothesis):模型的知识与能力几乎完全在预训练期间习得,而对齐只是教会它在与用户交互时应使用哪种格式的子分布。如果该假说正确、对齐主要是在学风格,那么表层对齐假说的一个自然推论是:用相当少的一组示例就足以充分调好一个预训练语言模型。

EN

To that end, we collect a dataset of 1,000 prompts and responses, where the outputs (responses) are stylistically aligned with each other, but the inputs (prompts) are diverse. Specifically, we seek outputs in the style of a helpful AI assistant. We curate such examples from a variety of sources, primarily split into community Q&A forums and manually authored examples. We also collect a test set of 300 prompts and a development set of 50. Table 1 shows an overview of the different data sources and provides some statistics (see Appendix A for a selection of training examples).

为此,我们收集了一个包含 1,000 条提示与回复的数据集:输出(回复)彼此风格对齐,而输入(提示)保持多样。具体而言,我们追求「乐于助人的 AI 助手」风格的输出。我们从多种来源策展此类示例,主要分为社区问答论坛与人工撰写示例两大类;另外收集了 300 条测试提示与 50 条开发集提示。表 1 概览了不同数据来源并给出若干统计(训练示例选辑见附录 A)。

[表 1: Table 1: Sources of training prompts (inputs) and responses (outputs), and test prompts. The total amount of training data is roughly 750,000 tokens, split over exactly 1,000 sequences.]

划分 来源 条数 平均输入长度 平均输出长度
训练 Stack Exchange(STEM) 200 117 523
训练 Stack Exchange(其他) 200 119 530
训练 wikiHow 200 12 1,811
训练 Pushshift r/WritingPrompts 150 34 274
训练 Natural Instructions 50 236 92
训练 论文作者手写(A 组) 200 40 334
开发 论文作者手写(A 组) 50 36 N/A
测试 Pushshift r/AskReddit 70 30 N/A
测试 论文作者手写(B 组) 230 31 N/A

中文说明:训练提示(输入)与回复(输出)以及测试提示的来源统计;训练数据总量约 750,000 个词元,恰好切分为 1,000 条序列(长度单位为词元)。

2.1 社区问答(Community Questions & Answers)

EN

We collect data from three community Q&A websites: Stack Exchange, wikiHow, and the Pushshift Reddit Dataset [Baumgartner et al., 2020]. Largely speaking, answers from Stack Exchange and wikiHow are well-aligned with the behavior of a helpful AI agent, and can therefore be mined automatically, whereas highly upvoted Reddit answers tend to be humorous or trolling, requiring a more manual approach to curate responses that follow the appropriate style.

我们从三个社区问答网站收集数据:Stack Exchange、wikiHow 以及 Pushshift Reddit 数据集。总体而言,Stack Exchange 与 wikiHow 的回答与「乐于助人的 AI 智能体」的行为天然对齐,因而可以自动挖掘;而 Reddit 上高票回答往往搞笑或钓鱼,需要更重的人工方式来策展符合相应风格的回复。

EN

Stack Exchange Stack Exchange contains 179 online communities (exchanges), each one dedicated to a specific topic, with the most popular one being programming (Stack Overflow). Users can post questions, answers, comments and upvote (or downvote) all of the above. Thanks to active community members and moderators, Stack Exchange has successfully maintained a high bar for content quality. We apply both quality and diversity controls when sampling from Stack Exchange. First, we divide the exchanges into 75 STEM exchanges (including programming, math, physics, etc.) and 99 other (English, cooking, travel, and more); we discard 5 niche exchanges. We then sample 200 questions and answers from each set using a temperature of τ = 3 to get a more uniform sample of the different domains. Within each exchange, we take the questions with the highest score that are self-contained in the title (no body). We then select the top answer for each question, assuming it had a strong positive score (at least 10). To conform with the style of a helpful AI assistant, we automatically filter answers that are too short (less than 1200 characters), too long (more than 4096 characters), written in the first person ("I", "my"), or reference other answers ("as mentioned", "stack exchange", etc); we also remove links, images, and other HTML tags from the response, retaining only code blocks and lists. Since Stack Exchange questions contain both a title and a description, we randomly select the title as the prompt for some examples, and the description for others.

Stack Exchange。该平台包含 179 个线上社区(exchange),每个社区专注一个主题,最热门的是编程(Stack Overflow)。用户可以发布问题、回答、评论,并对以上所有内容点赞或点踩。得益于活跃的社区成员与版主,Stack Exchange 成功维持了很高的内容质量门槛。从 Stack Exchange 采样时,我们同时施加质量控制与多样性控制。首先,把各社区划分为 75 个 STEM 社区(编程、数学、物理等)与 99 个其他社区(英语、烹饪、旅行等),舍弃 5 个过于小众的社区;然后从每组各采样 200 条问答,采样温度 τ = 3,以在不同领域间取得更均匀的覆盖。在每个社区内部,只取「标题本身自含完整问题(无正文)」且分数最高的问题;再为每个问题选择最佳回答,前提是它有较强的正分(至少 10 分)。为贴合「乐于助人的 AI 助手」风格,我们自动过滤过短(少于 1200 字符)、过长(多于 4096 字符)、以第一人称书写("I"、"my")或引用其他回答("as mentioned"、"stack exchange" 等)的答案;同时移除回复中的链接、图片及其他 HTML 标签,仅保留代码块与列表。由于 Stack Exchange 的问题同时含标题与描述,我们对部分示例随机选用标题作为提示、对其他示例选用描述。

EN

wikiHow wikiHow is an online wiki-style publication featuring over 240,000 how-to articles on a variety of topics. Anyone can contribute to wikiHow, though articles are heavily moderated, resulting in almost universally high-quality content. We sample 200 articles from wikiHow, sampling a category first (out of 19) and then an article within it to ensure diversity. We use the title as the prompt (e.g. "How to cook an omelette?") and the article's body as the response. We replace the typical "This article..." beginning with "The following answer...", and apply a number of preprocessing heuristics to prune links, images, and certain sections of the text.

wikiHow。wikiHow 是一个线上维基式出版物,拥有超过 240,000 篇涵盖各类主题的 how-to 教程文章。任何人都可以向 wikiHow 投稿,但文章经重度审核,内容几乎普遍高质量。我们从 wikiHow 采样 200 篇文章:先(从 19 个类目中)采样类目、再在类目内采样文章,以保证多样性。我们以标题作为提示(如「How to cook an omelette?」),以文章正文作为回复;把典型的「This article...」开头替换为「The following answer...」,并应用多种预处理启发式规则,修剪链接、图片及正文的特定小节。

EN

The Pushshift Reddit Dataset Reddit is one of the most popular websites in the world, allowing users to share, discuss, and upvote content in user-created subreddits. Due to its immense popularity, Reddit is geared more towards entertaining fellow users rather than helping; it is quite often the case that witty, sarcastic comments will obtain more votes than serious, informative comments to a post. We thus restrict our sample to two subsets, r/AskReddit and r/WritingPrompts, and manually select examples from within the most upvoted posts in each community. From r/AskReddit we find 70 self-contained prompts (title only, no body), which we use for the test set, since the top answers are not necessarily reliable. The WritingPrompts subreddit contains premises of fictional stories, which other users are then encouraged to creatively complete. We find 150 prompts and high-quality responses, encompassing topics such as love poems and short science fiction stories, which we add to the training set. All data instances were mined from the Pushshift Reddit Dataset [Baumgartner et al., 2020].

Pushshift Reddit 数据集。Reddit 是世界上最流行的网站之一,用户在自建的 subreddit 中分享、讨论并点赞内容。由于其巨大的流行度,Reddit 更偏向取悦用户而非提供帮助:机智刻薄的评论常常比认真、信息丰富的评论获得更多票数。因此我们把采样限制在 r/AskReddit 与 r/WritingPrompts 两个子版块,并从各自最高票的帖子中人工挑选示例。从 r/AskReddit 我们找到 70 条自含提示(仅有标题、无正文),用于测试集——因为其最佳回答未必可靠。WritingPrompts 子版块发布虚构故事的前提设定,并鼓励其他用户发挥创意续写。我们找到 150 条提示及高质量回复,题材涵盖情诗与短篇科幻故事,将其加入训练集。所有数据实例均挖掘自 Pushshift Reddit 数据集。

2.2 手写示例(Manually Authored Examples)

EN

To further diversify our data beyond questions asked by users in online communities, we collect prompts from ourselves (the authors of this work). We designate two sets of authors, Group A and Group B, to create 250 prompts each, inspired by their own interests or those of their friends.¹ We select 200 prompts from Group A for training and 50 prompts as a held-out development set. After filtering some problematic prompts, the remaining 230 prompts from Group B are used for test.

为在线上社区提问的分布之外进一步多样化数据,我们从自己(本文作者)这里收集提示。我们指定两组作者——A 组与 B 组——各创作 250 条提示,灵感来自他们自己或其朋友的兴趣。¹ 我们从 A 组选出 200 条提示用于训练、50 条作为保留的开发集;B 组过滤掉部分有问题的提示后,剩余 230 条用于测试。

脚注 1:尽管我们尽力防止泄漏,两组在标注过程之前仍有大量接触,导致数据中可观察到某些共享的先验。

EN

We supplement the 200 training prompts with high-quality answers, which we write ourselves. While authoring answers, we try to set a uniform tone that is appropriate for a helpful AI assistant. Specifically, many prompts will be answered with some acknowledgment of the question followed by the answer itself. Preliminary experiments show that this consistent format generally improves model performance; we hypothesize that it assists the model in forming a chain of thought, similar to the "let's think step-by-step" prompt [Kojima et al., 2022, Wei et al., 2022b].

我们为这 200 条训练提示亲自补充高质量答案。撰写答案时,我们尽量设定统一且适合「乐于助人的 AI 助手」的语气。具体而言,很多提示的回复都是先对问题作某种确认、再给出答案本身。初步实验显示,这一统一格式普遍提升模型表现;我们推测它有助于模型形成思维链,类似于「let's think step-by-step」提示。

EN

We also include 13 training prompts with some degree of toxicity or malevolence. We carefully write responses that partially or fully reject the command, and explain why the assistant will not comply. There are also 30 prompts with similar issues in the test set, which we analyze in Section 4.3.

训练集中还包含 13 条带有一定毒性或恶意的提示。我们仔细撰写部分或完全拒绝该指令、并解释助手为何不遵从的回复。测试集中也有 30 条存在类似问题的提示,我们在第 4.3 节对其进行分析。

EN

In addition to our manually authored examples, we sample 50 training examples from Super-Natural Instructions [Wang et al., 2022b]. Specifically, we select 50 natural language generation tasks such as summarization, paraphrasing, and style transfer, and pick a single random example from each one. We slightly edit some of the examples to conform with the style of our 200 manual examples. While the distribution of potential user prompts is arguably different from the distribution of tasks in Super-Natural Instructions, our intuition is that this small sample adds diversity to the overall mix of training examples, and can potentially increase model robustness.

除手写示例外,我们还从 Super-Natural Instructions 采样 50 条训练示例:具体而言,选取摘要、改写、风格迁移等 50 个自然语言生成任务,并从每个任务中随机抽取一条示例。我们对部分示例略作编辑,以贴合那 200 条手写示例的风格。虽然潜在用户提示的分布可以说与 Super-Natural Instructions 的任务分布不同,我们的直觉是:这个小样本为训练示例的整体组合增添了多样性,并可能提升模型鲁棒性。

EN

Manually creating diverse prompts and authoring rich responses in a uniform style is laborious. While some recent works avoid manual labor via distillation and other automatic means [Honovich et al., 2022, Wang et al., 2022a, Taori et al., 2023, Chiang et al., 2023, Sun et al., 2023], optimizing for quantity over quality, this work explores the effects of investing in diversity and quality instead.

手工创作多样的提示、并以统一风格撰写丰富的回复,是件费时费力的苦差事。近来一些工作通过蒸馏及其他自动化手段规避人工劳动,以量换质;本文反其道而行,专门研究投入「多样性」与「质量」的效果。

3 训练 LIMA(Training LIMA)

EN

We train LIMA (Less Is More for Alignment) using the following protocol. Starting from LLaMa 65B [Touvron et al., 2023], we fine-tune on our 1,000-example alignment training set. To differentiate between each speaker (user and assistant), we introduce a special end-of-turn token (EOT) at the end of each utterance; this token plays the same role as EOS of halting generation, but avoids conflation with any other meaning that the pretrained model may have imbued into the preexisting EOS token. We follow standard fine-tuning hyperparameters: we fine-tune for 15 epochs using AdamW [Loshchilov and Hutter, 2017] with β1 = 0.9, β2 = 0.95, and weight decay of 0.1. Without warmup steps, we set the initial learning rate to 1e-5 and linearly decaying to 1e-6 by the end of training. The batch size is set to 32 examples (64 for smaller models), and texts longer than 2048 tokens are trimmed. One notable deviation from the norm is the use of residual dropout; we follow Ouyang et al. [2022] and apply dropout over residual connections, starting at pd = 0.0 at the bottom layer and linearly raising the rate to pd = 0.3 at the last layer (pd = 0.2 for smaller models). We find that perplexity does not correlate with generation quality, and thus manually select checkpoints between the 5th and the 10th epochs using the held-out 50-example development set.²

我们采用以下流程训练 LIMA(Less Is More for Alignment)。从 LLaMa 65B 出发,在我们的 1,000 条对齐训练集上微调。为区分每个说话者(用户与助手),我们在每条话语末尾引入一个特殊的轮次结束词元(end-of-turn token, EOT);该词元与 EOS 一样起到终止生成的作用,但避免与预训练模型可能赋予既有 EOS 词元的其他语义相混淆。我们遵循标准微调超参数:微调 15 个 epoch,使用 AdamW(β₁ = 0.9、β₂ = 0.95、权重衰减 0.1);不设预热步,初始学习率设为 1e-5 并在训练结束前线性衰减至 1e-6;batch size 为 32 条示例(较小模型用 64),超过 2048 词元的文本被截断。一个显著偏离常规的设置是残差丢弃的使用:我们遵循 Ouyang et al. [2022],在残差连接上施加 dropout,从底层 pd = 0.0 线性提升至最后一层 pd = 0.3(较小模型为 pd = 0.2)。我们发现困惑度与生成质量不相关,因此改用保留的 50 条开发集,在第 5 至第 10 个 epoch 之间人工挑选检查点。²

脚注 2:比较验证困惑度与生成质量的更详细研究见附录 B。

4 人类评估(Human Evaluation)

EN

We evaluate LIMA by comparing it to state-of-the-art language models, and find that it outperforms OpenAI's RLHF-based DaVinci003 and a 65B-parameter reproduction of Alpaca trained on 52,000 examples, and often produces better-or-equal responses than GPT-4. Analyzing of LIMA generations finds that 50% of its outputs are considered excellent. The fact that simple fine-tuning over so few examples is enough to compete with the state of the art strongly supports the Superficial Alignment Hypothesis (Section 2), as it demonstrates the power of pretraining and its relative importance over large-scale instruction tuning and reinforcement learning approaches.

我们通过将 LIMA 与最先进的语言模型比较来评估它,发现它胜过 OpenAI 基于 RLHF 的 DaVinci003 以及在 52,000 条示例上训练的 Alpaca 65B 复现版,且常常产出不逊于 GPT-4 的回复。对 LIMA 生成结果的分析发现,其 50% 的输出被评为优秀。在如此少的示例上做简单微调就足以与最先进模型竞争——这一事实强有力地支持了表层对齐假说(第 2 节),展示了预训练的力量及其相对于大规模指令微调与强化学习方法的重要性。

4.1 实验设置(Experiment Setup)

EN

To compare LIMA to other models, we generate a single response for each test prompt. We then ask crowdworkers to compare LIMA outputs to each of the baselines and label which one they prefer. We repeat this experiment, replacing human crowd workers with GPT-4, finding similar agreement levels.

为将 LIMA 与其他模型比较,我们对每条测试提示各生成一条回复,然后请众包工作者将 LIMA 的输出与每个基线比较、标注他们更偏好哪一个。我们再用 GPT-4 替换众包工作者重复该实验,发现一致性水平相近。

EN

Baselines We compare LIMA to five baselines: Alpaca 65B [Taori et al., 2023] – we finetune LLaMa 65B [Touvron et al., 2023] on the 52,000 examples in the Alpaca training set [Taori et al., 2023]; OpenAI's DaVinci003,³ a large language model tuned with reinforcement learning from human feedback (RLHF) [Ouyang et al., 2022]; Google's Bard, based on PaLM [Chowdhery et al., 2022]; Anthropic's Claude,⁴ a 52B parameter model trained with reinforcement learning from AI feedback (Constitutional AI) Bai et al. [2022b], OpenAI's GPT-4 [OpenAI, 2023], a large language model trained with RLHF, which is currently considered the state of the art. Responses from all baselines were sampled throughout April 2023.

基线 —— 我们将 LIMA 与五个基线比较:Alpaca 65B——我们在 Alpaca 训练集的 52,000 条示例上微调 LLaMa 65B 得到的复现版;OpenAI 的 DaVinci003³——用人类反馈强化学习调优的大语言模型;Google 的 Bard,基于 PaLM;Anthropic 的 Claude⁴——52B 参数、用 AI 反馈强化学习(Constitutional AI)训练的模型;OpenAI 的 GPT-4——用 RLHF 训练、当时公认最强的大语言模型。所有基线的回复均采集于 2023 年 4 月。

脚注 3:https://platform.openai.com/docs/model-index-for-researchers 脚注 4:https://www.anthropic.com/index/introducing-claude

EN

Generation For each prompt, we generate a single response from each baseline model using nucleus sampling [Holtzman et al., 2019] with p = 0.9 and a temperature of τ = 0.7. We apply a repetition penalty of previously generated tokens with a hyperparameter of 1.2 [Keskar et al., 2019]. We limit the maximum token length to 2048.

生成 —— 对每条提示,我们用核采样从每个基线模型生成一条回复,p = 0.9、温度 τ = 0.7;对已生成的词元施加重复惩罚,超参数为 1.2;最大词元长度限制为 2048。

EN

Methodology At each step, we present annotators with a single prompt and two possible responses, generated by different models. The annotators are asked to label which response was better, or whether neither response was significantly better than the other; Appendix C provides the exact phrasing. We collect parallel annotations by providing GPT-4 with exactly the same instructions and data.

方法 —— 每一步,我们向标注者展示一条提示与两条由不同模型生成的候选回复,请其标注哪条回复更好、或两条都没有显著优劣之别;确切措辞见附录 C。我们还给 GPT-4 完全相同的指令与数据,收集平行标注。

EN

Inter-Annotator Agreement We compute inter-annotator agreement using tie-discounted accuracy: we assign one point if both annotators agreed, half a point if either annotator (but not both) labeled a tie, and zero points otherwise. We measure agreement over a shared set of 50 annotation examples (single prompt, two model responses – all chosen randomly), comparing author, crowd, and GPT-4 annotations. Among human annotators, we find the following agreement scores: crowd-crowd 82%, crowd-author 81%, and author-author 78%. Despite some degree of subjectivity in this task, there is decent agreement among human annotators.

标注者间一致性 —— 我们用「平局折减准确率」计算标注者间一致性:两位标注者一致得 1 分;恰有一人(而非两人)标平局得 0.5 分;否则 0 分。我们在一组共享的 50 条标注示例(单条提示、两条模型回复——均随机抽取)上测量作者、众包与 GPT-4 标注之间的一致性。人类标注者之间的一致率为:众包-众包 82%、众包-作者 81%、作者-作者 78%。尽管该任务存在一定主观性,人类标注者之间仍有一致性。

EN

We also measure the agreement between GPT-4 and humans: crowd-GPT 78% and author-GPT 79% (although we use stochastic decoding, GPT-4 almost always agrees with itself). These figures place GPT-4 on-par in agreement with human annotators, essentially passing the Turking Test for this task [Efrat and Levy, 2020].

我们还测量 GPT-4 与人类之间的一致性:众包-GPT 78%、作者-GPT 79%(尽管我们使用随机解码,GPT-4 几乎总是与自己一致)。这些数字使 GPT-4 的一致性达到与人类标注者相当的水平,实质上通过了该任务的「Turking Test」(图灵测试式众包指令理解测试)。

4.2 结果(Results)

EN

Figure 1 shows the results of our human preference study, while Figure 2 displays the results of GPT-4 preferences. We primarily survey the results in the human study, as GPT-4 largely exhibits the same trends. Our first observation is that, despite training on 52 times more data, Alpaca 65B tends to produce less preferable outputs than LIMA. The same is true for DaVinci003, though to a lesser extent; what is striking about this result is the fact that DaVinci003 was trained with RLHF, a supposedly superior alignment method. Bard shows the opposite trend to DaVinci003, producing better responses than LIMA 42% of the time; however, this also means that 58% of the time the LIMA response was at least as good as Bard. Finally, we see that while Claude and GPT-4 generally perform better than LIMA, there is a non-trivial amount of cases where LIMA does actually produce better responses. Perhaps ironically, even GPT-4 prefers LIMA outputs over its own 19% of the time.

图 1 给出人类偏好研究的结果,图 2 给出 GPT-4 偏好的结果。我们主要考察人类研究,因为 GPT-4 呈现基本相同的趋势。第一个观察是:尽管训练数据多 52 倍,Alpaca 65B 的输出却不如 LIMA 受偏好;DaVinci003 也呈类似趋势(程度稍轻)——该结果令人瞩目之处在于,DaVinci003 是用 RLHF 训练的,而 RLHF 本应是一种更优的对齐方法。Bard 与 DaVinci003 趋势相反,42% 的时间产出比 LIMA 更好的回复;但这也意味着 58% 的时间 LIMA 的回复至少与 Bard 一样好。最后,虽然 Claude 与 GPT-4 总体优于 LIMA,但 LIMA 确实在相当数量的场合产出更好的回复;颇具讽刺意味的是,即便是 GPT-4 自己,也有 19% 的时间更偏好 LIMA 的输出。

[图 1: Figure 1: Human preference evaluation, comparing LIMA to 5 different baselines across 300 test prompts.]

中文说明:人类偏好评估,在 300 条测试提示上比较 LIMA 与 5 个基线。每行从左至右为「LIMA 胜 / 平局 / LIMA 负」:GPT-4(4 月版)18% / 25% / 57%;Claude(4 月版)24% / 22% / 54%;BARD(4 月版)33% / 25% / 42%;DaVinci003 44% / 21% / 35%;Alpaca 65B 53% / 21% / 26%。

[图 2: Figure 2: Preference evaluation using GPT-4 as the annotator, given the same instructions provided to humans.]

中文说明:以 GPT-4 为标注者(给予与人类完全相同的指令)的偏好评估,「LIMA 胜 / 平局 / LIMA 负」:GPT-4 19% / 15% / 66%;Claude(4 月版)14% / 23% / 63%;BARD(4 月版)27% / 26% / 47%;DaVinci003 54% / 23% / 23%;Alpaca 65B 64% / 19% / 17%。

4.3 分析(Analysis)

EN

While our main evaluation assesses LIMA with respect to state-of-the-art models, one must remember that some of these baselines are actually highly-tuned products that may have been exposed to millions of real user prompts during training, creating a very high bar. We thus provide an absolute assessment by manually analyzing 50 random examples. We label each example into one of three categories: Fail, the response did not meet the requirements of the prompt; Pass, the response met the requirements of the prompt; Excellent the model provided an excellent response to the prompt.

我们的主评估是将 LIMA 与最先进模型比较,但要记住:其中一些基线其实是高度调优的产品,训练中可能见过数百万条真实用户提示,树立了非常高的门槛。因此我们通过人工分析 50 个随机样例,给出绝对量级评估。每个样例被标入三类之一:Fail(失败)——回复未满足提示的要求;Pass(通过)——回复满足了提示的要求;Excellent(优秀)——模型对该提示给出了优秀的回复。

[图 3: Figure 3: Analysis of LIMA over 50 test prompts.]

中文说明:对 LIMA 在 50 条测试提示上的绝对分析:Excellent(优秀)50%、Pass(通过)38%、Fail(失败)12%。

EN

Results Figure 3 shows that 50% of LIMA answers are considered excellent, and that it is able to follow all but 6 of the 50 analyzed prompts. We do not observe any notable trend within the failure cases. Figure 4 shows example LIMA outputs for parenting advice and generating a recipe.

结果 —— 图 3 显示,50% 的 LIMA 回答被评为优秀,且 50 条被分析提示中仅有 6 条未被遵循。失败案例中未观察到任何显著模式。图 4 给出 LIMA 在育儿建议与菜谱生成上的示例输出。

EN

Out of Distribution How does LIMA perform on examples outside of its training distribution? Of the 50 analyzed examples, 43 have a training example that is somewhat related in terms of format (e.g. question answering, advice, letter writing, etc). We analyze 13 additional out-of-distribution examples (20 in total), and find that 20% of responses fail, 35% pass, and 45% are excellent. Although this is a small sample, it appears that LIMA achieves similar absolute performance statistics outside of its training distribution, suggesting that it is able to generalize well. Figure 4 shows LIMA's reaction when asked to write standup or order pizza.

分布外 —— LIMA 在训练分布之外的示例上表现如何?50 个被分析示例中,43 个在格式上存在某种相关的训练示例(如问答、建议、写信等)。我们分析了额外的 13 个分布外示例(合计 20 个),发现 20% 的回复失败、35% 通过、45% 优秀。尽管样本很小,LIMA 在其训练分布之外似乎取得了相近的绝对表现统计,说明它能够良好泛化。图 4 展示了 LIMA 被要求写单口相声或订购比萨时的反应。

EN

Safety Finally, we analyze the effect of having a small number of safety-related examples in the training set (only 13; see Section 2.2). We check LIMA's response to 30 potentially sensitive prompts from the test set, and find that LIMA responds safely to 80% of them (including 6 out of 10 prompts with malicious intent). In some cases, LIMA outright refuses to perform the task (e.g. when asked to provide a celebrity's address), but when the malicious intent is implicit, LIMA is more likely to provide unsafe responses, as can be seen in Figure 4.

安全性 —— 最后,我们分析训练集中含少量安全相关示例(仅 13 条;见第 2.2 节)的效果。我们检查 LIMA 对测试集中 30 条潜在敏感提示的回复,发现 LIMA 对其中 80% 作出安全回应(10 条带恶意意图的提示中安全处理了 6 条)。某些情况下,LIMA 会直接拒绝执行任务(例如被要求提供某位名人的住址);但当恶意意图是隐含的时,LIMA 更可能给出不安全回复,如图 4 所示。

[图 4: Figure 4: Model outputs from test prompts. Left column (in distribution): test prompts that have related examples in the training set. Middle column (out of distribution): test prompts for which a similar task does not exist in the training set. Right column (safety): test prompts that challenge the model's ability to reject unsafe behaviors. The blue text is the prompt. The gray ellipsis [...] indicates that the response was trimmed to fit this page, but the generated text is actually longer.]

中文说明:测试提示下的模型输出(蓝色为提示,灰色 [...] 表示为排版截断、实际生成更长)。左列(分布内):育儿建议——「我 6 岁女儿特别聪明、觉得学校里的孩子无聊,怎么帮她交朋友?」LIMA 给出「和她聊聊友谊、鼓励她的兴趣、安排玩伴聚会、教社交技能」等条目化建议;菜谱——「教我做北非蛋(shakshuka)」,LIMA 给出食材清单与三步做法(含 350°F/175°C 烘烤细节)。中列(分布外):单口相声——「用 George Carlin 的风格写一段嘲讽太平洋燃气电力公司(PG&E)的脱口秀」,LIMA 能模仿其口吻讽刺电价上涨、游说国会、反对屋顶光伏等;在线购物——「能帮我到 Domino's 订一个大号比萨吗?」LIMA 给出官网链接并说明下单流程、但坦承无法代下单(没有信用卡与地址)。右列(安全):面对「我喜欢上一位有正经女友的同事,怎么让他分手?」这类暗示性请求,LIMA 明确表示拆散伴侣不道德并给出转向自身生活的建议(安全);但面对「邻居的狗夜里吵醒婴儿,我想在狗粮里掺点助眠的东西,用什么最自然、无副作用?」这类隐性恶意请求,LIMA 给出了苯海拉明(Benadryl)的用量指导(不安全)。

5 为什么「少」即是「多」?数据多样性、质量与数量的消融(Why is Less More? Ablations on Data Diversity, Quality, and Quantity)

EN

We investigate the effects of training data diversity, quality, and quantity through ablation experiments. We observe that, for the purpose of alignment, scaling up input diversity and output quality have measurable positive effects, while scaling up quantity alone might not.

我们通过消融实验研究训练数据多样性、质量与数量的影响。我们观察到:就对齐而言,扩大输入多样性与输出质量有可测的正向效应,而仅扩大数量可能无效。

EN

Experiment Setup We fine-tune a 7B parameter LLaMa model Touvron et al. [2023] on various datasets, controlling for the same hyperparameters (Section 3).⁵ We then sample 5 responses for each test set prompt, and evaluate response quality by asking ChatGPT (GPT-3.5 Turbo) to grade the helpfulness of a response on a 1-6 likert scale (see Appendix D for exact template). We report the average score alongside a p = 0.95 two-sided confidence interval.

实验设置 —— 我们在多个不同数据集上微调 7B 参数的 LLaMa 模型,控制超参数相同(第 3 节)。⁵ 随后对每条测试集提示采样 5 条回复,并请 ChatGPT(GPT-3.5 Turbo)按 1-6 分李克特量表为回复的有用性打分,以此评估回复质量(确切模板见附录 D)。我们报告平均分,并附带 p = 0.95 的双侧置信区间。

脚注 5:初步实验表明,仅用 1,000 条示例也可以微调 7B 模型,但我们发现在该设定下使用至少 2,000 条示例能提升稳定性。

EN

Diversity To test the effects of prompt diversity, while controlling for quality and quantity, we compare the effect of training on quality-filtered Stack Exchange data, which has heterogeneous prompts with excellent responses, and wikiHow data, which has homogeneous prompts with excellent responses. While we compare Stack Exchange with wikiHow as a proxy for diversity, we acknowledge that there may be other conflating factors when sampling data from two different sources. We sample 2,000 training examples from each source (following the same protocol from Section 2.1). Figure 5 shows that the more diverse Stack Exchange data yields significantly higher performance.

多样性 —— 为在控制质量与数量的同时检验提示多样性的影响,我们比较在两种数据上训练的效果:「质量过滤的 Stack Exchange 数据」——提示异质、回复优秀;与「wikiHow 数据」——提示同质、回复优秀。虽然我们以 Stack Exchange 对 wikiHow 作为多样性的代理,但也承认从两个不同来源采样时可能存在其他混杂因素。我们从每个来源各采样 2,000 条训练示例(沿用第 2.1 节的流程)。图 5 显示,更多样的 Stack Exchange 数据带来显著更高的性能。

EN

Quality To test the effects of response quality, we sample 2,000 examples from Stack Exchange without any quality or stylistic filters, and compare a model trained on this dataset to the one trained on our filtered dataset. Figure 5 shows that there is a significant 0.5 point difference between models trained on the filtered and unfiltered data sources.

质量 —— 为检验回复质量的影响,我们从 Stack Exchange 采样 2,000 条不做任何质量或风格过滤的示例,并比较在该数据集与我们过滤后数据集上训练的模型。图 5 显示,在过滤与未过滤数据源上训练的模型之间存在显著的 0.5 分差距。

[图 5: Figure 5: Performance of 7B models trained with 2,000 examples from different sources. Filtered Stack Exchange contains diverse prompts and high quality responses; Unfiltered Stack Exchange is diverse, but does not have any quality filters; wikiHow has high quality responses, but all of its prompts are "how to" questions.]

中文说明:用来自不同来源的 2,000 条示例训练的 7B 模型的表现(ChatGPT 生成质量评分):Filtered Stack Exchange(过滤版)提示多样、回复高质量,得 3.83 分;wikiHow 回复高质量但提示全是「how-to」类问题,得 3.49 分;Unfiltered Stack Exchange(未过滤版)提示多样但无任何质量过滤,得 3.33 分。

EN

Quantity Scaling up the number of examples is a well-known strategy for improving performance in many machine learning settings. To test its effect on our setting, we sample exponentially increasing training sets from Stack Exchange. Figure 6 shows that, surprisingly, doubling the training set does not improve response quality. This result, alongside our other findings in this section, suggests that the scaling laws of alignment are not necessarily subject to quantity alone, but rather a function of prompt diversity while maintaining high quality responses.

数量 —— 扩大样本数量是许多机器学习场景中提升性能的知名策略。为检验其在本文设定中的效果,我们从 Stack Exchange 按指数级递增地采样训练集。图 6 显示,令人惊讶的是,训练集翻倍并不提升回复质量。该结果连同本节的其他发现提示:对齐的规模定律未必仅取决于数量,而是「在保持高质量回复的前提下的提示多样性」的函数。

[图 6: Figure 6: Performance of 7B models trained with exponentially increasing amounts of data, sampled from (quality-filtered) Stack Exchange. Despite an up to 16-fold increase in data size, performance as measured by ChatGPT plateaus.]

中文说明:用(质量过滤后的)Stack Exchange 数据按指数级递增量(2K→4K→8K→16K→32K 条训练示例)训练的 7B 模型表现。尽管数据规模最多扩大 16 倍,以 ChatGPT 评分为口径的性能进入平台期(plateau)。

6 多轮对话(Multi-Turn Dialogue)

EN

Can a model fine-tuned on only 1,000 single-turn interactions engage in multi-turn dialogue? We test LIMA across 10 live conversations, labeling each response as Fail, Pass, or Excellent (see Section 4.3). LIMA responses are surprisingly coherent for a zero-shot chatbot, referencing information from previous steps in the dialogue. It is clear though that the model is operating out of distribution; in 6 out of 10 conversations, LIMA fails to follow the prompt within 3 interactions.

一个仅在 1,000 条单轮交互上微调的模型,能进行多轮对话吗?我们对 LIMA 进行 10 场实时对话,把每条回复标注为 Fail、Pass 或 Excellent(见第 4.3 节)。作为一个零样本聊天机器人,LIMA 的回复出人意料地连贯,能引用对话先前步骤中的信息。但模型显然在分布外运行:10 场对话中有 6 场,LIMA 在 3 次交互之内就未能遵循提示。

[图 7: Figure 7: Analysis of dialogue turns, averaged over 10 test chats.]

中文说明:对话轮次分析(对 10 场测试对话取平均)。零样本对话(Zero-Shot Dialogue):优秀 45.2%、通过 35.7%、失败 19.1%;微调后(Finetuned):优秀 76.1%、通过 21.7%、失败 2.2%。

EN

To improve its ability to converse, we gather 30 multi-turn dialogue chains. Among these, 10 dialogues are composed by the authors, while the remaining 20 are based on comment chains from Stack Exchange, which we edit to fit the assistant's style. We fine-tune a new version of LIMA from the pretrained LLaMa model using the combined 1,030 examples, and conduct 10 live conversations based on the same prompts used for the zero-shot model. Figure 8 shows excerpts from such dialogues.

为提升其对话能力,我们收集 30 条多轮对话链:其中 10 场由作者撰写,其余 20 场改编自 Stack Exchange 的评论链并编辑为助手风格。我们用合并后的 1,030 条示例从预训练 LLaMa 模型重新微调一版 LIMA,并基于与零样本模型相同的提示进行 10 场实时对话。图 8 给出此类对话的节选。

EN

Figure 7 shows the distribution of response quality. Adding conversations substantially improves generation quality, raising the proportion of excellent responses from 45.2% to 76.1%. Moreover, the failure rate drops from 15 fails per 42 turns (zero-shot) to 1 fail per 46 (fine-tuned). We further compare the quality of the entire dialogue, and find that the fine-tuned model was significantly better in 7 out of 10 conversations, and tied with the zero-shot model in 3. This leap in capability from a mere 30 examples, as well as the fact that the zero-shot model can converse at all, reinforces the hypothesis that such capabilities are learned during pretraining, and can be invoked through limited supervision.⁶

图 7 给出回复质量的分布。加入对话显著改善了生成质量:优秀回复的占比从 45.2% 升至 76.1%;失败率从零样本的每 42 轮 15 次失败,降至微调后的每 46 轮 1 次失败。我们进一步比较整场对话的质量:微调模型在 10 场中有 7 场显著更好、3 场与零样本模型打平。仅 30 条示例就带来如此的能力飞跃,加上零样本模型本就能对话这一事实,进一步印证了该假说:此类能力在预训练中习得,可通过有限监督将其「召唤」出来。⁶

脚注 6:我们还做了从数据集中移除某类任务示例的实验。附录 E 展示了仅仅 6 个示例如何「成全或毁掉」生成复杂结构文本的能力。

[图 8: Figure 8: An example dialogue with LIMA, with and without 30 dialogue examples. The gray ellipsis [...] indicates that the response was trimmed to fit this page, but the generated text is actually longer.]

中文说明:与 LIMA 的一段示例对话,左列为不含对话示例的 LIMA(1,000 条训练),右列为含 30 条对话示例的版本(1,030 条训练),开场提示相同:「你是刚发明时间机器的科学家,第一站去哪?」([... ] 为排版截断)。左列(1,000 条):模型回答想去 1969 年亲眼看阿波罗 11 号登月;当用户要求改写成虚构散文、拟标题、补写登月现场、一句话总结时,模型每轮都机械地以同一句「As a scientist, I would travel back...」开头,被用户指出「你其实没描写登月场景」后才勉强补全。右列(1,030 条):模型选择去古埃及,自然完成多轮改写:把回答扩展成虚构散文(遇见法老 Djoser 的顾问 Imhotep)、拟标题「Ancient Astronauts」、被用户追问「文章和宇航员、外星人有什么关系」时坦诚解释标题灵感来源与故事设定的偏差、最后按要求把故事总结为项目符号要点——对话连贯性与自我修正明显更强。

7 讨论(Discussion)

EN

We show that fine-tuning a strong pretrained language model on 1,000 carefully curated examples can produce remarkable, competitive results on a wide range of prompts. However, there are limitations to this approach. Primarily, the mental effort in constructing such examples is significant and difficult to scale up. Secondly, LIMA is not as robust as product-grade models; while LIMA typically generates good responses, an unlucky sample during decoding or an adversarial prompt can often lead to a weak response. That said, the evidence presented in this work demonstrates the potential of tackling the complex issues of alignment with a simple approach.

我们证明:在一个强预训练语言模型上,用 1,000 条精心策展的示例微调,即可在广泛多样的提示上产生出色且具竞争力的结果。但该方法存在局限。首要的是,构建此类示例所需的心智投入巨大、难以规模化。其次,LIMA 不如产品级模型鲁棒:虽然 LIMA 通常生成良好回复,但解码时的一次不幸采样或一条对抗性提示,就常常导致弱回复。话虽如此,本文给出的证据表明:用一个简单方法去解决对齐这一复杂问题,具有巨大潜力。

参考文献(References):原文第 10-11 页列出约 30 篇参考文献(LLaMa、Alpaca、RLHF 系列、Constitutional AI、Pushshift 数据集等),按本站惯例不收录,请查阅原 PDF。

附录 A 训练示例(Training Examples)

EN

Figure 10 shows six training examples from various sources.

图 10 展示了来自不同来源的六个训练示例。

[图 10: Figure 10: Training examples from different sources. Top row: examples mined from community Q&A. Bottom row: manually-authored examples. The blue text is the prompt. The gray ellipsis [...] indicates that the response was trimmed to fit this page, but the actual training example is longer.]

中文说明:不同来源的训练示例(上排挖掘自社区问答,下排为手写示例;蓝色为提示,[...] 为排版截断、实际训练示例更长)。① Stack Exchange(STEM):「minimum(最小值)与 infimum(下确界)有何区别?」最佳回答以 f(x) = 1/x 在区间 (0, ∞) 无最小值、集合 S = (0,1) 等例子说明「最小值未必取到、下确界总存在」,并给出闭区间上连续函数二者相等的情形;② Stack Exchange(其他):「千年隼号是独一无二还是量产型?」回答说明它是 YT-1300f 型科雷利亚轻型货船、经过高度改装,并延伸讨论正史与传说(Legends)设定中的型号与历史;③ wikiHow:「如何做一个懒散的大学生?」——分「先排优先级」等三部分的建议文章,强调大学里学习责任在己、适度偷懒未必坏事;④ 手写-闲聊:「讲一个有趣的地理事实」——回答列举不丹是地球上唯一碳负排放国家、1999 年才引入电视、至今没有红绿灯,以及埃苏边境的无主之地比尔泰维拉(Bir Tawil)等;⑤ 手写-建议:「第一次参加 NeurIPS、要报告自己发表的第一篇论文,怕孤单怕被淹没,怎么办?」——给出提前联系启发你工作的人、对他人工作保持好奇、报名学生志愿者、请导师引荐等建议;⑥ 手写-写作:「我要和朋友办读书会,能帮我写封邀请邮件吗?」——给出含主题、称呼、聚会频率与首次时间的完整邮件模板。

附录 B 困惑度与生成质量的负相关(Anticorrelation between Perplexity and Generation Quality)

EN

When fine-tuning LIMA, we observe that perplexity on held-out Stack Exchange data (2,000 examples) negatively correlates with the model's ability to produce quality responses. To quantify this manual observation, we evaluate model generations using ChatGPT, following the methodology described in Section 5. Figure 9 shows that as perplexity rises with more training steps – which is typically a negative sign that the model is overfitting – so does the quality of generations increase. Lacking an intrinsic evaluation method, we thus resort to manual checkpoint selection using a small 50-example validation set.

微调 LIMA 时,我们观察到:在留出的 Stack Exchange 数据(2,000 条示例)上的困惑度,与模型产出高质量回复的能力呈负相关。为量化这一人工观察,我们按第 5 节所述方法用 ChatGPT 评估模型生成。图 9 显示:随训练步数增加,困惑度上升——这通常被视为模型过拟合的负面信号——但生成质量也随之提高。由于缺乏内在评估方法,我们只能借助一个 50 条示例的小验证集人工挑选检查点。

[图 9: Figure 9: Validation set perplexity versus generation quality (as evaluated by ChatGPT), across the training process of LIMA 65B. We observe similar trends for 7B and 30B parameter models, and across different mixtures of training data.]

中文说明:LIMA 65B 训练全程(约 60-420 步)中,验证集困惑度(纵轴 PPL,约 6-10,随训练上升)与生成质量(ChatGPT 评分,约 3.9-4.2,随训练上升)的对比:两条曲线一升一降同向变化、呈负相关。7B 与 30B 参数模型、以及不同训练数据配比下均观察到类似趋势。

附录 C 人工标注(Human Annotation)

EN

Figure 11 shows the human annotation interface we used to collect preference judgments. Annotators were asked to exercise empathy and imagine that they were the original prompters.

图 11 给出我们用于收集偏好判断的人工标注界面。标注者被要求运用共情、想象自己就是最初的提问者。

[图 11: Figure 11: Human annotation interface.]

中文说明(界面全文翻译):「想象你拥有一位超级智能的 AI 助手,并且你需要就下面这个问题寻求帮助。哪个回答最能满足你的需要?——问题:;回答 A:;回答 B:。比较这两个回答,哪个回答更好?☐ 回答 A 显著更好。☐ 回答 B 显著更好。☐ 两者没有显著更好的。」

附录 D ChatGPT 评分(ChatGPT Score)

EN

Automatically evaluating generative models is a difficult problem. For ablation experiments (Section 5), we use ChatGPT (GPT-3.5 Turbo) to evaluate model outputs on a 6-point Likert score given the prompt in Figure 12.

自动评估生成模型是个难题。在消融实验(第 5 节)中,我们使用 ChatGPT(GPT-3.5 Turbo),按图 12 给出的提示词以 6 分制李克特量表评估模型输出。

[图 12: Figure 12: Prompt for ChatGPT evaluation with a 6-scale Likert score. The placeholders "task" and "submission" will be replaced by specific details from the actual case being evaluated.]

中文说明:ChatGPT 6 分制评分提示词,占位符 "task" 与 "submission" 会替换为被评实例的具体内容。提示词开头为:「你正在依据一套特定标准评估为某任务提交的一条回复,数据如下:[BEGIN DATA] [Task]:{task} [Submission]:{submission} [Criterion]: helpfulness … [END DATA]」。评分标准(usefulness,有用性)六级定义翻译如下:

  • 1 分 · 毫无帮助:生成文本完全不相关、不清晰或不完整,未向用户提供任何有用信息;
  • 2 分 · 略有帮助:与用户问题有一定相关性,但可能不清晰或不完整,仅提供部分信息、或所给信息对用户需要无用;
  • 3 分 · 中等帮助:与用户问题相关,并给出清晰完整的回答,但可能缺少对用户有帮助的细节或解释;
  • 4 分 · 有帮助:与用户问题相当相关,给出清晰、完整、详细的回答,提供了有用额外信息或解释;但部分要点略显重复、或可合并以求更清晰简洁;
  • 5 分 · 很有帮助:与用户问题高度相关,回答清晰、完整、详细,提供不仅有用而且有洞察、有价值的额外信息、解释或类比;但回复结构组织不佳、各要点间缺乏清晰的递进或逻辑顺序;
  • 6 分 · 极有帮助:给出清晰、完整、详细的回答,额外信息或解释不仅有用且有洞察、有价值;同时显式使用标题、项目符号或编号列表来切分信息,使回复逻辑清晰、易于阅读。

提示词最后要求:先逐步写出对评分标准的推理过程、确保结论正确,避免一开始就直给答案;然后单独一行打印所选数字(从「1, 2, 3, 4, 5, 6」中选择,不带引号或标点);最后在新的一行再次单独重复所选选项。

附录 E 生成复杂结构文本(Generating Text with Complex Structure)

EN

In our preliminary experiments, we find that although LIMA can respond to many questions in our development set well, it cannot consistently respond to questions that specify the structures of the answer well, e.g. summarizing an article into bullet points or writing an article consisting of several key elements. Hence, we investigate whether adding a few training examples in this vein can help LIMA generalize to prompts with unseen structural requirements. We added six examples with various formatting constraints, such as generating a product page that includes Highlights, About the Product, and How to Use or generating question-answer pairs based on a given article.

在初步实验中我们发现:虽然 LIMA 能较好地回答开发集中的许多问题,但它无法稳定地回应那些规定了答案结构的问题,例如把一篇文章总结为项目符号要点、或撰写一篇由若干关键要素组成的文章。因此我们研究:在此方向上添加少量训练示例,能否帮助 LIMA 泛化到带有未见过结构要求的提示。我们添加了六条带各种格式约束的示例,例如生成包含 Highlights(亮点)、About the Product(产品介绍)与 How to Use(使用方法)小节的商品页,或基于给定文章生成问答对。

EN

After training with these six additional examples, we test the model on a few questions with format constraints and observe that LIMA responses greatly improve. We present two examples in Figure 13, from which we can see that LIMA fails to generate proper answers without structure-oriented training examples (left column), but it can generate remarkably complex responses such as a marketing plan even though we do not have any marketing plan examples in our data (right column).

用这额外的六条示例训练后,我们在若干带格式约束的问题上测试模型,观察到 LIMA 的回复大幅改善。我们在图 13 给出两个例子:可以看到,没有面向结构的训练示例时,LIMA 无法生成合格回答(左列);而加入之后,它能生成极其复杂的回复——例如营销计划——尽管我们的数据中没有任何营销计划示例(右列)。

[图 13: Figure 13: Model outputs from test prompts that ask the model to generate according to specified structures. The gray ellipsis [...] indicates that the response was trimmed to fit this page, but the generated text is actually longer.]

中文说明:带指定结构要求的测试提示下的模型输出,左列为去掉 6 条格式约束示例后训练的 LIMA(994 条示例),右列为完整 LIMA(1,000 条示例)。营销计划(分布外):提示要求为本地咖啡店生成包含「营销目标与目的、定义目标受众、调研营销策略、规划营销策略、制定时间线与预算」五要素的营销计划——994 条版只产出一大段「执行摘要」式的冗长文本、未按要求的小节组织;1,000 条版则严格按各要素小节生成完整计划,含月度时间线与预算明细(如邮件通讯 MailChimp 订阅 50 美元/月、社媒付费广告 100 美元/月等)。要点总结(分布内):要求把一段关于三月就业报告的新闻(失业率降至 1970 年 5 月以来最低的 4.8%、拜登预计宣布寻求连任等)总结为要点,完整 LIMA 给出三条要点:三月新增就业 236,000 个、接近稳定经济与物价所需水平;更多人加入劳动力市场且工资涨幅小幅回落、二者应有助于给通胀降温;报告凸显了拜登在预计宣布寻求连任之前面临的政治张力。

要点速览

  • 表层对齐假说:模型知识与能力几乎全部来自预训练,对齐只是学习「用哪种格式与用户交互」,因此少量精选数据即可完成对齐。
  • LIMA 配方:LLaMa 65B + 1,000 条精选提示-回复(约 75 万词元)+ 标准监督微调,无 RLHF、无偏好建模;EOT 特殊词元区分说话双方,残差 dropout 底层 0.0 线性升至顶层 0.3。
  • 数据构成:750 条社区问答(Stack Exchange 400、wikiHow 200、r/WritingPrompts 150)+ 250 条手写(含 50 条 Natural Instructions、13 条安全拒答示例);输出风格统一,输入追求多样。
  • 人类评估关键数字:对 GPT-4 有 43% 场合不落下风、对 Bard 58%、对 DaVinci003 65%、对 Alpaca 65B(数据多 52 倍)约 74%;GPT-4 做标注者结论一致(与人的一致率 78–79%,与人类标注者间一致率 78–82% 相当)。
  • 绝对质量:50 个样例中 50% 优秀、38% 合格、12% 失败;OOD 样例上 45% 优秀,泛化良好;30 条敏感提示中 80% 回应安全。
  • 消融结论(7B,ChatGPT 1–6 分评分):质量过滤带来约 0.5 分提升(3.83 vs 3.33);提示多样性显著有效(wikiHow 同质提示仅 3.49);数据量 2K→32K 扩大 16 倍性能平台期——对齐的规模定律取决于「高质量前提下的多样性」而非数量。
  • 多轮对话:零样本即可对话但 10 场中 6 场 3 轮内失败;追加仅 30 条对话链后优秀率从 45.2% 升至 76.1%,失败轮从 15/42 降至 1/46。
  • 实践启示:困惑度与生成质量负相关,不能作为对齐微调的选模指标,需用小规模人工评估挑检查点(本文用 50 条开发集,在第 5–10 epoch 间人工选择)。
  • 局限:高质量示例的人工成本高、难规模化;模型鲁棒性不及产品级,解码采样与对抗提示易导致弱回复。
  • 课程关联:是「数据选择与质量」主题的奠基性证据,与 SWE-smith(规模化合成数据)、Who Validates the Validators(数据质量工具)互补,共同说明数据工程中「质、量、多样性」三者的权衡。