<?xml version='1.0' encoding='utf-8'?>
<feed xmlns="http://www.w3.org/2005/Atom"><title>Malladi Lab — Writing</title><id>https://sadhikamalladi.github.io/feed.xml</id><link href="https://sadhikamalladi.github.io/feed.xml" rel="self" type="application/atom+xml" /><link href="https://sadhikamalladi.github.io/blog/" rel="alternate" /><updated>2026-10-07T18:14:59.258963+00:00</updated><entry><title>The Hidden Infinity in Preference Learning</title><id>https://sadhikamalladi.github.io/blog/2024/07/09/dpo-infinity/</id><link href="https://sadhikamalladi.github.io/blog/2024/07/09/dpo-infinity/" rel="alternate" type="text/html" /><published>2024-07-09T00:00:00-07:00</published><updated>2024-07-09T00:00:00-07:00</updated><author><name>Sadhika Malladi</name></author><summary>An illustration of how length normalization aids learning from model-annotated data</summary><content type="html">
        &lt;p&gt;
          &lt;strong&gt;TL;DR&lt;/strong&gt;: I demonstrate from first principles how
          offline preference learning algorithms (e.g.,
          &lt;a href="https://arxiv.org/abs/2405.14734"&gt;SimPO&lt;/a&gt;) can benefit from
          length normalization, especially when training on model-annotated
          preference data. The derivation also lends insight into a subtle but
          important challenge in training LMs: infinite strings.
        &lt;/p&gt;

        &lt;hr /&gt;

        &lt;p&gt;
          Aligning language models using human feedback has become increasingly
          popular. The process of collecting this feedback (i.e., a human rater
          reading and ranking pieces of text) is non-differentiable, so people
          turned to reinforcement learning (RL).
        &lt;/p&gt;
        &lt;h1 id="learning-from-human-feedback-as-rl"&gt;
          Learning from Human Feedback as RL
        &lt;/h1&gt;
        &lt;p&gt;
          It turns out that it is confusing to translate the problem of
          improving language models via human feedback to the RL setting. First
          question: how should we define the states and the actions?
        &lt;/p&gt;

        &lt;p&gt;
          Originally,
          &lt;a href="https://proceedings.mlr.press/v70/jaques17a.html"
            &gt;Jaques et al., 2017&lt;/a
          &gt;
          designed the states in the problem to be the partial context, usually
          some combination of pre-filled prompt tokens and generated tokens, so
          the action was to pick a single token to generate next. I’ll refer to
          this as the &lt;strong&gt;token-level formulation&lt;/strong&gt;. On the other
          hand, later works (&lt;a href="https://arxiv.org/abs/1909.08593"
            &gt;Ziegler et al., 2019&lt;/a
          &gt;;
          &lt;a href="https://arxiv.org/abs/2009.01325"&gt;Stiennon et al., 2020&lt;/a&gt;)
          framed the problem in a bandit setting, where the state is the
          pre-filled prompt and the action is to select which sequence to
          generate out of a few candidates. I’ll refer to this as the
          &lt;strong&gt;bandit formulation&lt;/strong&gt;. Both of these designs are not
          perfectly suited to our setting, because we get human feedback on a
          sequence level but ask our LM to act on (i.e., generate) individual
          tokens. More recently, some works have designed objectives that
          incorporate token-level rewards or constraints alongside
          sequence-level feedback (&lt;a href="https://arxiv.org/abs/2404.11999"
            &gt;Zeng et al., 2024&lt;/a
          &gt;; &lt;a href="https://arxiv.org/abs/2406.18629"&gt;Lai et al., 2024&lt;/a&gt;).
          On the more theoretical and conceptual side,
          &lt;a href="https://arxiv.org/abs/2404.12358"&gt;Rafailov et al., 2024&lt;/a&gt;
          has a great illustration of this topic and draws meaningful
          connections between the two settings, and I’m sure that many RL papers
          are dedicated to parsing this gap between trajectory-level rewards and
          step-level actions.
        &lt;/p&gt;

        &lt;p&gt;
          The reason I bring up these two settings is to see how each one deals
          with the case of a model producing infinite-length strings. And yes, I
          know that it’s impossible in practice for LMs to actually generate an
          infinite-length string. But they can certainly put a tiny amount of
          probability mass on an infinite-length string. And if they do so for
          many such strings, then this ends up being a huge amount of
          probability mass lost to sequences that we’ll never see.
        &lt;/p&gt;

        &lt;p&gt;
          In the token-level formulation, it’s clear that an infinite-length
          string corresponds to essentially an infinite-length trajectory rolled
          out from the policy. One may think the bandit formulation somehow
          circumvents this problem because it provides only two bounded-length
          sequences for the model to choose from. But, when deriving optimal
          solutions to the bandit setting, you still need to deal with the
          problem of infinite “arms” (i.e., possible responses sampled from your
          policy). For example, Appendix A.1 of the
          &lt;a href="https://arxiv.org/abs/2305.18290"&gt;DPO paper&lt;/a&gt; glosses over
          the fact that the partition function $Z_\theta$ may not exist if you
          don’t bound the length of the generations or somehow assume that the
          reward decays with length. If you’re still not convinced that
          infinite-length strings matter, then stay tuned to see how this idea
          eventually plays out into a rigorous understanding of the role of
          length normalization.
        &lt;/p&gt;

        &lt;h1 id="length-normalization"&gt;Length Normalization&lt;/h1&gt;
        &lt;p&gt;
          We can now derive how length normalization might make sense from the
          RL perspective. My disclaimer here is that I definitely worked
          backwards from the idea of length normalization to arrive at this
          justification, so there may be many other algorithms that are equally
          or better suited to operating in the setting I described. This
          derivation is also hand-wavy and meant more to provide intuition more
          than anything else. I’ll use the bandit-level formulation.
        &lt;/p&gt;

        &lt;p&gt;
          Suppose we have a prompt $x$ and two candidate completions $y_1$ and
          $y_2$. We will use $\Pr[y_1\succ y_2\mid x]$ to denote the proportion
          of people who prefer $y_1$ over $y_2$ as a completion to $x$.
        &lt;/p&gt;

        &lt;p&gt;
          The first thing we need to do is make some assumption about the reward
          structure. The Bradley-Terry (BT) model was derived to rigorously
          describe the distribution of human preferences:
        &lt;/p&gt;

        &lt;p&gt;
          &lt;span class="katex-display"
            &gt;&lt;span class="katex"
              &gt;&lt;span class="katex-mathml"
                &gt;&lt;math xmlns="http://www.w3.org/1998/Math/MathML"
                  &gt;&lt;semantics
                    &gt;&lt;mrow
                      &gt;&lt;mi&gt;Pr&lt;/mi&gt;&lt;mo&gt;⁡&lt;/mo&gt;&lt;mo stretchy="false"&gt;[&lt;/mo
                      &gt;&lt;msub&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/msub
                      &gt;&lt;mo&gt;≻&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mn&gt;2&lt;/mn&gt;&lt;/msub
                      &gt;&lt;mo&gt;∣&lt;/mo&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mo stretchy="false"&gt;]&lt;/mo&gt;&lt;mo&gt;=&lt;/mo
                      &gt;&lt;mfrac
                        &gt;&lt;mrow
                          &gt;&lt;mi&gt;exp&lt;/mi&gt;&lt;mo&gt;⁡&lt;/mo&gt;&lt;mo stretchy="false"&gt;(&lt;/mo
                          &gt;&lt;msub&gt;&lt;mi&gt;r&lt;/mi&gt;&lt;mtext&gt;h&lt;/mtext&gt;&lt;/msub
                          &gt;&lt;mo stretchy="false"&gt;(&lt;/mo&gt;&lt;mi&gt;x&lt;/mi
                          &gt;&lt;mo separator="true"&gt;,&lt;/mo
                          &gt;&lt;msub&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/msub
                          &gt;&lt;mo stretchy="false"&gt;)&lt;/mo
                          &gt;&lt;mo stretchy="false"&gt;)&lt;/mo&gt;&lt;/mrow
                        &gt;&lt;mrow
                          &gt;&lt;mi&gt;exp&lt;/mi&gt;&lt;mo&gt;⁡&lt;/mo&gt;&lt;mo stretchy="false"&gt;(&lt;/mo
                          &gt;&lt;msub&gt;&lt;mi&gt;r&lt;/mi&gt;&lt;mtext&gt;h&lt;/mtext&gt;&lt;/msub
                          &gt;&lt;mo stretchy="false"&gt;(&lt;/mo&gt;&lt;mi&gt;x&lt;/mi
                          &gt;&lt;mo separator="true"&gt;,&lt;/mo
                          &gt;&lt;msub&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/msub
                          &gt;&lt;mo stretchy="false"&gt;)&lt;/mo&gt;&lt;mo stretchy="false"&gt;)&lt;/mo
                          &gt;&lt;mo&gt;+&lt;/mo&gt;&lt;mi&gt;exp&lt;/mi&gt;&lt;mo&gt;⁡&lt;/mo
                          &gt;&lt;mo stretchy="false"&gt;(&lt;/mo
                          &gt;&lt;msub&gt;&lt;mi&gt;r&lt;/mi&gt;&lt;mtext&gt;h&lt;/mtext&gt;&lt;/msub
                          &gt;&lt;mo stretchy="false"&gt;(&lt;/mo&gt;&lt;mi&gt;x&lt;/mi
                          &gt;&lt;mo separator="true"&gt;,&lt;/mo
                          &gt;&lt;msub&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mn&gt;2&lt;/mn&gt;&lt;/msub
                          &gt;&lt;mo stretchy="false"&gt;)&lt;/mo
                          &gt;&lt;mo stretchy="false"&gt;)&lt;/mo&gt;&lt;/mrow
                        &gt;&lt;/mfrac
                      &gt;&lt;/mrow
                    &gt;&lt;annotation encoding="application/x-tex"&gt;
                      \Pr[y_1\succ y_2\mid x] = \frac{\exp(r_\text{h}(x,
                      y_1))}{\exp(r_\text{h}(x, y_1)) + \exp(r_\text{h}(x,
                      y_2))}&lt;/annotation
                    &gt;&lt;/semantics
                  &gt;&lt;/math
                &gt;&lt;/span
              &gt;&lt;span class="katex-html" aria-hidden="true"
                &gt;&lt;span class="base"
                  &gt;&lt;span
                    class="strut"
                    style="height: 1em; vertical-align: -0.25em"
                  &gt;&lt;/span
                  &gt;&lt;span class="mop"&gt;Pr&lt;/span&gt;&lt;span class="mopen"&gt;[&lt;/span
                  &gt;&lt;span class="mord"
                    &gt;&lt;span
                      class="mord mathdefault"
                      style="margin-right: 0.03588em"
                      &gt;y&lt;/span
                    &gt;&lt;span class="msupsub"
                      &gt;&lt;span class="vlist-t vlist-t2"
                        &gt;&lt;span class="vlist-r"
                          &gt;&lt;span
                            class="vlist"
                            style="height: 0.30110799999999993em"
                            &gt;&lt;span
                              style="
                                top: -2.5500000000000003em;
                                margin-left: -0.03588em;
                                margin-right: 0.05em;
                              "
                              &gt;&lt;span class="pstrut" style="height: 2.7em"&gt;&lt;/span
                              &gt;&lt;span class="sizing reset-size6 size3 mtight"
                                &gt;&lt;span class="mord mtight"&gt;1&lt;/span&gt;&lt;/span
                              &gt;&lt;/span
                            &gt;&lt;/span
                          &gt;&lt;span class="vlist-s"&gt;​&lt;/span&gt;&lt;/span
                        &gt;&lt;span class="vlist-r"
                          &gt;&lt;span class="vlist" style="height: 0.15em"
                            &gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span
                  &gt;&lt;span
                    class="mspace"
                    style="margin-right: 0.2777777777777778em"
                  &gt;&lt;/span
                  &gt;&lt;span class="mrel"&gt;≻&lt;/span
                  &gt;&lt;span
                    class="mspace"
                    style="margin-right: 0.2777777777777778em"
                  &gt;&lt;/span&gt;&lt;/span
                &gt;&lt;span class="base"
                  &gt;&lt;span
                    class="strut"
                    style="height: 1em; vertical-align: -0.25em"
                  &gt;&lt;/span
                  &gt;&lt;span class="mord"
                    &gt;&lt;span
                      class="mord mathdefault"
                      style="margin-right: 0.03588em"
                      &gt;y&lt;/span
                    &gt;&lt;span class="msupsub"
                      &gt;&lt;span class="vlist-t vlist-t2"
                        &gt;&lt;span class="vlist-r"
                          &gt;&lt;span
                            class="vlist"
                            style="height: 0.30110799999999993em"
                            &gt;&lt;span
                              style="
                                top: -2.5500000000000003em;
                                margin-left: -0.03588em;
                                margin-right: 0.05em;
                              "
                              &gt;&lt;span class="pstrut" style="height: 2.7em"&gt;&lt;/span
                              &gt;&lt;span class="sizing reset-size6 size3 mtight"
                                &gt;&lt;span class="mord mtight"&gt;2&lt;/span&gt;&lt;/span
                              &gt;&lt;/span
                            &gt;&lt;/span
                          &gt;&lt;span class="vlist-s"&gt;​&lt;/span&gt;&lt;/span
                        &gt;&lt;span class="vlist-r"
                          &gt;&lt;span class="vlist" style="height: 0.15em"
                            &gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span
                  &gt;&lt;span
                    class="mspace"
                    style="margin-right: 0.2777777777777778em"
                  &gt;&lt;/span
                  &gt;&lt;span class="mrel"&gt;∣&lt;/span
                  &gt;&lt;span
                    class="mspace"
                    style="margin-right: 0.2777777777777778em"
                  &gt;&lt;/span&gt;&lt;/span
                &gt;&lt;span class="base"
                  &gt;&lt;span
                    class="strut"
                    style="height: 1em; vertical-align: -0.25em"
                  &gt;&lt;/span
                  &gt;&lt;span class="mord mathdefault"&gt;x&lt;/span
                  &gt;&lt;span class="mclose"&gt;]&lt;/span
                  &gt;&lt;span
                    class="mspace"
                    style="margin-right: 0.2777777777777778em"
                  &gt;&lt;/span
                  &gt;&lt;span class="mrel"&gt;=&lt;/span
                  &gt;&lt;span
                    class="mspace"
                    style="margin-right: 0.2777777777777778em"
                  &gt;&lt;/span&gt;&lt;/span
                &gt;&lt;span class="base"
                  &gt;&lt;span
                    class="strut"
                    style="height: 2.363em; vertical-align: -0.936em"
                  &gt;&lt;/span
                  &gt;&lt;span class="mord"
                    &gt;&lt;span class="mopen nulldelimiter"&gt;&lt;/span
                    &gt;&lt;span class="mfrac"
                      &gt;&lt;span class="vlist-t vlist-t2"
                        &gt;&lt;span class="vlist-r"
                          &gt;&lt;span class="vlist" style="height: 1.427em"
                            &gt;&lt;span style="top: -2.314em"
                              &gt;&lt;span class="pstrut" style="height: 3em"&gt;&lt;/span
                              &gt;&lt;span class="mord"
                                &gt;&lt;span class="mop"&gt;exp&lt;/span
                                &gt;&lt;span class="mopen"&gt;(&lt;/span
                                &gt;&lt;span class="mord"
                                  &gt;&lt;span
                                    class="mord mathdefault"
                                    style="margin-right: 0.02778em"
                                    &gt;r&lt;/span
                                  &gt;&lt;span class="msupsub"
                                    &gt;&lt;span class="vlist-t vlist-t2"
                                      &gt;&lt;span class="vlist-r"
                                        &gt;&lt;span
                                          class="vlist"
                                          style="height: 0.33610799999999996em"
                                          &gt;&lt;span
                                            style="
                                              top: -2.5500000000000003em;
                                              margin-left: -0.02778em;
                                              margin-right: 0.05em;
                                            "
                                            &gt;&lt;span
                                              class="pstrut"
                                              style="height: 2.7em"
                                            &gt;&lt;/span
                                            &gt;&lt;span
                                              class="sizing reset-size6 size3 mtight"
                                              &gt;&lt;span class="mord text mtight"
                                                &gt;&lt;span class="mord mtight"
                                                  &gt;h&lt;/span
                                                &gt;&lt;/span
                                              &gt;&lt;/span
                                            &gt;&lt;/span
                                          &gt;&lt;/span
                                        &gt;&lt;span class="vlist-s"&gt;​&lt;/span&gt;&lt;/span
                                      &gt;&lt;span class="vlist-r"
                                        &gt;&lt;span
                                          class="vlist"
                                          style="height: 0.15em"
                                          &gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span
                                &gt;&lt;span class="mopen"&gt;(&lt;/span
                                &gt;&lt;span class="mord mathdefault"&gt;x&lt;/span
                                &gt;&lt;span class="mpunct"&gt;,&lt;/span
                                &gt;&lt;span
                                  class="mspace"
                                  style="margin-right: 0.16666666666666666em"
                                &gt;&lt;/span
                                &gt;&lt;span class="mord"
                                  &gt;&lt;span
                                    class="mord mathdefault"
                                    style="margin-right: 0.03588em"
                                    &gt;y&lt;/span
                                  &gt;&lt;span class="msupsub"
                                    &gt;&lt;span class="vlist-t vlist-t2"
                                      &gt;&lt;span class="vlist-r"
                                        &gt;&lt;span
                                          class="vlist"
                                          style="height: 0.30110799999999993em"
                                          &gt;&lt;span
                                            style="
                                              top: -2.5500000000000003em;
                                              margin-left: -0.03588em;
                                              margin-right: 0.05em;
                                            "
                                            &gt;&lt;span
                                              class="pstrut"
                                              style="height: 2.7em"
                                            &gt;&lt;/span
                                            &gt;&lt;span
                                              class="sizing reset-size6 size3 mtight"
                                              &gt;&lt;span class="mord mtight"
                                                &gt;1&lt;/span
                                              &gt;&lt;/span
                                            &gt;&lt;/span
                                          &gt;&lt;/span
                                        &gt;&lt;span class="vlist-s"&gt;​&lt;/span&gt;&lt;/span
                                      &gt;&lt;span class="vlist-r"
                                        &gt;&lt;span
                                          class="vlist"
                                          style="height: 0.15em"
                                          &gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span
                                &gt;&lt;span class="mclose"&gt;)&lt;/span
                                &gt;&lt;span class="mclose"&gt;)&lt;/span
                                &gt;&lt;span
                                  class="mspace"
                                  style="margin-right: 0.2222222222222222em"
                                &gt;&lt;/span
                                &gt;&lt;span class="mbin"&gt;+&lt;/span
                                &gt;&lt;span
                                  class="mspace"
                                  style="margin-right: 0.2222222222222222em"
                                &gt;&lt;/span
                                &gt;&lt;span class="mop"&gt;exp&lt;/span
                                &gt;&lt;span class="mopen"&gt;(&lt;/span
                                &gt;&lt;span class="mord"
                                  &gt;&lt;span
                                    class="mord mathdefault"
                                    style="margin-right: 0.02778em"
                                    &gt;r&lt;/span
                                  &gt;&lt;span class="msupsub"
                                    &gt;&lt;span class="vlist-t vlist-t2"
                                      &gt;&lt;span class="vlist-r"
                                        &gt;&lt;span
                                          class="vlist"
                                          style="height: 0.33610799999999996em"
                                          &gt;&lt;span
                                            style="
                                              top: -2.5500000000000003em;
                                              margin-left: -0.02778em;
                                              margin-right: 0.05em;
                                            "
                                            &gt;&lt;span
                                              class="pstrut"
                                              style="height: 2.7em"
                                            &gt;&lt;/span
                                            &gt;&lt;span
                                              class="sizing reset-size6 size3 mtight"
                                              &gt;&lt;span class="mord text mtight"
                                                &gt;&lt;span class="mord mtight"
                                                  &gt;h&lt;/span
                                                &gt;&lt;/span
                                              &gt;&lt;/span
                                            &gt;&lt;/span
                                          &gt;&lt;/span
                                        &gt;&lt;span class="vlist-s"&gt;​&lt;/span&gt;&lt;/span
                                      &gt;&lt;span class="vlist-r"
                                        &gt;&lt;span
                                          class="vlist"
                                          style="height: 0.15em"
                                          &gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span
                                &gt;&lt;span class="mopen"&gt;(&lt;/span
                                &gt;&lt;span class="mord mathdefault"&gt;x&lt;/span
                                &gt;&lt;span class="mpunct"&gt;,&lt;/span
                                &gt;&lt;span
                                  class="mspace"
                                  style="margin-right: 0.16666666666666666em"
                                &gt;&lt;/span
                                &gt;&lt;span class="mord"
                                  &gt;&lt;span
                                    class="mord mathdefault"
                                    style="margin-right: 0.03588em"
                                    &gt;y&lt;/span
                                  &gt;&lt;span class="msupsub"
                                    &gt;&lt;span class="vlist-t vlist-t2"
                                      &gt;&lt;span class="vlist-r"
                                        &gt;&lt;span
                                          class="vlist"
                                          style="height: 0.30110799999999993em"
                                          &gt;&lt;span
                                            style="
                                              top: -2.5500000000000003em;
                                              margin-left: -0.03588em;
                                              margin-right: 0.05em;
                                            "
                                            &gt;&lt;span
                                              class="pstrut"
                                              style="height: 2.7em"
                                            &gt;&lt;/span
                                            &gt;&lt;span
                                              class="sizing reset-size6 size3 mtight"
                                              &gt;&lt;span class="mord mtight"
                                                &gt;2&lt;/span
                                              &gt;&lt;/span
                                            &gt;&lt;/span
                                          &gt;&lt;/span
                                        &gt;&lt;span class="vlist-s"&gt;​&lt;/span&gt;&lt;/span
                                      &gt;&lt;span class="vlist-r"
                                        &gt;&lt;span
                                          class="vlist"
                                          style="height: 0.15em"
                                          &gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span
                                &gt;&lt;span class="mclose"&gt;)&lt;/span
                                &gt;&lt;span class="mclose"&gt;)&lt;/span&gt;&lt;/span
                              &gt;&lt;/span
                            &gt;&lt;span style="top: -3.23em"
                              &gt;&lt;span class="pstrut" style="height: 3em"&gt;&lt;/span
                              &gt;&lt;span
                                class="frac-line"
                                style="border-bottom-width: 0.04em"
                              &gt;&lt;/span&gt;&lt;/span
                            &gt;&lt;span style="top: -3.677em"
                              &gt;&lt;span class="pstrut" style="height: 3em"&gt;&lt;/span
                              &gt;&lt;span class="mord"
                                &gt;&lt;span class="mop"&gt;exp&lt;/span
                                &gt;&lt;span class="mopen"&gt;(&lt;/span
                                &gt;&lt;span class="mord"
                                  &gt;&lt;span
                                    class="mord mathdefault"
                                    style="margin-right: 0.02778em"
                                    &gt;r&lt;/span
                                  &gt;&lt;span class="msupsub"
                                    &gt;&lt;span class="vlist-t vlist-t2"
                                      &gt;&lt;span class="vlist-r"
                                        &gt;&lt;span
                                          class="vlist"
                                          style="height: 0.33610799999999996em"
                                          &gt;&lt;span
                                            style="
                                              top: -2.5500000000000003em;
                                              margin-left: -0.02778em;
                                              margin-right: 0.05em;
                                            "
                                            &gt;&lt;span
                                              class="pstrut"
                                              style="height: 2.7em"
                                            &gt;&lt;/span
                                            &gt;&lt;span
                                              class="sizing reset-size6 size3 mtight"
                                              &gt;&lt;span class="mord text mtight"
                                                &gt;&lt;span class="mord mtight"
                                                  &gt;h&lt;/span
                                                &gt;&lt;/span
                                              &gt;&lt;/span
                                            &gt;&lt;/span
                                          &gt;&lt;/span
                                        &gt;&lt;span class="vlist-s"&gt;​&lt;/span&gt;&lt;/span
                                      &gt;&lt;span class="vlist-r"
                                        &gt;&lt;span
                                          class="vlist"
                                          style="height: 0.15em"
                                          &gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span
                                &gt;&lt;span class="mopen"&gt;(&lt;/span
                                &gt;&lt;span class="mord mathdefault"&gt;x&lt;/span
                                &gt;&lt;span class="mpunct"&gt;,&lt;/span
                                &gt;&lt;span
                                  class="mspace"
                                  style="margin-right: 0.16666666666666666em"
                                &gt;&lt;/span
                                &gt;&lt;span class="mord"
                                  &gt;&lt;span
                                    class="mord mathdefault"
                                    style="margin-right: 0.03588em"
                                    &gt;y&lt;/span
                                  &gt;&lt;span class="msupsub"
                                    &gt;&lt;span class="vlist-t vlist-t2"
                                      &gt;&lt;span class="vlist-r"
                                        &gt;&lt;span
                                          class="vlist"
                                          style="height: 0.30110799999999993em"
                                          &gt;&lt;span
                                            style="
                                              top: -2.5500000000000003em;
                                              margin-left: -0.03588em;
                                              margin-right: 0.05em;
                                            "
                                            &gt;&lt;span
                                              class="pstrut"
                                              style="height: 2.7em"
                                            &gt;&lt;/span
                                            &gt;&lt;span
                                              class="sizing reset-size6 size3 mtight"
                                              &gt;&lt;span class="mord mtight"
                                                &gt;1&lt;/span
                                              &gt;&lt;/span
                                            &gt;&lt;/span
                                          &gt;&lt;/span
                                        &gt;&lt;span class="vlist-s"&gt;​&lt;/span&gt;&lt;/span
                                      &gt;&lt;span class="vlist-r"
                                        &gt;&lt;span
                                          class="vlist"
                                          style="height: 0.15em"
                                          &gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span
                                &gt;&lt;span class="mclose"&gt;)&lt;/span
                                &gt;&lt;span class="mclose"&gt;)&lt;/span&gt;&lt;/span
                              &gt;&lt;/span
                            &gt;&lt;/span
                          &gt;&lt;span class="vlist-s"&gt;​&lt;/span&gt;&lt;/span
                        &gt;&lt;span class="vlist-r"
                          &gt;&lt;span class="vlist" style="height: 0.936em"
                            &gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span
                    &gt;&lt;span
                      class="mclose nulldelimiter"
                    &gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span
          &gt;&lt;/span&gt;
        &lt;/p&gt;

        &lt;p&gt;
          where $r_\text{h}(x,y)$ is the human ground-truth reward function that
          we don’t have access to. We already arrive at our first complication.
          In practice, some datasets are constructed by using an auxiliary model
          (e.g., GPT-4 or PairRM by
          &lt;a href="https://aclanthology.org/2023.acl-long.792/"
            &gt;Jiang et al., 2023&lt;/a
          &gt;) to pick which response is preferred. So, let’s call the reward
          function of such a language model $r_\text{LM}(x,y)$ and now consider
          the case where our data is annotated by a model (i.e., “RLAIF”).
        &lt;/p&gt;

        &lt;p&gt;
          One thing we do kind of know about $r_\text{LM}(x,y)$ is that it
          usually grows with the length of the response $y$. That is, models
          tend to favor longer responses. And in fact, one popular paper (&lt;a
            href="https://arxiv.org/abs/2310.03716"
            &gt;Singhal et al., 2023&lt;/a
          &gt;) showed that for some models, the reward grows
          &lt;em&gt;linearly&lt;/em&gt; with the length of $y$. But in the case of
          alignment, it’s not clear if we always want our models to output
          longer sequences. For example, if we are trying to prevent harmful
          behavior, length is likely a spurious correlation. On the other hand,
          if we want to promote helpful behavior, increasing the length may
          somewhat correlate with providing useful information. Regardless,
          let’s disentangle length from the reward that we want to maximize and
          define $r^\star(x,y)$ such that
        &lt;/p&gt;

        &lt;p&gt;
          &lt;span class="katex-display"
            &gt;&lt;span class="katex"
              &gt;&lt;span class="katex-mathml"
                &gt;&lt;math xmlns="http://www.w3.org/1998/Math/MathML"
                  &gt;&lt;semantics
                    &gt;&lt;mrow
                      &gt;&lt;msub&gt;&lt;mi&gt;r&lt;/mi&gt;&lt;mtext&gt;LM&lt;/mtext&gt;&lt;/msub
                      &gt;&lt;mo stretchy="false"&gt;(&lt;/mo&gt;&lt;mi&gt;x&lt;/mi
                      &gt;&lt;mo separator="true"&gt;,&lt;/mo&gt;&lt;mi&gt;y&lt;/mi
                      &gt;&lt;mo stretchy="false"&gt;)&lt;/mo&gt;&lt;mo&gt;=&lt;/mo
                      &gt;&lt;mi mathvariant="normal"&gt;∣&lt;/mi&gt;&lt;mi&gt;y&lt;/mi
                      &gt;&lt;mi mathvariant="normal"&gt;∣&lt;/mi
                      &gt;&lt;msup&gt;&lt;mi&gt;r&lt;/mi&gt;&lt;mo&gt;⋆&lt;/mo&gt;&lt;/msup
                      &gt;&lt;mo stretchy="false"&gt;(&lt;/mo&gt;&lt;mi&gt;x&lt;/mi
                      &gt;&lt;mo separator="true"&gt;,&lt;/mo&gt;&lt;mi&gt;y&lt;/mi
                      &gt;&lt;mo stretchy="false"&gt;)&lt;/mo&gt;&lt;/mrow
                    &gt;&lt;annotation encoding="application/x-tex"
                      &gt;r_\text{LM}(x,y) = |y| r^\star(x,y)&lt;/annotation
                    &gt;&lt;/semantics
                  &gt;&lt;/math
                &gt;&lt;/span
              &gt;&lt;span class="katex-html" aria-hidden="true"
                &gt;&lt;span class="base"
                  &gt;&lt;span
                    class="strut"
                    style="height: 1em; vertical-align: -0.25em"
                  &gt;&lt;/span
                  &gt;&lt;span class="mord"
                    &gt;&lt;span
                      class="mord mathdefault"
                      style="margin-right: 0.02778em"
                      &gt;r&lt;/span
                    &gt;&lt;span class="msupsub"
                      &gt;&lt;span class="vlist-t vlist-t2"
                        &gt;&lt;span class="vlist-r"
                          &gt;&lt;span
                            class="vlist"
                            style="height: 0.32833099999999993em"
                            &gt;&lt;span
                              style="
                                top: -2.5500000000000003em;
                                margin-left: -0.02778em;
                                margin-right: 0.05em;
                              "
                              &gt;&lt;span class="pstrut" style="height: 2.7em"&gt;&lt;/span
                              &gt;&lt;span class="sizing reset-size6 size3 mtight"
                                &gt;&lt;span class="mord text mtight"
                                  &gt;&lt;span class="mord mtight"&gt;LM&lt;/span&gt;&lt;/span
                                &gt;&lt;/span
                              &gt;&lt;/span
                            &gt;&lt;/span
                          &gt;&lt;span class="vlist-s"&gt;​&lt;/span&gt;&lt;/span
                        &gt;&lt;span class="vlist-r"
                          &gt;&lt;span class="vlist" style="height: 0.15em"
                            &gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span
                  &gt;&lt;span class="mopen"&gt;(&lt;/span
                  &gt;&lt;span class="mord mathdefault"&gt;x&lt;/span
                  &gt;&lt;span class="mpunct"&gt;,&lt;/span
                  &gt;&lt;span
                    class="mspace"
                    style="margin-right: 0.16666666666666666em"
                  &gt;&lt;/span
                  &gt;&lt;span
                    class="mord mathdefault"
                    style="margin-right: 0.03588em"
                    &gt;y&lt;/span
                  &gt;&lt;span class="mclose"&gt;)&lt;/span
                  &gt;&lt;span
                    class="mspace"
                    style="margin-right: 0.2777777777777778em"
                  &gt;&lt;/span
                  &gt;&lt;span class="mrel"&gt;=&lt;/span
                  &gt;&lt;span
                    class="mspace"
                    style="margin-right: 0.2777777777777778em"
                  &gt;&lt;/span&gt;&lt;/span
                &gt;&lt;span class="base"
                  &gt;&lt;span
                    class="strut"
                    style="height: 1em; vertical-align: -0.25em"
                  &gt;&lt;/span
                  &gt;&lt;span class="mord"&gt;∣&lt;/span
                  &gt;&lt;span
                    class="mord mathdefault"
                    style="margin-right: 0.03588em"
                    &gt;y&lt;/span
                  &gt;&lt;span class="mord"&gt;∣&lt;/span
                  &gt;&lt;span class="mord"
                    &gt;&lt;span
                      class="mord mathdefault"
                      style="margin-right: 0.02778em"
                      &gt;r&lt;/span
                    &gt;&lt;span class="msupsub"
                      &gt;&lt;span class="vlist-t"
                        &gt;&lt;span class="vlist-r"
                          &gt;&lt;span class="vlist" style="height: 0.738696em"
                            &gt;&lt;span style="top: -3.113em; margin-right: 0.05em"
                              &gt;&lt;span class="pstrut" style="height: 2.7em"&gt;&lt;/span
                              &gt;&lt;span class="sizing reset-size6 size3 mtight"
                                &gt;&lt;span class="mbin mtight"&gt;⋆&lt;/span&gt;&lt;/span
                              &gt;&lt;/span
                            &gt;&lt;/span
                          &gt;&lt;/span
                        &gt;&lt;/span
                      &gt;&lt;/span
                    &gt;&lt;/span
                  &gt;&lt;span class="mopen"&gt;(&lt;/span
                  &gt;&lt;span class="mord mathdefault"&gt;x&lt;/span
                  &gt;&lt;span class="mpunct"&gt;,&lt;/span
                  &gt;&lt;span
                    class="mspace"
                    style="margin-right: 0.16666666666666666em"
                  &gt;&lt;/span
                  &gt;&lt;span
                    class="mord mathdefault"
                    style="margin-right: 0.03588em"
                    &gt;y&lt;/span
                  &gt;&lt;span class="mclose"&gt;)&lt;/span&gt;&lt;/span
                &gt;&lt;/span
              &gt;&lt;/span
            &gt;&lt;/span
          &gt;
        &lt;/p&gt;

        &lt;p&gt;
          Now, we want to learn a model $\pi^\star$ that is trained with access
          to only data annotated according to $r_\text{GPT}$ but maximizes the
          length-normalized reward $r^\star$. In other words, we can write:
        &lt;/p&gt;

        &lt;p&gt;
          &lt;span class="katex-display"
            &gt;&lt;span class="katex"
              &gt;&lt;span class="katex-mathml"
                &gt;&lt;math xmlns="http://www.w3.org/1998/Math/MathML"
                  &gt;&lt;semantics
                    &gt;&lt;mrow
                      &gt;&lt;msup&gt;&lt;mi&gt;r&lt;/mi&gt;&lt;mo&gt;⋆&lt;/mo&gt;&lt;/msup
                      &gt;&lt;mo stretchy="false"&gt;(&lt;/mo&gt;&lt;mi&gt;x&lt;/mi
                      &gt;&lt;mo separator="true"&gt;,&lt;/mo&gt;&lt;mi&gt;y&lt;/mi
                      &gt;&lt;mo stretchy="false"&gt;)&lt;/mo&gt;&lt;mo&gt;=&lt;/mo
                      &gt;&lt;mfrac
                        &gt;&lt;mrow
                          &gt;&lt;mi&gt;log&lt;/mi&gt;&lt;mo&gt;⁡&lt;/mo&gt;&lt;mo stretchy="false"&gt;(&lt;/mo
                          &gt;&lt;msup&gt;&lt;mi&gt;π&lt;/mi&gt;&lt;mo&gt;⋆&lt;/mo&gt;&lt;/msup
                          &gt;&lt;mo stretchy="false"&gt;(&lt;/mo&gt;&lt;mi&gt;x&lt;/mi
                          &gt;&lt;mo separator="true"&gt;,&lt;/mo&gt;&lt;mi&gt;y&lt;/mi
                          &gt;&lt;mo stretchy="false"&gt;)&lt;/mo
                          &gt;&lt;mo stretchy="false"&gt;)&lt;/mo&gt;&lt;/mrow
                        &gt;&lt;mrow
                          &gt;&lt;mi mathvariant="normal"&gt;∣&lt;/mi&gt;&lt;mi&gt;y&lt;/mi
                          &gt;&lt;mi mathvariant="normal"&gt;∣&lt;/mi&gt;&lt;/mrow
                        &gt;&lt;/mfrac
                      &gt;&lt;/mrow
                    &gt;&lt;annotation encoding="application/x-tex"&gt;
                      r^\star(x,y) = \frac{\log(\pi^\star(x,y))}{|y|}
                    &lt;/annotation&gt;&lt;/semantics
                  &gt;&lt;/math
                &gt;&lt;/span
              &gt;&lt;span class="katex-html" aria-hidden="true"
                &gt;&lt;span class="base"
                  &gt;&lt;span
                    class="strut"
                    style="height: 1em; vertical-align: -0.25em"
                  &gt;&lt;/span
                  &gt;&lt;span class="mord"
                    &gt;&lt;span
                      class="mord mathdefault"
                      style="margin-right: 0.02778em"
                      &gt;r&lt;/span
                    &gt;&lt;span class="msupsub"
                      &gt;&lt;span class="vlist-t"
                        &gt;&lt;span class="vlist-r"
                          &gt;&lt;span class="vlist" style="height: 0.738696em"
                            &gt;&lt;span style="top: -3.113em; margin-right: 0.05em"
                              &gt;&lt;span class="pstrut" style="height: 2.7em"&gt;&lt;/span
                              &gt;&lt;span class="sizing reset-size6 size3 mtight"
                                &gt;&lt;span class="mbin mtight"&gt;⋆&lt;/span&gt;&lt;/span
                              &gt;&lt;/span
                            &gt;&lt;/span
                          &gt;&lt;/span
                        &gt;&lt;/span
                      &gt;&lt;/span
                    &gt;&lt;/span
                  &gt;&lt;span class="mopen"&gt;(&lt;/span
                  &gt;&lt;span class="mord mathdefault"&gt;x&lt;/span
                  &gt;&lt;span class="mpunct"&gt;,&lt;/span
                  &gt;&lt;span
                    class="mspace"
                    style="margin-right: 0.16666666666666666em"
                  &gt;&lt;/span
                  &gt;&lt;span
                    class="mord mathdefault"
                    style="margin-right: 0.03588em"
                    &gt;y&lt;/span
                  &gt;&lt;span class="mclose"&gt;)&lt;/span
                  &gt;&lt;span
                    class="mspace"
                    style="margin-right: 0.2777777777777778em"
                  &gt;&lt;/span
                  &gt;&lt;span class="mrel"&gt;=&lt;/span
                  &gt;&lt;span
                    class="mspace"
                    style="margin-right: 0.2777777777777778em"
                  &gt;&lt;/span&gt;&lt;/span
                &gt;&lt;span class="base"
                  &gt;&lt;span
                    class="strut"
                    style="height: 2.363em; vertical-align: -0.936em"
                  &gt;&lt;/span
                  &gt;&lt;span class="mord"
                    &gt;&lt;span class="mopen nulldelimiter"&gt;&lt;/span
                    &gt;&lt;span class="mfrac"
                      &gt;&lt;span class="vlist-t vlist-t2"
                        &gt;&lt;span class="vlist-r"
                          &gt;&lt;span class="vlist" style="height: 1.427em"
                            &gt;&lt;span style="top: -2.314em"
                              &gt;&lt;span class="pstrut" style="height: 3em"&gt;&lt;/span
                              &gt;&lt;span class="mord"
                                &gt;&lt;span class="mord"&gt;∣&lt;/span
                                &gt;&lt;span
                                  class="mord mathdefault"
                                  style="margin-right: 0.03588em"
                                  &gt;y&lt;/span
                                &gt;&lt;span class="mord"&gt;∣&lt;/span&gt;&lt;/span
                              &gt;&lt;/span
                            &gt;&lt;span style="top: -3.23em"
                              &gt;&lt;span class="pstrut" style="height: 3em"&gt;&lt;/span
                              &gt;&lt;span
                                class="frac-line"
                                style="border-bottom-width: 0.04em"
                              &gt;&lt;/span&gt;&lt;/span
                            &gt;&lt;span style="top: -3.677em"
                              &gt;&lt;span class="pstrut" style="height: 3em"&gt;&lt;/span
                              &gt;&lt;span class="mord"
                                &gt;&lt;span class="mop"
                                  &gt;lo&lt;span style="margin-right: 0.01389em"
                                    &gt;g&lt;/span
                                  &gt;&lt;/span
                                &gt;&lt;span class="mopen"&gt;(&lt;/span
                                &gt;&lt;span class="mord"
                                  &gt;&lt;span
                                    class="mord mathdefault"
                                    style="margin-right: 0.03588em"
                                    &gt;π&lt;/span
                                  &gt;&lt;span class="msupsub"
                                    &gt;&lt;span class="vlist-t"
                                      &gt;&lt;span class="vlist-r"
                                        &gt;&lt;span
                                          class="vlist"
                                          style="height: 0.688696em"
                                          &gt;&lt;span
                                            style="
                                              top: -3.063em;
                                              margin-right: 0.05em;
                                            "
                                            &gt;&lt;span
                                              class="pstrut"
                                              style="height: 2.7em"
                                            &gt;&lt;/span
                                            &gt;&lt;span
                                              class="sizing reset-size6 size3 mtight"
                                              &gt;&lt;span class="mbin mtight"
                                                &gt;⋆&lt;/span
                                              &gt;&lt;/span
                                            &gt;&lt;/span
                                          &gt;&lt;/span
                                        &gt;&lt;/span
                                      &gt;&lt;/span
                                    &gt;&lt;/span
                                  &gt;&lt;/span
                                &gt;&lt;span class="mopen"&gt;(&lt;/span
                                &gt;&lt;span class="mord mathdefault"&gt;x&lt;/span
                                &gt;&lt;span class="mpunct"&gt;,&lt;/span
                                &gt;&lt;span
                                  class="mspace"
                                  style="margin-right: 0.16666666666666666em"
                                &gt;&lt;/span
                                &gt;&lt;span
                                  class="mord mathdefault"
                                  style="margin-right: 0.03588em"
                                  &gt;y&lt;/span
                                &gt;&lt;span class="mclose"&gt;)&lt;/span
                                &gt;&lt;span class="mclose"&gt;)&lt;/span&gt;&lt;/span
                              &gt;&lt;/span
                            &gt;&lt;/span
                          &gt;&lt;span class="vlist-s"&gt;​&lt;/span&gt;&lt;/span
                        &gt;&lt;span class="vlist-r"
                          &gt;&lt;span class="vlist" style="height: 0.936em"
                            &gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span
                    &gt;&lt;span
                      class="mclose nulldelimiter"
                    &gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span
          &gt;&lt;/span&gt;
        &lt;/p&gt;

        &lt;p&gt;
          where I’m keeping with the standard ML practice of considering the
          log-likelihood of a softmax-parametrized model instead of the
          likelihood directly. We can now define a score $\Lambda$ using
          $r^\star$.
        &lt;/p&gt;

        &lt;p&gt;
          &lt;span class="katex-display"
            &gt;&lt;span class="katex"
              &gt;&lt;span class="katex-mathml"
                &gt;&lt;math xmlns="http://www.w3.org/1998/Math/MathML"
                  &gt;&lt;semantics
                    &gt;&lt;mrow
                      &gt;&lt;mi mathvariant="normal"&gt;Λ&lt;/mi&gt;&lt;mo stretchy="false"&gt;[&lt;/mo
                      &gt;&lt;msub&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/msub
                      &gt;&lt;mo&gt;≻&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mn&gt;2&lt;/mn&gt;&lt;/msub
                      &gt;&lt;mo&gt;∣&lt;/mo&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mo stretchy="false"&gt;]&lt;/mo&gt;&lt;mo&gt;=&lt;/mo
                      &gt;&lt;mi&gt;σ&lt;/mi&gt;&lt;mo stretchy="false"&gt;(&lt;/mo
                      &gt;&lt;msup&gt;&lt;mi&gt;r&lt;/mi&gt;&lt;mo&gt;⋆&lt;/mo&gt;&lt;/msup
                      &gt;&lt;mo stretchy="false"&gt;(&lt;/mo&gt;&lt;mi&gt;x&lt;/mi
                      &gt;&lt;mo separator="true"&gt;,&lt;/mo
                      &gt;&lt;msub&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/msub
                      &gt;&lt;mo stretchy="false"&gt;)&lt;/mo&gt;&lt;mo&gt;−&lt;/mo
                      &gt;&lt;msup&gt;&lt;mi&gt;r&lt;/mi&gt;&lt;mo&gt;⋆&lt;/mo&gt;&lt;/msup
                      &gt;&lt;mo stretchy="false"&gt;(&lt;/mo&gt;&lt;mi&gt;x&lt;/mi
                      &gt;&lt;mo separator="true"&gt;,&lt;/mo
                      &gt;&lt;msub&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mn&gt;2&lt;/mn&gt;&lt;/msub
                      &gt;&lt;mo stretchy="false"&gt;)&lt;/mo
                      &gt;&lt;mo stretchy="false"&gt;)&lt;/mo&gt;&lt;/mrow
                    &gt;&lt;annotation encoding="application/x-tex"&gt;
                      \Lambda[y_1\succ y_2 \mid x] = \sigma(r^\star(x,y_1) -
                      r^\star(x,y_2))
                    &lt;/annotation&gt;&lt;/semantics
                  &gt;&lt;/math
                &gt;&lt;/span
              &gt;&lt;span class="katex-html" aria-hidden="true"
                &gt;&lt;span class="base"
                  &gt;&lt;span
                    class="strut"
                    style="height: 1em; vertical-align: -0.25em"
                  &gt;&lt;/span
                  &gt;&lt;span class="mord"&gt;Λ&lt;/span&gt;&lt;span class="mopen"&gt;[&lt;/span
                  &gt;&lt;span class="mord"
                    &gt;&lt;span
                      class="mord mathdefault"
                      style="margin-right: 0.03588em"
                      &gt;y&lt;/span
                    &gt;&lt;span class="msupsub"
                      &gt;&lt;span class="vlist-t vlist-t2"
                        &gt;&lt;span class="vlist-r"
                          &gt;&lt;span
                            class="vlist"
                            style="height: 0.30110799999999993em"
                            &gt;&lt;span
                              style="
                                top: -2.5500000000000003em;
                                margin-left: -0.03588em;
                                margin-right: 0.05em;
                              "
                              &gt;&lt;span class="pstrut" style="height: 2.7em"&gt;&lt;/span
                              &gt;&lt;span class="sizing reset-size6 size3 mtight"
                                &gt;&lt;span class="mord mtight"&gt;1&lt;/span&gt;&lt;/span
                              &gt;&lt;/span
                            &gt;&lt;/span
                          &gt;&lt;span class="vlist-s"&gt;​&lt;/span&gt;&lt;/span
                        &gt;&lt;span class="vlist-r"
                          &gt;&lt;span class="vlist" style="height: 0.15em"
                            &gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span
                  &gt;&lt;span
                    class="mspace"
                    style="margin-right: 0.2777777777777778em"
                  &gt;&lt;/span
                  &gt;&lt;span class="mrel"&gt;≻&lt;/span
                  &gt;&lt;span
                    class="mspace"
                    style="margin-right: 0.2777777777777778em"
                  &gt;&lt;/span&gt;&lt;/span
                &gt;&lt;span class="base"
                  &gt;&lt;span
                    class="strut"
                    style="height: 1em; vertical-align: -0.25em"
                  &gt;&lt;/span
                  &gt;&lt;span class="mord"
                    &gt;&lt;span
                      class="mord mathdefault"
                      style="margin-right: 0.03588em"
                      &gt;y&lt;/span
                    &gt;&lt;span class="msupsub"
                      &gt;&lt;span class="vlist-t vlist-t2"
                        &gt;&lt;span class="vlist-r"
                          &gt;&lt;span
                            class="vlist"
                            style="height: 0.30110799999999993em"
                            &gt;&lt;span
                              style="
                                top: -2.5500000000000003em;
                                margin-left: -0.03588em;
                                margin-right: 0.05em;
                              "
                              &gt;&lt;span class="pstrut" style="height: 2.7em"&gt;&lt;/span
                              &gt;&lt;span class="sizing reset-size6 size3 mtight"
                                &gt;&lt;span class="mord mtight"&gt;2&lt;/span&gt;&lt;/span
                              &gt;&lt;/span
                            &gt;&lt;/span
                          &gt;&lt;span class="vlist-s"&gt;​&lt;/span&gt;&lt;/span
                        &gt;&lt;span class="vlist-r"
                          &gt;&lt;span class="vlist" style="height: 0.15em"
                            &gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span
                  &gt;&lt;span
                    class="mspace"
                    style="margin-right: 0.2777777777777778em"
                  &gt;&lt;/span
                  &gt;&lt;span class="mrel"&gt;∣&lt;/span
                  &gt;&lt;span
                    class="mspace"
                    style="margin-right: 0.2777777777777778em"
                  &gt;&lt;/span&gt;&lt;/span
                &gt;&lt;span class="base"
                  &gt;&lt;span
                    class="strut"
                    style="height: 1em; vertical-align: -0.25em"
                  &gt;&lt;/span
                  &gt;&lt;span class="mord mathdefault"&gt;x&lt;/span
                  &gt;&lt;span class="mclose"&gt;]&lt;/span
                  &gt;&lt;span
                    class="mspace"
                    style="margin-right: 0.2777777777777778em"
                  &gt;&lt;/span
                  &gt;&lt;span class="mrel"&gt;=&lt;/span
                  &gt;&lt;span
                    class="mspace"
                    style="margin-right: 0.2777777777777778em"
                  &gt;&lt;/span&gt;&lt;/span
                &gt;&lt;span class="base"
                  &gt;&lt;span
                    class="strut"
                    style="height: 1em; vertical-align: -0.25em"
                  &gt;&lt;/span
                  &gt;&lt;span
                    class="mord mathdefault"
                    style="margin-right: 0.03588em"
                    &gt;σ&lt;/span
                  &gt;&lt;span class="mopen"&gt;(&lt;/span
                  &gt;&lt;span class="mord"
                    &gt;&lt;span
                      class="mord mathdefault"
                      style="margin-right: 0.02778em"
                      &gt;r&lt;/span
                    &gt;&lt;span class="msupsub"
                      &gt;&lt;span class="vlist-t"
                        &gt;&lt;span class="vlist-r"
                          &gt;&lt;span class="vlist" style="height: 0.738696em"
                            &gt;&lt;span style="top: -3.113em; margin-right: 0.05em"
                              &gt;&lt;span class="pstrut" style="height: 2.7em"&gt;&lt;/span
                              &gt;&lt;span class="sizing reset-size6 size3 mtight"
                                &gt;&lt;span class="mbin mtight"&gt;⋆&lt;/span&gt;&lt;/span
                              &gt;&lt;/span
                            &gt;&lt;/span
                          &gt;&lt;/span
                        &gt;&lt;/span
                      &gt;&lt;/span
                    &gt;&lt;/span
                  &gt;&lt;span class="mopen"&gt;(&lt;/span
                  &gt;&lt;span class="mord mathdefault"&gt;x&lt;/span
                  &gt;&lt;span class="mpunct"&gt;,&lt;/span
                  &gt;&lt;span
                    class="mspace"
                    style="margin-right: 0.16666666666666666em"
                  &gt;&lt;/span
                  &gt;&lt;span class="mord"
                    &gt;&lt;span
                      class="mord mathdefault"
                      style="margin-right: 0.03588em"
                      &gt;y&lt;/span
                    &gt;&lt;span class="msupsub"
                      &gt;&lt;span class="vlist-t vlist-t2"
                        &gt;&lt;span class="vlist-r"
                          &gt;&lt;span
                            class="vlist"
                            style="height: 0.30110799999999993em"
                            &gt;&lt;span
                              style="
                                top: -2.5500000000000003em;
                                margin-left: -0.03588em;
                                margin-right: 0.05em;
                              "
                              &gt;&lt;span class="pstrut" style="height: 2.7em"&gt;&lt;/span
                              &gt;&lt;span class="sizing reset-size6 size3 mtight"
                                &gt;&lt;span class="mord mtight"&gt;1&lt;/span&gt;&lt;/span
                              &gt;&lt;/span
                            &gt;&lt;/span
                          &gt;&lt;span class="vlist-s"&gt;​&lt;/span&gt;&lt;/span
                        &gt;&lt;span class="vlist-r"
                          &gt;&lt;span class="vlist" style="height: 0.15em"
                            &gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span
                  &gt;&lt;span class="mclose"&gt;)&lt;/span
                  &gt;&lt;span
                    class="mspace"
                    style="margin-right: 0.2222222222222222em"
                  &gt;&lt;/span
                  &gt;&lt;span class="mbin"&gt;−&lt;/span
                  &gt;&lt;span
                    class="mspace"
                    style="margin-right: 0.2222222222222222em"
                  &gt;&lt;/span&gt;&lt;/span
                &gt;&lt;span class="base"
                  &gt;&lt;span
                    class="strut"
                    style="height: 1em; vertical-align: -0.25em"
                  &gt;&lt;/span
                  &gt;&lt;span class="mord"
                    &gt;&lt;span
                      class="mord mathdefault"
                      style="margin-right: 0.02778em"
                      &gt;r&lt;/span
                    &gt;&lt;span class="msupsub"
                      &gt;&lt;span class="vlist-t"
                        &gt;&lt;span class="vlist-r"
                          &gt;&lt;span class="vlist" style="height: 0.738696em"
                            &gt;&lt;span style="top: -3.113em; margin-right: 0.05em"
                              &gt;&lt;span class="pstrut" style="height: 2.7em"&gt;&lt;/span
                              &gt;&lt;span class="sizing reset-size6 size3 mtight"
                                &gt;&lt;span class="mbin mtight"&gt;⋆&lt;/span&gt;&lt;/span
                              &gt;&lt;/span
                            &gt;&lt;/span
                          &gt;&lt;/span
                        &gt;&lt;/span
                      &gt;&lt;/span
                    &gt;&lt;/span
                  &gt;&lt;span class="mopen"&gt;(&lt;/span
                  &gt;&lt;span class="mord mathdefault"&gt;x&lt;/span
                  &gt;&lt;span class="mpunct"&gt;,&lt;/span
                  &gt;&lt;span
                    class="mspace"
                    style="margin-right: 0.16666666666666666em"
                  &gt;&lt;/span
                  &gt;&lt;span class="mord"
                    &gt;&lt;span
                      class="mord mathdefault"
                      style="margin-right: 0.03588em"
                      &gt;y&lt;/span
                    &gt;&lt;span class="msupsub"
                      &gt;&lt;span class="vlist-t vlist-t2"
                        &gt;&lt;span class="vlist-r"
                          &gt;&lt;span
                            class="vlist"
                            style="height: 0.30110799999999993em"
                            &gt;&lt;span
                              style="
                                top: -2.5500000000000003em;
                                margin-left: -0.03588em;
                                margin-right: 0.05em;
                              "
                              &gt;&lt;span class="pstrut" style="height: 2.7em"&gt;&lt;/span
                              &gt;&lt;span class="sizing reset-size6 size3 mtight"
                                &gt;&lt;span class="mord mtight"&gt;2&lt;/span&gt;&lt;/span
                              &gt;&lt;/span
                            &gt;&lt;/span
                          &gt;&lt;span class="vlist-s"&gt;​&lt;/span&gt;&lt;/span
                        &gt;&lt;span class="vlist-r"
                          &gt;&lt;span class="vlist" style="height: 0.15em"
                            &gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span
                  &gt;&lt;span class="mclose"&gt;)&lt;/span
                  &gt;&lt;span class="mclose"&gt;)&lt;/span&gt;&lt;/span
                &gt;&lt;/span
              &gt;&lt;/span
            &gt;&lt;/span
          &gt;
        &lt;/p&gt;

        &lt;p&gt;
          where $\sigma$ is the sigmoid function. Note that the DPO derivation
          does not rely on pulling this score out of thin air because they used
          the Bradley-Terry assumption on $r_\text{h}$ to show that, if the data
          is human-annotated, then $\sigma(r_\text{h}(x, y_1) - r_\text{h}(x,
          y_2))$ is exactly equal to the ground-truth probability that $y_1$ is
          preferred over $y_2$. Here, however, I want to make it clear that we
          have no reason to believe that $r^\star$ obeys the Bradley-Terry
          assumption. But, optimistically, if it did, then $\Lambda$ would
          define a valid probability distribution capturing how much the
          &lt;em&gt;model&lt;/em&gt; prefers $y_1$ over $y_2$ if we removed the bias that
          favored long responses.
        &lt;/p&gt;

        &lt;p&gt;
          Now, we want to train the model to maximize this score in the case
          that $y_1 = y_w$ is the winning response and $y_2=y_l$ is the losing
          one. So we can write our length-normalized objective as
        &lt;/p&gt;

        &lt;p&gt;
          &lt;span class="katex-display"
            &gt;&lt;span class="katex"
              &gt;&lt;span class="katex-mathml"
                &gt;&lt;math xmlns="http://www.w3.org/1998/Math/MathML"
                  &gt;&lt;semantics
                    &gt;&lt;mrow
                      &gt;&lt;mi mathvariant="script"&gt;L&lt;/mi&gt;&lt;mo stretchy="false"&gt;(&lt;/mo
                      &gt;&lt;msub&gt;&lt;mi&gt;π&lt;/mi&gt;&lt;mi&gt;θ&lt;/mi&gt;&lt;/msub
                      &gt;&lt;mo separator="true"&gt;,&lt;/mo&gt;&lt;mi mathvariant="script"&gt;D&lt;/mi
                      &gt;&lt;mo stretchy="false"&gt;)&lt;/mo&gt;&lt;mo&gt;=&lt;/mo
                      &gt;&lt;msub
                        &gt;&lt;mo&gt;&lt;mi mathvariant="double-struck"&gt;E&lt;/mi&gt;&lt;/mo
                        &gt;&lt;mrow
                          &gt;&lt;mo stretchy="false"&gt;(&lt;/mo&gt;&lt;mi&gt;x&lt;/mi
                          &gt;&lt;mo separator="true"&gt;,&lt;/mo
                          &gt;&lt;msub&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mi&gt;w&lt;/mi&gt;&lt;/msub
                          &gt;&lt;mo separator="true"&gt;,&lt;/mo
                          &gt;&lt;msub&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mi&gt;l&lt;/mi&gt;&lt;/msub
                          &gt;&lt;mo stretchy="false"&gt;)&lt;/mo&gt;&lt;mo&gt;∼&lt;/mo
                          &gt;&lt;mi mathvariant="script"&gt;D&lt;/mi&gt;&lt;/mrow
                        &gt;&lt;/msub
                      &gt;&lt;mrow
                        &gt;&lt;mo fence="true"&gt;[&lt;/mo&gt;&lt;mi&gt;log&lt;/mi&gt;&lt;mo&gt;⁡&lt;/mo&gt;&lt;mi&gt;σ&lt;/mi
                        &gt;&lt;mrow
                          &gt;&lt;mo fence="true"&gt;(&lt;/mo
                          &gt;&lt;mfrac
                            &gt;&lt;mrow
                              &gt;&lt;mi&gt;log&lt;/mi&gt;&lt;mo&gt;⁡&lt;/mo
                              &gt;&lt;msub&gt;&lt;mi&gt;π&lt;/mi&gt;&lt;mi&gt;θ&lt;/mi&gt;&lt;/msub
                              &gt;&lt;mo stretchy="false"&gt;(&lt;/mo
                              &gt;&lt;msub&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mi&gt;w&lt;/mi&gt;&lt;/msub
                              &gt;&lt;mo&gt;∣&lt;/mo&gt;&lt;mi&gt;x&lt;/mi
                              &gt;&lt;mo stretchy="false"&gt;)&lt;/mo&gt;&lt;/mrow
                            &gt;&lt;mrow
                              &gt;&lt;mi mathvariant="normal"&gt;∣&lt;/mi
                              &gt;&lt;msub&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mi&gt;w&lt;/mi&gt;&lt;/msub
                              &gt;&lt;mi mathvariant="normal"&gt;∣&lt;/mi&gt;&lt;/mrow
                            &gt;&lt;/mfrac
                          &gt;&lt;mo&gt;−&lt;/mo
                          &gt;&lt;mfrac
                            &gt;&lt;mrow
                              &gt;&lt;mi&gt;log&lt;/mi&gt;&lt;mo&gt;⁡&lt;/mo
                              &gt;&lt;msub&gt;&lt;mi&gt;π&lt;/mi&gt;&lt;mi&gt;θ&lt;/mi&gt;&lt;/msub
                              &gt;&lt;mo stretchy="false"&gt;(&lt;/mo
                              &gt;&lt;msub&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mi&gt;l&lt;/mi&gt;&lt;/msub
                              &gt;&lt;mo&gt;∣&lt;/mo&gt;&lt;mi&gt;x&lt;/mi
                              &gt;&lt;mo stretchy="false"&gt;)&lt;/mo&gt;&lt;/mrow
                            &gt;&lt;mrow
                              &gt;&lt;mi mathvariant="normal"&gt;∣&lt;/mi
                              &gt;&lt;msub&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mi&gt;l&lt;/mi&gt;&lt;/msub
                              &gt;&lt;mi mathvariant="normal"&gt;∣&lt;/mi&gt;&lt;/mrow
                            &gt;&lt;/mfrac
                          &gt;&lt;mo fence="true"&gt;)&lt;/mo&gt;&lt;/mrow
                        &gt;&lt;mo fence="true"&gt;]&lt;/mo&gt;&lt;/mrow
                      &gt;&lt;/mrow
                    &gt;&lt;annotation encoding="application/x-tex"&gt;
                      \mathcal{L}(\pi_\theta, \mathcal{D}) =
                      \mathop{\mathbb{E}}_{(x, y_w, y_l)\sim\mathcal{D}}
                      \left[\log \sigma \left(\frac{\log \pi_\theta (y_w\mid
                      x)}{|y_w|} - \frac{\log \pi_\theta(y_l\mid
                      x)}{|y_l|}\right)\right]&lt;/annotation
                    &gt;&lt;/semantics
                  &gt;&lt;/math
                &gt;&lt;/span
              &gt;&lt;span class="katex-html" aria-hidden="true"
                &gt;&lt;span class="base"
                  &gt;&lt;span
                    class="strut"
                    style="height: 1em; vertical-align: -0.25em"
                  &gt;&lt;/span
                  &gt;&lt;span class="mord"&gt;&lt;span class="mord mathcal"&gt;L&lt;/span&gt;&lt;/span
                  &gt;&lt;span class="mopen"&gt;(&lt;/span
                  &gt;&lt;span class="mord"
                    &gt;&lt;span
                      class="mord mathdefault"
                      style="margin-right: 0.03588em"
                      &gt;π&lt;/span
                    &gt;&lt;span class="msupsub"
                      &gt;&lt;span class="vlist-t vlist-t2"
                        &gt;&lt;span class="vlist-r"
                          &gt;&lt;span
                            class="vlist"
                            style="height: 0.33610799999999996em"
                            &gt;&lt;span
                              style="
                                top: -2.5500000000000003em;
                                margin-left: -0.03588em;
                                margin-right: 0.05em;
                              "
                              &gt;&lt;span class="pstrut" style="height: 2.7em"&gt;&lt;/span
                              &gt;&lt;span class="sizing reset-size6 size3 mtight"
                                &gt;&lt;span
                                  class="mord mathdefault mtight"
                                  style="margin-right: 0.02778em"
                                  &gt;θ&lt;/span
                                &gt;&lt;/span
                              &gt;&lt;/span
                            &gt;&lt;/span
                          &gt;&lt;span class="vlist-s"&gt;​&lt;/span&gt;&lt;/span
                        &gt;&lt;span class="vlist-r"
                          &gt;&lt;span class="vlist" style="height: 0.15em"
                            &gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span
                  &gt;&lt;span class="mpunct"&gt;,&lt;/span
                  &gt;&lt;span
                    class="mspace"
                    style="margin-right: 0.16666666666666666em"
                  &gt;&lt;/span
                  &gt;&lt;span class="mord"
                    &gt;&lt;span class="mord mathcal" style="margin-right: 0.02778em"
                      &gt;D&lt;/span
                    &gt;&lt;/span
                  &gt;&lt;span class="mclose"&gt;)&lt;/span
                  &gt;&lt;span
                    class="mspace"
                    style="margin-right: 0.2777777777777778em"
                  &gt;&lt;/span
                  &gt;&lt;span class="mrel"&gt;=&lt;/span
                  &gt;&lt;span
                    class="mspace"
                    style="margin-right: 0.2777777777777778em"
                  &gt;&lt;/span&gt;&lt;/span
                &gt;&lt;span class="base"
                  &gt;&lt;span
                    class="strut"
                    style="height: 2.40003em; vertical-align: -0.95003em"
                  &gt;&lt;/span
                  &gt;&lt;span class="mop"
                    &gt;&lt;span class="mop"
                      &gt;&lt;span class="mord"
                        &gt;&lt;span class="mord mathbb"&gt;E&lt;/span&gt;&lt;/span
                      &gt;&lt;/span
                    &gt;&lt;span class="msupsub"
                      &gt;&lt;span class="vlist-t vlist-t2"
                        &gt;&lt;span class="vlist-r"
                          &gt;&lt;span
                            class="vlist"
                            style="height: 0.34480000000000005em"
                            &gt;&lt;span style="top: -2.5198em; margin-right: 0.05em"
                              &gt;&lt;span class="pstrut" style="height: 2.7em"&gt;&lt;/span
                              &gt;&lt;span class="sizing reset-size6 size3 mtight"
                                &gt;&lt;span class="mord mtight"
                                  &gt;&lt;span class="mopen mtight"&gt;(&lt;/span
                                  &gt;&lt;span class="mord mathdefault mtight"&gt;x&lt;/span
                                  &gt;&lt;span class="mpunct mtight"&gt;,&lt;/span
                                  &gt;&lt;span class="mord mtight"
                                    &gt;&lt;span
                                      class="mord mathdefault mtight"
                                      style="margin-right: 0.03588em"
                                      &gt;y&lt;/span
                                    &gt;&lt;span class="msupsub"
                                      &gt;&lt;span class="vlist-t vlist-t2"
                                        &gt;&lt;span class="vlist-r"
                                          &gt;&lt;span
                                            class="vlist"
                                            style="
                                              height: 0.16454285714285719em;
                                            "
                                            &gt;&lt;span
                                              style="
                                                top: -2.357em;
                                                margin-left: -0.03588em;
                                                margin-right: 0.07142857142857144em;
                                              "
                                              &gt;&lt;span
                                                class="pstrut"
                                                style="height: 2.5em"
                                              &gt;&lt;/span
                                              &gt;&lt;span
                                                class="sizing reset-size3 size1 mtight"
                                                &gt;&lt;span
                                                  class="mord mathdefault mtight"
                                                  style="
                                                    margin-right: 0.02691em;
                                                  "
                                                  &gt;w&lt;/span
                                                &gt;&lt;/span
                                              &gt;&lt;/span
                                            &gt;&lt;/span
                                          &gt;&lt;span class="vlist-s"&gt;​&lt;/span&gt;&lt;/span
                                        &gt;&lt;span class="vlist-r"
                                          &gt;&lt;span
                                            class="vlist"
                                            style="height: 0.143em"
                                            &gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span
                                  &gt;&lt;span class="mpunct mtight"&gt;,&lt;/span
                                  &gt;&lt;span class="mord mtight"
                                    &gt;&lt;span
                                      class="mord mathdefault mtight"
                                      style="margin-right: 0.03588em"
                                      &gt;y&lt;/span
                                    &gt;&lt;span class="msupsub"
                                      &gt;&lt;span class="vlist-t vlist-t2"
                                        &gt;&lt;span class="vlist-r"
                                          &gt;&lt;span
                                            class="vlist"
                                            style="height: 0.3448em"
                                            &gt;&lt;span
                                              style="
                                                top: -2.3487714285714287em;
                                                margin-left: -0.03588em;
                                                margin-right: 0.07142857142857144em;
                                              "
                                              &gt;&lt;span
                                                class="pstrut"
                                                style="height: 2.5em"
                                              &gt;&lt;/span
                                              &gt;&lt;span
                                                class="sizing reset-size3 size1 mtight"
                                                &gt;&lt;span
                                                  class="mord mathdefault mtight"
                                                  style="
                                                    margin-right: 0.01968em;
                                                  "
                                                  &gt;l&lt;/span
                                                &gt;&lt;/span
                                              &gt;&lt;/span
                                            &gt;&lt;/span
                                          &gt;&lt;span class="vlist-s"&gt;​&lt;/span&gt;&lt;/span
                                        &gt;&lt;span class="vlist-r"
                                          &gt;&lt;span
                                            class="vlist"
                                            style="
                                              height: 0.15122857142857138em;
                                            "
                                            &gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span
                                  &gt;&lt;span class="mclose mtight"&gt;)&lt;/span
                                  &gt;&lt;span class="mrel mtight"&gt;∼&lt;/span
                                  &gt;&lt;span class="mord mtight"
                                    &gt;&lt;span
                                      class="mord mathcal mtight"
                                      style="margin-right: 0.02778em"
                                      &gt;D&lt;/span
                                    &gt;&lt;/span
                                  &gt;&lt;/span
                                &gt;&lt;/span
                              &gt;&lt;/span
                            &gt;&lt;/span
                          &gt;&lt;span class="vlist-s"&gt;​&lt;/span&gt;&lt;/span
                        &gt;&lt;span class="vlist-r"
                          &gt;&lt;span
                            class="vlist"
                            style="height: 0.3551999999999999em"
                            &gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span
                  &gt;&lt;span
                    class="mspace"
                    style="margin-right: 0.16666666666666666em"
                  &gt;&lt;/span
                  &gt;&lt;span class="minner"
                    &gt;&lt;span class="mopen delimcenter" style="top: 0em"
                      &gt;&lt;span class="delimsizing size3"&gt;[&lt;/span&gt;&lt;/span
                    &gt;&lt;span class="mop"
                      &gt;lo&lt;span style="margin-right: 0.01389em"&gt;g&lt;/span&gt;&lt;/span
                    &gt;&lt;span
                      class="mspace"
                      style="margin-right: 0.16666666666666666em"
                    &gt;&lt;/span
                    &gt;&lt;span
                      class="mord mathdefault"
                      style="margin-right: 0.03588em"
                      &gt;σ&lt;/span
                    &gt;&lt;span
                      class="mspace"
                      style="margin-right: 0.16666666666666666em"
                    &gt;&lt;/span
                    &gt;&lt;span class="minner"
                      &gt;&lt;span class="mopen delimcenter" style="top: 0em"
                        &gt;&lt;span class="delimsizing size3"&gt;(&lt;/span&gt;&lt;/span
                      &gt;&lt;span class="mord"
                        &gt;&lt;span class="mopen nulldelimiter"&gt;&lt;/span
                        &gt;&lt;span class="mfrac"
                          &gt;&lt;span class="vlist-t vlist-t2"
                            &gt;&lt;span class="vlist-r"
                              &gt;&lt;span class="vlist" style="height: 1.427em"
                                &gt;&lt;span style="top: -2.314em"
                                  &gt;&lt;span
                                    class="pstrut"
                                    style="height: 3em"
                                  &gt;&lt;/span
                                  &gt;&lt;span class="mord"
                                    &gt;&lt;span class="mord"&gt;∣&lt;/span
                                    &gt;&lt;span class="mord"
                                      &gt;&lt;span
                                        class="mord mathdefault"
                                        style="margin-right: 0.03588em"
                                        &gt;y&lt;/span
                                      &gt;&lt;span class="msupsub"
                                        &gt;&lt;span class="vlist-t vlist-t2"
                                          &gt;&lt;span class="vlist-r"
                                            &gt;&lt;span
                                              class="vlist"
                                              style="height: 0.151392em"
                                              &gt;&lt;span
                                                style="
                                                  top: -2.5500000000000003em;
                                                  margin-left: -0.03588em;
                                                  margin-right: 0.05em;
                                                "
                                                &gt;&lt;span
                                                  class="pstrut"
                                                  style="height: 2.7em"
                                                &gt;&lt;/span
                                                &gt;&lt;span
                                                  class="sizing reset-size6 size3 mtight"
                                                  &gt;&lt;span
                                                    class="mord mathdefault mtight"
                                                    style="
                                                      margin-right: 0.02691em;
                                                    "
                                                    &gt;w&lt;/span
                                                  &gt;&lt;/span
                                                &gt;&lt;/span
                                              &gt;&lt;/span
                                            &gt;&lt;span class="vlist-s"
                                              &gt;​&lt;/span
                                            &gt;&lt;/span
                                          &gt;&lt;span class="vlist-r"
                                            &gt;&lt;span
                                              class="vlist"
                                              style="height: 0.15em"
                                              &gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span
                                    &gt;&lt;span class="mord"&gt;∣&lt;/span&gt;&lt;/span
                                  &gt;&lt;/span
                                &gt;&lt;span style="top: -3.23em"
                                  &gt;&lt;span
                                    class="pstrut"
                                    style="height: 3em"
                                  &gt;&lt;/span
                                  &gt;&lt;span
                                    class="frac-line"
                                    style="border-bottom-width: 0.04em"
                                  &gt;&lt;/span&gt;&lt;/span
                                &gt;&lt;span style="top: -3.677em"
                                  &gt;&lt;span
                                    class="pstrut"
                                    style="height: 3em"
                                  &gt;&lt;/span
                                  &gt;&lt;span class="mord"
                                    &gt;&lt;span class="mop"
                                      &gt;lo&lt;span style="margin-right: 0.01389em"
                                        &gt;g&lt;/span
                                      &gt;&lt;/span
                                    &gt;&lt;span
                                      class="mspace"
                                      style="
                                        margin-right: 0.16666666666666666em;
                                      "
                                    &gt;&lt;/span
                                    &gt;&lt;span class="mord"
                                      &gt;&lt;span
                                        class="mord mathdefault"
                                        style="margin-right: 0.03588em"
                                        &gt;π&lt;/span
                                      &gt;&lt;span class="msupsub"
                                        &gt;&lt;span class="vlist-t vlist-t2"
                                          &gt;&lt;span class="vlist-r"
                                            &gt;&lt;span
                                              class="vlist"
                                              style="
                                                height: 0.33610799999999996em;
                                              "
                                              &gt;&lt;span
                                                style="
                                                  top: -2.5500000000000003em;
                                                  margin-left: -0.03588em;
                                                  margin-right: 0.05em;
                                                "
                                                &gt;&lt;span
                                                  class="pstrut"
                                                  style="height: 2.7em"
                                                &gt;&lt;/span
                                                &gt;&lt;span
                                                  class="sizing reset-size6 size3 mtight"
                                                  &gt;&lt;span
                                                    class="mord mathdefault mtight"
                                                    style="
                                                      margin-right: 0.02778em;
                                                    "
                                                    &gt;θ&lt;/span
                                                  &gt;&lt;/span
                                                &gt;&lt;/span
                                              &gt;&lt;/span
                                            &gt;&lt;span class="vlist-s"
                                              &gt;​&lt;/span
                                            &gt;&lt;/span
                                          &gt;&lt;span class="vlist-r"
                                            &gt;&lt;span
                                              class="vlist"
                                              style="height: 0.15em"
                                              &gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span
                                    &gt;&lt;span class="mopen"&gt;(&lt;/span
                                    &gt;&lt;span class="mord"
                                      &gt;&lt;span
                                        class="mord mathdefault"
                                        style="margin-right: 0.03588em"
                                        &gt;y&lt;/span
                                      &gt;&lt;span class="msupsub"
                                        &gt;&lt;span class="vlist-t vlist-t2"
                                          &gt;&lt;span class="vlist-r"
                                            &gt;&lt;span
                                              class="vlist"
                                              style="height: 0.151392em"
                                              &gt;&lt;span
                                                style="
                                                  top: -2.5500000000000003em;
                                                  margin-left: -0.03588em;
                                                  margin-right: 0.05em;
                                                "
                                                &gt;&lt;span
                                                  class="pstrut"
                                                  style="height: 2.7em"
                                                &gt;&lt;/span
                                                &gt;&lt;span
                                                  class="sizing reset-size6 size3 mtight"
                                                  &gt;&lt;span
                                                    class="mord mathdefault mtight"
                                                    style="
                                                      margin-right: 0.02691em;
                                                    "
                                                    &gt;w&lt;/span
                                                  &gt;&lt;/span
                                                &gt;&lt;/span
                                              &gt;&lt;/span
                                            &gt;&lt;span class="vlist-s"
                                              &gt;​&lt;/span
                                            &gt;&lt;/span
                                          &gt;&lt;span class="vlist-r"
                                            &gt;&lt;span
                                              class="vlist"
                                              style="height: 0.15em"
                                              &gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span
                                    &gt;&lt;span
                                      class="mspace"
                                      style="margin-right: 0.2777777777777778em"
                                    &gt;&lt;/span
                                    &gt;&lt;span class="mrel"&gt;∣&lt;/span
                                    &gt;&lt;span
                                      class="mspace"
                                      style="margin-right: 0.2777777777777778em"
                                    &gt;&lt;/span
                                    &gt;&lt;span class="mord mathdefault"&gt;x&lt;/span
                                    &gt;&lt;span class="mclose"&gt;)&lt;/span&gt;&lt;/span
                                  &gt;&lt;/span
                                &gt;&lt;/span
                              &gt;&lt;span class="vlist-s"&gt;​&lt;/span&gt;&lt;/span
                            &gt;&lt;span class="vlist-r"
                              &gt;&lt;span class="vlist" style="height: 0.936em"
                                &gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span
                        &gt;&lt;span class="mclose nulldelimiter"&gt;&lt;/span&gt;&lt;/span
                      &gt;&lt;span
                        class="mspace"
                        style="margin-right: 0.2222222222222222em"
                      &gt;&lt;/span
                      &gt;&lt;span class="mbin"&gt;−&lt;/span
                      &gt;&lt;span
                        class="mspace"
                        style="margin-right: 0.2222222222222222em"
                      &gt;&lt;/span
                      &gt;&lt;span class="mord"
                        &gt;&lt;span class="mopen nulldelimiter"&gt;&lt;/span
                        &gt;&lt;span class="mfrac"
                          &gt;&lt;span class="vlist-t vlist-t2"
                            &gt;&lt;span class="vlist-r"
                              &gt;&lt;span class="vlist" style="height: 1.427em"
                                &gt;&lt;span style="top: -2.314em"
                                  &gt;&lt;span
                                    class="pstrut"
                                    style="height: 3em"
                                  &gt;&lt;/span
                                  &gt;&lt;span class="mord"
                                    &gt;&lt;span class="mord"&gt;∣&lt;/span
                                    &gt;&lt;span class="mord"
                                      &gt;&lt;span
                                        class="mord mathdefault"
                                        style="margin-right: 0.03588em"
                                        &gt;y&lt;/span
                                      &gt;&lt;span class="msupsub"
                                        &gt;&lt;span class="vlist-t vlist-t2"
                                          &gt;&lt;span class="vlist-r"
                                            &gt;&lt;span
                                              class="vlist"
                                              style="
                                                height: 0.33610799999999996em;
                                              "
                                              &gt;&lt;span
                                                style="
                                                  top: -2.5500000000000003em;
                                                  margin-left: -0.03588em;
                                                  margin-right: 0.05em;
                                                "
                                                &gt;&lt;span
                                                  class="pstrut"
                                                  style="height: 2.7em"
                                                &gt;&lt;/span
                                                &gt;&lt;span
                                                  class="sizing reset-size6 size3 mtight"
                                                  &gt;&lt;span
                                                    class="mord mathdefault mtight"
                                                    style="
                                                      margin-right: 0.01968em;
                                                    "
                                                    &gt;l&lt;/span
                                                  &gt;&lt;/span
                                                &gt;&lt;/span
                                              &gt;&lt;/span
                                            &gt;&lt;span class="vlist-s"
                                              &gt;​&lt;/span
                                            &gt;&lt;/span
                                          &gt;&lt;span class="vlist-r"
                                            &gt;&lt;span
                                              class="vlist"
                                              style="height: 0.15em"
                                              &gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span
                                    &gt;&lt;span class="mord"&gt;∣&lt;/span&gt;&lt;/span
                                  &gt;&lt;/span
                                &gt;&lt;span style="top: -3.23em"
                                  &gt;&lt;span
                                    class="pstrut"
                                    style="height: 3em"
                                  &gt;&lt;/span
                                  &gt;&lt;span
                                    class="frac-line"
                                    style="border-bottom-width: 0.04em"
                                  &gt;&lt;/span&gt;&lt;/span
                                &gt;&lt;span style="top: -3.677em"
                                  &gt;&lt;span
                                    class="pstrut"
                                    style="height: 3em"
                                  &gt;&lt;/span
                                  &gt;&lt;span class="mord"
                                    &gt;&lt;span class="mop"
                                      &gt;lo&lt;span style="margin-right: 0.01389em"
                                        &gt;g&lt;/span
                                      &gt;&lt;/span
                                    &gt;&lt;span
                                      class="mspace"
                                      style="
                                        margin-right: 0.16666666666666666em;
                                      "
                                    &gt;&lt;/span
                                    &gt;&lt;span class="mord"
                                      &gt;&lt;span
                                        class="mord mathdefault"
                                        style="margin-right: 0.03588em"
                                        &gt;π&lt;/span
                                      &gt;&lt;span class="msupsub"
                                        &gt;&lt;span class="vlist-t vlist-t2"
                                          &gt;&lt;span class="vlist-r"
                                            &gt;&lt;span
                                              class="vlist"
                                              style="
                                                height: 0.33610799999999996em;
                                              "
                                              &gt;&lt;span
                                                style="
                                                  top: -2.5500000000000003em;
                                                  margin-left: -0.03588em;
                                                  margin-right: 0.05em;
                                                "
                                                &gt;&lt;span
                                                  class="pstrut"
                                                  style="height: 2.7em"
                                                &gt;&lt;/span
                                                &gt;&lt;span
                                                  class="sizing reset-size6 size3 mtight"
                                                  &gt;&lt;span
                                                    class="mord mathdefault mtight"
                                                    style="
                                                      margin-right: 0.02778em;
                                                    "
                                                    &gt;θ&lt;/span
                                                  &gt;&lt;/span
                                                &gt;&lt;/span
                                              &gt;&lt;/span
                                            &gt;&lt;span class="vlist-s"
                                              &gt;​&lt;/span
                                            &gt;&lt;/span
                                          &gt;&lt;span class="vlist-r"
                                            &gt;&lt;span
                                              class="vlist"
                                              style="height: 0.15em"
                                              &gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span
                                    &gt;&lt;span class="mopen"&gt;(&lt;/span
                                    &gt;&lt;span class="mord"
                                      &gt;&lt;span
                                        class="mord mathdefault"
                                        style="margin-right: 0.03588em"
                                        &gt;y&lt;/span
                                      &gt;&lt;span class="msupsub"
                                        &gt;&lt;span class="vlist-t vlist-t2"
                                          &gt;&lt;span class="vlist-r"
                                            &gt;&lt;span
                                              class="vlist"
                                              style="
                                                height: 0.33610799999999996em;
                                              "
                                              &gt;&lt;span
                                                style="
                                                  top: -2.5500000000000003em;
                                                  margin-left: -0.03588em;
                                                  margin-right: 0.05em;
                                                "
                                                &gt;&lt;span
                                                  class="pstrut"
                                                  style="height: 2.7em"
                                                &gt;&lt;/span
                                                &gt;&lt;span
                                                  class="sizing reset-size6 size3 mtight"
                                                  &gt;&lt;span
                                                    class="mord mathdefault mtight"
                                                    style="
                                                      margin-right: 0.01968em;
                                                    "
                                                    &gt;l&lt;/span
                                                  &gt;&lt;/span
                                                &gt;&lt;/span
                                              &gt;&lt;/span
                                            &gt;&lt;span class="vlist-s"
                                              &gt;​&lt;/span
                                            &gt;&lt;/span
                                          &gt;&lt;span class="vlist-r"
                                            &gt;&lt;span
                                              class="vlist"
                                              style="height: 0.15em"
                                              &gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span
                                    &gt;&lt;span
                                      class="mspace"
                                      style="margin-right: 0.2777777777777778em"
                                    &gt;&lt;/span
                                    &gt;&lt;span class="mrel"&gt;∣&lt;/span
                                    &gt;&lt;span
                                      class="mspace"
                                      style="margin-right: 0.2777777777777778em"
                                    &gt;&lt;/span
                                    &gt;&lt;span class="mord mathdefault"&gt;x&lt;/span
                                    &gt;&lt;span class="mclose"&gt;)&lt;/span&gt;&lt;/span
                                  &gt;&lt;/span
                                &gt;&lt;/span
                              &gt;&lt;span class="vlist-s"&gt;​&lt;/span&gt;&lt;/span
                            &gt;&lt;span class="vlist-r"
                              &gt;&lt;span class="vlist" style="height: 0.936em"
                                &gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span
                        &gt;&lt;span class="mclose nulldelimiter"&gt;&lt;/span&gt;&lt;/span
                      &gt;&lt;span class="mclose delimcenter" style="top: 0em"
                        &gt;&lt;span class="delimsizing size3"&gt;)&lt;/span&gt;&lt;/span
                      &gt;&lt;/span
                    &gt;&lt;span class="mclose delimcenter" style="top: 0em"
                      &gt;&lt;span class="delimsizing size3"&gt;]&lt;/span&gt;&lt;/span
                    &gt;&lt;/span
                  &gt;&lt;/span
                &gt;&lt;/span
              &gt;&lt;/span
            &gt;&lt;/span
          &gt;
        &lt;/p&gt;

        &lt;p&gt;
          And so we arrive at a length-normalized preference learning objective.
          There are some interesting takeaways from this derivation.
        &lt;/p&gt;
        &lt;ol&gt;
          &lt;li&gt;
            If we’re going to annotate preference data using other models, then
            we can actually adjust our objectives to counteract known biases in
            the scoring. Length normalization is one example of this strategy.
            If and when other ones come to light, we could feasibly use this
            approach to allow our models to learn from imperfectly annotated
            data. We can analogously use this approach to correct biases in
            human-annotated data as well, though we ought to exercise a lot more
            caution and thought when shifting or otherwise modifying
            human-expressed preferences.
          &lt;/li&gt;
          &lt;li&gt;
            Note that no matter what I do, this current strategy cannot link
            $r_{\text{LM}}$ to $r_\text{h}$ – that is, I can’t really describe
            how the rewards that the model provides relate to the ground-truth
            human rewards. Instead, all I can say is that length-normalized
            objectives provide one possible strategy for mitigating the
            absorption of biases from the models used to annotate the data. By
            the same token, it doesn’t matter if the annotator LM was trained on
            human data or not. These nuances are something I hope to explore in
            future work!
          &lt;/li&gt;
          &lt;li&gt;
            Connection to &lt;strong&gt;average reward maximization&lt;/strong&gt;: Another
            way to see $r^\star(x,y)$ is that it’s the average reward of each
            token in the sequence. This links length normalization to a long
            line of RL algorithms that maximize the average reward instead of
            the total reward (see a survey in
            &lt;a href="https://link.springer.com/article/10.1007/BF00114727"
              &gt;Mahadevan, 1996&lt;/a
            &gt;). This formulation is especially useful when dealing with cyclical
            or infinite-horizon problems. Note that another way to deal with
            infinite-horizon problems is to discount rewards that will arrive in
            the future – if you were to do this, then you would arrive at the
            length-regularized DPO objective (&lt;a
              href="https://arxiv.org/abs/2403.19159"
              &gt;Park et al., 2024&lt;/a
            &gt;).
          &lt;/li&gt;
          &lt;li&gt;
            If we just look at the final length-normalized objective, then we
            can see that it explicitly requires $\pi_\theta(y_w\mid x)$ to
            increase much more than it would if there was no normalization.
            Other modifications to the objective, like target reward margins, do
            not &lt;em&gt;explicitly&lt;/em&gt; ensure that this will happen, because they
            operate on the distance between the two likelihoods. After all is
            said and done, we usually just use $\pi_\theta$ to generate
            sequences, so directly modifying it seems like it would be more
            effective.
          &lt;/li&gt;
        &lt;/ol&gt;

        &lt;h1 id="discussion"&gt;Discussion&lt;/h1&gt;
        &lt;p&gt;
          The derivation above demonstrates how we might adjust our objective
          functions to enable learning from imperfectly (e.g., automatically)
          annotated data. I think that using other models to grade or annotate
          preference data will become increasingly prevalent, due to growing
          evidence that preference learning with on-policy data is more
          effective (&lt;a href="https://arxiv.org/abs/2404.14367"
            &gt;Tajwar et al., 2024&lt;/a
          &gt;) and the rising cost of finding and hiring qualified human
          annotators. In this post, I considered the case where the preference
          data was model-annotated and the model had a bias favoring longer
          responses. The analysis provides some insight into one particular
          situation in which length normalization makes sense, but it by no
          means provides a general prescriptive guideline for objective design.
        &lt;/p&gt;

        &lt;p&gt;
          I was originally motivated to derive this after reading about SimPO
          (&lt;a href="https://arxiv.org/abs/2405.14734"&gt;Meng et al., 2024&lt;/a&gt;),
          which proposes a length-normalized preference learning objective. If
          you are curious about the other modifications in SimPO – removing the
          reference model and setting a global target reward margin – then you
          might be interested in reading
          &lt;a href="https://arxiv.org/abs/2405.19534"&gt;our recent paper&lt;/a&gt; on how
          the reference model can prevent DPO from aligning effectively.
        &lt;/p&gt;

        &lt;p&gt;
          The possibility of infinite-length strings has been explored in
          several prior works, mainly with a focus on characterizing how prone a
          particular model architecture (e.g., transformer or RNN) is to leaking
          probability mass to infinite-length strings and how certain objectives
          might exacerbate or mitigate this issue (&lt;a
            href="https://aclanthology.org/2020.emnlp-main.448/"
            &gt;Welleck et al., 2020&lt;/a
          &gt;;
          &lt;a href="https://aclanthology.org/2023.acl-long.543/"
            &gt;Du et al., 2023&lt;/a
          &gt;).
        &lt;/p&gt;

        &lt;p&gt;
          There is also a big leap missing from what I have derived so far. We
          may train a policy $\pi^\star$ to maximize rewards but when we sample
          from the model (e.g., via greedy decoding), we don’t actually obtain a
          sequence that maximizes the log likelihood prescribed by the model.
          This was, after all, the motivation for developing many different
          decoding strategies, including ones that implicitly plan or search. In
          the context of post-training, I find this observation fascinating,
          given how the field has primarily focused on next-token prediction
          objectives as a means to constrain or guide long-form generations.&lt;sup
            id="fnref:1"
            role="doc-noteref"
            &gt;&lt;a href="#fn:1" class="footnote" rel="footnote"&gt;1&lt;/a&gt;&lt;/sup
          &gt;
          Some of my research going forward will be on this topic of the gap
          between next-token prediction and long-form generation. If you are
          interested in this direction, I enjoyed reading several of
          &lt;a href="https://cimeister.github.io/publication/"
            &gt;Clara Meister’s papers&lt;/a
          &gt;.
        &lt;/p&gt;

        &lt;p&gt;
          Since this is a blog post and not a paper, I’m not doing a full
          literature survey on these topics. As I mentioned, there are many RL
          papers on these topics, and my main research area is not RL. That
          being said, if I didn’t mention your paper and you think it is
          relevant, kindly send me an email!
        &lt;/p&gt;

        &lt;p&gt;
          &lt;strong&gt;Acknowledgments&lt;/strong&gt;: Thank you (in alphabetical order) to
          Angelica Chen, Xinyi Chen, Yu Meng, and Mengzhou Xia for discussions
          that helped me refine my derivation and for suggesting several
          relevant papers. And also thank you to Adithya Bhaskar, Tianyu Gao,
          Lucy He, and Kaifeng Lyu for helping me proofread this post. Thanks
          also to Ben Eysenbach for pointing me in the direction of average
          reward maximization.
        &lt;/p&gt;

        &lt;p&gt;
          &lt;strong&gt;Citation&lt;/strong&gt;: If you find this blog post to be helpful in
          your work, please use the following bibtex citation.
        &lt;/p&gt;
        &lt;div class="language-plaintext highlighter-rouge"&gt;
          &lt;div class="highlight"&gt;
            &lt;pre class="highlight"&gt;&lt;code&gt;@misc{malladi2024hiddeninfinity,
  title   = {The Hidden Infinity in Preference Learning},
  author  = {Malladi, Sadhika},
  year    = {2024},
  month   = {July},
  url     = {https://www.cs.princeton.edu/~smalladi/blog/2024/06/27/dpo-infinity/},
}
&lt;/code&gt;&lt;/pre&gt;
          &lt;/div&gt;
        &lt;/div&gt;

        &lt;hr /&gt;

        &lt;div class="footnotes" role="doc-endnotes"&gt;
          &lt;ol&gt;
            &lt;li id="fn:1" role="doc-endnote"&gt;
              &lt;p&gt;
                OK, to be fair, there are papers that now train models to
                predict several tokens into the future (&lt;a
                  href="https://arxiv.org/abs/2404.19737"
                  &gt;Gloeckle et al., 2024&lt;/a
                &gt;), but I would argue this is just a patch over the underlying
                problem. Similarly, there are papers that allow a model to
                access and train on its own generations (e.g., STaR,
                &lt;a href="https://arxiv.org/abs/2203.14465"
                  &gt;Zelikman et al., 2022&lt;/a
                &gt;), but I’m not fully convinced that this closes the gap between
                training and inference. &lt;a
                  href="#fnref:1"
                  class="reversefootnote"
                  role="doc-backlink"
                  &gt;&amp;#8617;&lt;/a
                &gt;
              &lt;/p&gt;
            &lt;/li&gt;
          &lt;/ol&gt;
        &lt;/div&gt;
      </content></entry><entry><title>Using LESS Data to Tune Models</title><id>https://sadhikamalladi.github.io/blog/2024/04/04/dataselection/</id><link href="https://sadhikamalladi.github.io/blog/2024/04/04/dataselection/" rel="alternate" type="text/html" /><published>2024-04-04T00:00:00-07:00</published><updated>2024-04-04T00:00:00-07:00</updated><author><name>Mengzhou Xia and Sadhika Malladi</name></author><summary>Data Selection in the Era of LLMs</summary><content type="html">
        &lt;p&gt;
          &lt;strong&gt;TL;DR&lt;/strong&gt;: We describe how data selection for modern-day
          LLMs differs from prior settings and how our algorithm,
          &lt;strong&gt;LESS&lt;/strong&gt;, effectively selects relevant data to cultivate
          specific capabilities in models during instruction tuning.
        &lt;/p&gt;

        &lt;p&gt;
          &lt;strong&gt;Paper:&lt;/strong&gt;
          &lt;a href="https://arxiv.org/abs/2402.04333"
            &gt;https://arxiv.org/abs/2402.04333&lt;/a
          &gt;
        &lt;/p&gt;

        &lt;p&gt;
          &lt;strong&gt;Code:&lt;/strong&gt;
          &lt;a href="https://github.com/princeton-nlp/LESS/tree/main"
            &gt;https://github.com/princeton-nlp/LESS/&lt;/a
          &gt;
        &lt;/p&gt;

        &lt;hr /&gt;

        &lt;p&gt;This post will take the following structure:&lt;/p&gt;

        &lt;ul&gt;
          &lt;li&gt;
            We introduce the motivation for data selection and describe how the
            criteria for “good” data depends heavily on the
            &lt;a href="#motivation"&gt;setting&lt;/a&gt;. One can either try to identify
            &lt;strong&gt;representative&lt;/strong&gt; datapoints for the in-domain setting
            or &lt;strong&gt;relevant&lt;/strong&gt; ones for the transfer setting.
          &lt;/li&gt;
          &lt;li&gt;
            Our algorithm, &lt;a href="#less"&gt;LESS&lt;/a&gt;, effectively selects
            relevant data to induce capabilities in the instruction tuning
            setting. LESS identifies 5% of the dataset that induces stronger
            performance than training on the full dataset.
          &lt;/li&gt;
          &lt;li&gt;
            We conduct an
            &lt;a href="#prior-work"&gt;in-depth analysis of prior works&lt;/a&gt; on data
            selection for various settings and provide insights into their
            technical details, strengths, and limitations.
          &lt;/li&gt;
          &lt;li&gt;
            We conclude by identifying trends in data selection in the era of
            LLMs.
          &lt;/li&gt;
        &lt;/ul&gt;

        &lt;hr /&gt;

        &lt;h1 id="motivation"&gt;Motivation&lt;/h1&gt;

        &lt;p&gt;
          The training dataset is a crucial design choice when building a
          machine learning model. Dataset choice can drive the
          &lt;strong&gt;capabilities&lt;/strong&gt; of the resulting model in various ways
          (see, for example,
          &lt;a href="https://arxiv.org/abs/2308.12950"&gt;CodeLLaMA&lt;/a&gt;, and
          &lt;a href="https://arxiv.org/abs/2402.03300"&gt;DeepSeek-Math&lt;/a&gt;). Also,
          training models can be expensive, and the cost usually scales with the
          size of the dataset, so dataset selection offers one way to improve
          &lt;strong&gt;efficiency&lt;/strong&gt; and reduce cost&lt;strong&gt;.&lt;/strong&gt;
        &lt;/p&gt;

        &lt;center&gt;
          &lt;div class="figure"&gt;
            &lt;img
              src="/assets/dataselection_img/settings.svg"
              alt="Cartoon of coreset selection vs transfer data selection."
              style="margin-left: auto; margin-right: auto; width: 95%"
            /&gt;
            &lt;br /&gt;

            &lt;div class="caption"&gt;
              &lt;span class="caption-label"
                &gt;&lt;i&gt;Coreset selection&lt;/i&gt; selects data such that the selected
                subset represents the full dataset. &lt;br /&gt;&lt;i
                  &gt;Transfer data selection&lt;/i
                &gt;
                selects the subset that is closest to the target data points.
              &lt;/span&gt;
            &lt;/div&gt;
          &lt;/div&gt;
        &lt;/center&gt;

        &lt;p&gt;
          We distinguish two settings for data selection: in-domain data
          selection and transfer data selection. In the former, the selected
          data is drawn from the same distribution as the evaluation data&lt;sup
            id="fnref:1"
            role="doc-noteref"
            &gt;&lt;a href="#fn:1" class="footnote" rel="footnote"&gt;1&lt;/a&gt;&lt;/sup
          &gt;
          , whereas in the latter, evaluation is performed on different data.
        &lt;/p&gt;

        &lt;ul&gt;
          &lt;li&gt;
            &lt;em&gt;In-domain data selection&lt;/em&gt; aims to identify the most
            &lt;strong&gt;representative&lt;/strong&gt; subset of data, often referred to as
            a coreset (&lt;a href="https://arxiv.org/abs/1708.00489"
              &gt;Sener &amp;amp; Severese 2018&lt;/a
            &gt;), from a large in-domain training dataset. The selection criteria
            include representativeness, coverage, correctness, and more. We talk
            more about this setting latter.
          &lt;/li&gt;
          &lt;li&gt;
            &lt;em&gt;Transfer data selection&lt;/em&gt; seeks to choose the most
            &lt;strong&gt;relevant&lt;/strong&gt; data from a broad training pool that
            closely align with the target examples. The selection criterion
            focuses on relevancy.
          &lt;/li&gt;
        &lt;/ul&gt;

        &lt;p&gt;
          Our work focuses on &lt;strong&gt;transfer data selection for&lt;/strong&gt;
          &lt;strong&gt;instruction tuning&lt;/strong&gt;. Instruction tuning has proven to
          be a highly effective way to quickly adapt language models to follow
          human instructions. Depending on the data used, models can be tuned to
          be general-purpose instruction followers (e.g.,
          &lt;a href="https://crfm.stanford.edu/2023/03/13/alpaca.html"&gt;Alpaca&lt;/a&gt;,
          &lt;a href="https://lmsys.org/blog/2023-03-30-vicuna/"&gt;Vicuna&lt;/a&gt;,
          &lt;a href="https://arxiv.org/abs/2310.16944"&gt;Zephyr&lt;/a&gt;) or solve more
          structured tasks per human instructions (i.e.,
          &lt;strong&gt;targeted instruction tuning&lt;/strong&gt;). Our work focuses on
          selecting data for the latter case, where models are tuned to perform
          particular types of reasoning (e.g., using a passage to answer a
          question). In this case, the data selection problem can be understood
          as bootstrapping a few examples to identify relevant data to solve a
          task.
        &lt;/p&gt;

        &lt;h1 id="less"&gt;
          Selecting LESS Datapoints for Targeted Instruction Tuning
        &lt;/h1&gt;

        &lt;p&gt;
          The targeted instruction tuning setting poses the following research
          question:
        &lt;/p&gt;

        &lt;center&gt;
          &lt;i
            &gt;Given just a few handwritten examples of a query type and a
            particular pre-trained model, how can we identify the most relevant
            data to train on out of a large pool of available instruction tuning
            data?&lt;/i
          &gt;
        &lt;/center&gt;

        &lt;p&gt;
          Several data selection strategies have been developed for
          pre-training, such as continued training on domain-specific data (&lt;a
            href="https://arxiv.org/abs/2004.10964"
            &gt;Gururangan et al., 2020&lt;/a
          &gt;) and using n-gram statistics to identify relevant data (&lt;a
            href="https://arxiv.org/abs/2302.03169"
            &gt;Xie et al., 2023&lt;/a
          &gt;). However, the instruction tuning setting is unique in that using
          all of the available data can hurt the development of specific
          capabilities (&lt;a href="https://arxiv.org/abs/2306.04751"
            &gt;Wang et al., 2023&lt;/a
          &gt;), and one wants to somehow account for the properties of the
          pre-trained model when selecting data. So, we avoid using heuristic
          definitions of useful data and instead frame data selection as a
          rigorous optimization problem. As we describe in the next section,
          LESS selects training data to minimize the loss on the target data
          (i.e., the few handwritten validation examples). Check out our paper
          &lt;a href="https://arxiv.org/abs/2402.04333" target="_blank"&gt;here&lt;/a&gt;
          and play with the code on
          &lt;a href="https://github.com/princeton-nlp/LESS" target="_blank"
            &gt;GitHub&lt;/a
          &gt;!
        &lt;/p&gt;

        &lt;h2 id="conceptual-approach"&gt;Conceptual Approach&lt;/h2&gt;

        &lt;p&gt;
          Suppose we have a handwritten validation example $z$ and a huge
          dataset of candidate training points $\mathcal{D}$. At the heart of
          any transfer data selection algorithm is the same question: how does
          training on some point $x\in\mathcal{D}$ affect the model’s
          performance on $z$? We explicitly formulate this by approximating how
          the validation loss $\ell(z;\theta)$ changes when we take one training
          step (i.e., update the model from $\theta_t$ to $\theta_{t+1}$) on a
          candidate datapoint $x$:
        &lt;/p&gt;

        &lt;p&gt;
          &lt;span class="katex-display"
            &gt;&lt;span class="katex"
              &gt;&lt;span class="katex-mathml"
                &gt;&lt;math xmlns="http://www.w3.org/1998/Math/MathML"
                  &gt;&lt;semantics
                    &gt;&lt;mrow
                      &gt;&lt;mi mathvariant="normal"&gt;ℓ&lt;/mi&gt;&lt;mo stretchy="false"&gt;(&lt;/mo
                      &gt;&lt;mi&gt;z&lt;/mi&gt;&lt;mo separator="true"&gt;;&lt;/mo
                      &gt;&lt;msub
                        &gt;&lt;mi&gt;θ&lt;/mi
                        &gt;&lt;mrow&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;mo&gt;+&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/mrow&gt;&lt;/msub
                      &gt;&lt;mo stretchy="false"&gt;)&lt;/mo&gt;&lt;mo&gt;≈&lt;/mo
                      &gt;&lt;mi mathvariant="normal"&gt;ℓ&lt;/mi&gt;&lt;mo stretchy="false"&gt;(&lt;/mo
                      &gt;&lt;mi&gt;z&lt;/mi&gt;&lt;mo separator="true"&gt;;&lt;/mo
                      &gt;&lt;msub&gt;&lt;mi&gt;θ&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/msub
                      &gt;&lt;mo stretchy="false"&gt;)&lt;/mo&gt;&lt;mo&gt;+&lt;/mo
                      &gt;&lt;mo stretchy="false"&gt;⟨&lt;/mo&gt;&lt;mi mathvariant="normal"&gt;∇&lt;/mi
                      &gt;&lt;mi mathvariant="normal"&gt;ℓ&lt;/mi&gt;&lt;mo stretchy="false"&gt;(&lt;/mo
                      &gt;&lt;mi&gt;z&lt;/mi&gt;&lt;mo separator="true"&gt;;&lt;/mo
                      &gt;&lt;msub&gt;&lt;mi&gt;θ&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/msub
                      &gt;&lt;mo stretchy="false"&gt;)&lt;/mo&gt;&lt;mo separator="true"&gt;,&lt;/mo
                      &gt;&lt;msub
                        &gt;&lt;mi&gt;θ&lt;/mi
                        &gt;&lt;mrow&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;mo&gt;+&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/mrow&gt;&lt;/msub
                      &gt;&lt;mo&gt;−&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;θ&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/msub
                      &gt;&lt;mo stretchy="false"&gt;⟩&lt;/mo&gt;&lt;/mrow
                    &gt;&lt;annotation encoding="application/x-tex"&gt;
                      \ell(z;\theta_{t+1}) \approx \ell(z;\theta_t) + \langle
                      \nabla \ell (z;\theta_t), \theta_{t+1} - \theta_t \rangle
                    &lt;/annotation&gt;&lt;/semantics
                  &gt;&lt;/math
                &gt;&lt;/span
              &gt;&lt;span class="katex-html" aria-hidden="true"
                &gt;&lt;span class="base"
                  &gt;&lt;span
                    class="strut"
                    style="height: 1em; vertical-align: -0.25em"
                  &gt;&lt;/span
                  &gt;&lt;span class="mord"&gt;ℓ&lt;/span&gt;&lt;span class="mopen"&gt;(&lt;/span
                  &gt;&lt;span
                    class="mord mathdefault"
                    style="margin-right: 0.04398em"
                    &gt;z&lt;/span
                  &gt;&lt;span class="mpunct"&gt;;&lt;/span
                  &gt;&lt;span
                    class="mspace"
                    style="margin-right: 0.16666666666666666em"
                  &gt;&lt;/span
                  &gt;&lt;span class="mord"
                    &gt;&lt;span
                      class="mord mathdefault"
                      style="margin-right: 0.02778em"
                      &gt;θ&lt;/span
                    &gt;&lt;span class="msupsub"
                      &gt;&lt;span class="vlist-t vlist-t2"
                        &gt;&lt;span class="vlist-r"
                          &gt;&lt;span class="vlist" style="height: 0.301108em"
                            &gt;&lt;span
                              style="
                                top: -2.5500000000000003em;
                                margin-left: -0.02778em;
                                margin-right: 0.05em;
                              "
                              &gt;&lt;span class="pstrut" style="height: 2.7em"&gt;&lt;/span
                              &gt;&lt;span class="sizing reset-size6 size3 mtight"
                                &gt;&lt;span class="mord mtight"
                                  &gt;&lt;span class="mord mathdefault mtight"&gt;t&lt;/span
                                  &gt;&lt;span class="mbin mtight"&gt;+&lt;/span
                                  &gt;&lt;span class="mord mtight"&gt;1&lt;/span&gt;&lt;/span
                                &gt;&lt;/span
                              &gt;&lt;/span
                            &gt;&lt;/span
                          &gt;&lt;span class="vlist-s"&gt;​&lt;/span&gt;&lt;/span
                        &gt;&lt;span class="vlist-r"
                          &gt;&lt;span class="vlist" style="height: 0.208331em"
                            &gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span
                  &gt;&lt;span class="mclose"&gt;)&lt;/span
                  &gt;&lt;span
                    class="mspace"
                    style="margin-right: 0.2777777777777778em"
                  &gt;&lt;/span
                  &gt;&lt;span class="mrel"&gt;≈&lt;/span
                  &gt;&lt;span
                    class="mspace"
                    style="margin-right: 0.2777777777777778em"
                  &gt;&lt;/span&gt;&lt;/span
                &gt;&lt;span class="base"
                  &gt;&lt;span
                    class="strut"
                    style="height: 1em; vertical-align: -0.25em"
                  &gt;&lt;/span
                  &gt;&lt;span class="mord"&gt;ℓ&lt;/span&gt;&lt;span class="mopen"&gt;(&lt;/span
                  &gt;&lt;span
                    class="mord mathdefault"
                    style="margin-right: 0.04398em"
                    &gt;z&lt;/span
                  &gt;&lt;span class="mpunct"&gt;;&lt;/span
                  &gt;&lt;span
                    class="mspace"
                    style="margin-right: 0.16666666666666666em"
                  &gt;&lt;/span
                  &gt;&lt;span class="mord"
                    &gt;&lt;span
                      class="mord mathdefault"
                      style="margin-right: 0.02778em"
                      &gt;θ&lt;/span
                    &gt;&lt;span class="msupsub"
                      &gt;&lt;span class="vlist-t vlist-t2"
                        &gt;&lt;span class="vlist-r"
                          &gt;&lt;span
                            class="vlist"
                            style="height: 0.2805559999999999em"
                            &gt;&lt;span
                              style="
                                top: -2.5500000000000003em;
                                margin-left: -0.02778em;
                                margin-right: 0.05em;
                              "
                              &gt;&lt;span class="pstrut" style="height: 2.7em"&gt;&lt;/span
                              &gt;&lt;span class="sizing reset-size6 size3 mtight"
                                &gt;&lt;span class="mord mathdefault mtight"
                                  &gt;t&lt;/span
                                &gt;&lt;/span
                              &gt;&lt;/span
                            &gt;&lt;/span
                          &gt;&lt;span class="vlist-s"&gt;​&lt;/span&gt;&lt;/span
                        &gt;&lt;span class="vlist-r"
                          &gt;&lt;span class="vlist" style="height: 0.15em"
                            &gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span
                  &gt;&lt;span class="mclose"&gt;)&lt;/span
                  &gt;&lt;span
                    class="mspace"
                    style="margin-right: 0.2222222222222222em"
                  &gt;&lt;/span
                  &gt;&lt;span class="mbin"&gt;+&lt;/span
                  &gt;&lt;span
                    class="mspace"
                    style="margin-right: 0.2222222222222222em"
                  &gt;&lt;/span&gt;&lt;/span
                &gt;&lt;span class="base"
                  &gt;&lt;span
                    class="strut"
                    style="height: 1em; vertical-align: -0.25em"
                  &gt;&lt;/span
                  &gt;&lt;span class="mopen"&gt;⟨&lt;/span&gt;&lt;span class="mord"&gt;∇&lt;/span
                  &gt;&lt;span class="mord"&gt;ℓ&lt;/span&gt;&lt;span class="mopen"&gt;(&lt;/span
                  &gt;&lt;span
                    class="mord mathdefault"
                    style="margin-right: 0.04398em"
                    &gt;z&lt;/span
                  &gt;&lt;span class="mpunct"&gt;;&lt;/span
                  &gt;&lt;span
                    class="mspace"
                    style="margin-right: 0.16666666666666666em"
                  &gt;&lt;/span
                  &gt;&lt;span class="mord"
                    &gt;&lt;span
                      class="mord mathdefault"
                      style="margin-right: 0.02778em"
                      &gt;θ&lt;/span
                    &gt;&lt;span class="msupsub"
                      &gt;&lt;span class="vlist-t vlist-t2"
                        &gt;&lt;span class="vlist-r"
                          &gt;&lt;span
                            class="vlist"
                            style="height: 0.2805559999999999em"
                            &gt;&lt;span
                              style="
                                top: -2.5500000000000003em;
                                margin-left: -0.02778em;
                                margin-right: 0.05em;
                              "
                              &gt;&lt;span class="pstrut" style="height: 2.7em"&gt;&lt;/span
                              &gt;&lt;span class="sizing reset-size6 size3 mtight"
                                &gt;&lt;span class="mord mathdefault mtight"
                                  &gt;t&lt;/span
                                &gt;&lt;/span
                              &gt;&lt;/span
                            &gt;&lt;/span
                          &gt;&lt;span class="vlist-s"&gt;​&lt;/span&gt;&lt;/span
                        &gt;&lt;span class="vlist-r"
                          &gt;&lt;span class="vlist" style="height: 0.15em"
                            &gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span
                  &gt;&lt;span class="mclose"&gt;)&lt;/span&gt;&lt;span class="mpunct"&gt;,&lt;/span
                  &gt;&lt;span
                    class="mspace"
                    style="margin-right: 0.16666666666666666em"
                  &gt;&lt;/span
                  &gt;&lt;span class="mord"
                    &gt;&lt;span
                      class="mord mathdefault"
                      style="margin-right: 0.02778em"
                      &gt;θ&lt;/span
                    &gt;&lt;span class="msupsub"
                      &gt;&lt;span class="vlist-t vlist-t2"
                        &gt;&lt;span class="vlist-r"
                          &gt;&lt;span class="vlist" style="height: 0.301108em"
                            &gt;&lt;span
                              style="
                                top: -2.5500000000000003em;
                                margin-left: -0.02778em;
                                margin-right: 0.05em;
                              "
                              &gt;&lt;span class="pstrut" style="height: 2.7em"&gt;&lt;/span
                              &gt;&lt;span class="sizing reset-size6 size3 mtight"
                                &gt;&lt;span class="mord mtight"
                                  &gt;&lt;span class="mord mathdefault mtight"&gt;t&lt;/span
                                  &gt;&lt;span class="mbin mtight"&gt;+&lt;/span
                                  &gt;&lt;span class="mord mtight"&gt;1&lt;/span&gt;&lt;/span
                                &gt;&lt;/span
                              &gt;&lt;/span
                            &gt;&lt;/span
                          &gt;&lt;span class="vlist-s"&gt;​&lt;/span&gt;&lt;/span
                        &gt;&lt;span class="vlist-r"
                          &gt;&lt;span class="vlist" style="height: 0.208331em"
                            &gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span
                  &gt;&lt;span
                    class="mspace"
                    style="margin-right: 0.2222222222222222em"
                  &gt;&lt;/span
                  &gt;&lt;span class="mbin"&gt;−&lt;/span
                  &gt;&lt;span
                    class="mspace"
                    style="margin-right: 0.2222222222222222em"
                  &gt;&lt;/span&gt;&lt;/span
                &gt;&lt;span class="base"
                  &gt;&lt;span
                    class="strut"
                    style="height: 1em; vertical-align: -0.25em"
                  &gt;&lt;/span
                  &gt;&lt;span class="mord"
                    &gt;&lt;span
                      class="mord mathdefault"
                      style="margin-right: 0.02778em"
                      &gt;θ&lt;/span
                    &gt;&lt;span class="msupsub"
                      &gt;&lt;span class="vlist-t vlist-t2"
                        &gt;&lt;span class="vlist-r"
                          &gt;&lt;span
                            class="vlist"
                            style="height: 0.2805559999999999em"
                            &gt;&lt;span
                              style="
                                top: -2.5500000000000003em;
                                margin-left: -0.02778em;
                                margin-right: 0.05em;
                              "
                              &gt;&lt;span class="pstrut" style="height: 2.7em"&gt;&lt;/span
                              &gt;&lt;span class="sizing reset-size6 size3 mtight"
                                &gt;&lt;span class="mord mathdefault mtight"
                                  &gt;t&lt;/span
                                &gt;&lt;/span
                              &gt;&lt;/span
                            &gt;&lt;/span
                          &gt;&lt;span class="vlist-s"&gt;​&lt;/span&gt;&lt;/span
                        &gt;&lt;span class="vlist-r"
                          &gt;&lt;span class="vlist" style="height: 0.15em"
                            &gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span
                  &gt;&lt;span class="mclose"&gt;⟩&lt;/span&gt;&lt;/span
                &gt;&lt;/span
              &gt;&lt;/span
            &gt;&lt;/span
          &gt;
        &lt;/p&gt;

        &lt;p&gt;
          Assume that we were training with SGD with step size $\eta_t$, we can
          further derive the following formulation:
        &lt;/p&gt;

        &lt;p&gt;
          &lt;span class="katex-display"
            &gt;&lt;span class="katex"
              &gt;&lt;span class="katex-mathml"
                &gt;&lt;math xmlns="http://www.w3.org/1998/Math/MathML"
                  &gt;&lt;semantics
                    &gt;&lt;mrow
                      &gt;&lt;mi mathvariant="normal"&gt;ℓ&lt;/mi&gt;&lt;mo stretchy="false"&gt;(&lt;/mo
                      &gt;&lt;mi&gt;z&lt;/mi&gt;&lt;mo separator="true"&gt;;&lt;/mo
                      &gt;&lt;msub
                        &gt;&lt;mi&gt;θ&lt;/mi
                        &gt;&lt;mrow&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;mo&gt;+&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/mrow&gt;&lt;/msub
                      &gt;&lt;mo stretchy="false"&gt;)&lt;/mo&gt;&lt;mo&gt;−&lt;/mo
                      &gt;&lt;mi mathvariant="normal"&gt;ℓ&lt;/mi&gt;&lt;mo stretchy="false"&gt;(&lt;/mo
                      &gt;&lt;mi&gt;z&lt;/mi&gt;&lt;mo separator="true"&gt;;&lt;/mo
                      &gt;&lt;msub&gt;&lt;mi&gt;θ&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/msub
                      &gt;&lt;mo stretchy="false"&gt;)&lt;/mo&gt;&lt;mo&gt;≈&lt;/mo
                      &gt;&lt;msub&gt;&lt;mi&gt;η&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/msub
                      &gt;&lt;mo stretchy="false"&gt;⟨&lt;/mo&gt;&lt;mi mathvariant="normal"&gt;∇&lt;/mi
                      &gt;&lt;mi mathvariant="normal"&gt;ℓ&lt;/mi&gt;&lt;mo stretchy="false"&gt;(&lt;/mo
                      &gt;&lt;mi&gt;z&lt;/mi&gt;&lt;mo separator="true"&gt;;&lt;/mo
                      &gt;&lt;msub&gt;&lt;mi&gt;θ&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/msub
                      &gt;&lt;mo stretchy="false"&gt;)&lt;/mo&gt;&lt;mo separator="true"&gt;,&lt;/mo
                      &gt;&lt;mi mathvariant="normal"&gt;∇&lt;/mi
                      &gt;&lt;mi mathvariant="normal"&gt;ℓ&lt;/mi&gt;&lt;mo stretchy="false"&gt;(&lt;/mo
                      &gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mo separator="true"&gt;;&lt;/mo
                      &gt;&lt;msub&gt;&lt;mi&gt;θ&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/msub
                      &gt;&lt;mo stretchy="false"&gt;)&lt;/mo
                      &gt;&lt;mo stretchy="false"&gt;⟩&lt;/mo&gt;&lt;/mrow
                    &gt;&lt;annotation encoding="application/x-tex"&gt;
                      \ell(z;\theta_{t+1}) -\ell(z;\theta_t) \approx
                      \eta_t\langle \nabla \ell (z;\theta_t), \nabla \ell
                      (x;\theta_t)\rangle
                    &lt;/annotation&gt;&lt;/semantics
                  &gt;&lt;/math
                &gt;&lt;/span
              &gt;&lt;span class="katex-html" aria-hidden="true"
                &gt;&lt;span class="base"
                  &gt;&lt;span
                    class="strut"
                    style="height: 1em; vertical-align: -0.25em"
                  &gt;&lt;/span
                  &gt;&lt;span class="mord"&gt;ℓ&lt;/span&gt;&lt;span class="mopen"&gt;(&lt;/span
                  &gt;&lt;span
                    class="mord mathdefault"
                    style="margin-right: 0.04398em"
                    &gt;z&lt;/span
                  &gt;&lt;span class="mpunct"&gt;;&lt;/span
                  &gt;&lt;span
                    class="mspace"
                    style="margin-right: 0.16666666666666666em"
                  &gt;&lt;/span
                  &gt;&lt;span class="mord"
                    &gt;&lt;span
                      class="mord mathdefault"
                      style="margin-right: 0.02778em"
                      &gt;θ&lt;/span
                    &gt;&lt;span class="msupsub"
                      &gt;&lt;span class="vlist-t vlist-t2"
                        &gt;&lt;span class="vlist-r"
                          &gt;&lt;span class="vlist" style="height: 0.301108em"
                            &gt;&lt;span
                              style="
                                top: -2.5500000000000003em;
                                margin-left: -0.02778em;
                                margin-right: 0.05em;
                              "
                              &gt;&lt;span class="pstrut" style="height: 2.7em"&gt;&lt;/span
                              &gt;&lt;span class="sizing reset-size6 size3 mtight"
                                &gt;&lt;span class="mord mtight"
                                  &gt;&lt;span class="mord mathdefault mtight"&gt;t&lt;/span
                                  &gt;&lt;span class="mbin mtight"&gt;+&lt;/span
                                  &gt;&lt;span class="mord mtight"&gt;1&lt;/span&gt;&lt;/span
                                &gt;&lt;/span
                              &gt;&lt;/span
                            &gt;&lt;/span
                          &gt;&lt;span class="vlist-s"&gt;​&lt;/span&gt;&lt;/span
                        &gt;&lt;span class="vlist-r"
                          &gt;&lt;span class="vlist" style="height: 0.208331em"
                            &gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span
                  &gt;&lt;span class="mclose"&gt;)&lt;/span
                  &gt;&lt;span
                    class="mspace"
                    style="margin-right: 0.2222222222222222em"
                  &gt;&lt;/span
                  &gt;&lt;span class="mbin"&gt;−&lt;/span
                  &gt;&lt;span
                    class="mspace"
                    style="margin-right: 0.2222222222222222em"
                  &gt;&lt;/span&gt;&lt;/span
                &gt;&lt;span class="base"
                  &gt;&lt;span
                    class="strut"
                    style="height: 1em; vertical-align: -0.25em"
                  &gt;&lt;/span
                  &gt;&lt;span class="mord"&gt;ℓ&lt;/span&gt;&lt;span class="mopen"&gt;(&lt;/span
                  &gt;&lt;span
                    class="mord mathdefault"
                    style="margin-right: 0.04398em"
                    &gt;z&lt;/span
                  &gt;&lt;span class="mpunct"&gt;;&lt;/span
                  &gt;&lt;span
                    class="mspace"
                    style="margin-right: 0.16666666666666666em"
                  &gt;&lt;/span
                  &gt;&lt;span class="mord"
                    &gt;&lt;span
                      class="mord mathdefault"
                      style="margin-right: 0.02778em"
                      &gt;θ&lt;/span
                    &gt;&lt;span class="msupsub"
                      &gt;&lt;span class="vlist-t vlist-t2"
                        &gt;&lt;span class="vlist-r"
                          &gt;&lt;span
                            class="vlist"
                            style="height: 0.2805559999999999em"
                            &gt;&lt;span
                              style="
                                top: -2.5500000000000003em;
                                margin-left: -0.02778em;
                                margin-right: 0.05em;
                              "
                              &gt;&lt;span class="pstrut" style="height: 2.7em"&gt;&lt;/span
                              &gt;&lt;span class="sizing reset-size6 size3 mtight"
                                &gt;&lt;span class="mord mathdefault mtight"
                                  &gt;t&lt;/span
                                &gt;&lt;/span
                              &gt;&lt;/span
                            &gt;&lt;/span
                          &gt;&lt;span class="vlist-s"&gt;​&lt;/span&gt;&lt;/span
                        &gt;&lt;span class="vlist-r"
                          &gt;&lt;span class="vlist" style="height: 0.15em"
                            &gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span
                  &gt;&lt;span class="mclose"&gt;)&lt;/span
                  &gt;&lt;span
                    class="mspace"
                    style="margin-right: 0.2777777777777778em"
                  &gt;&lt;/span
                  &gt;&lt;span class="mrel"&gt;≈&lt;/span
                  &gt;&lt;span
                    class="mspace"
                    style="margin-right: 0.2777777777777778em"
                  &gt;&lt;/span&gt;&lt;/span
                &gt;&lt;span class="base"
                  &gt;&lt;span
                    class="strut"
                    style="height: 1em; vertical-align: -0.25em"
                  &gt;&lt;/span
                  &gt;&lt;span class="mord"
                    &gt;&lt;span
                      class="mord mathdefault"
                      style="margin-right: 0.03588em"
                      &gt;η&lt;/span
                    &gt;&lt;span class="msupsub"
                      &gt;&lt;span class="vlist-t vlist-t2"
                        &gt;&lt;span class="vlist-r"
                          &gt;&lt;span
                            class="vlist"
                            style="height: 0.2805559999999999em"
                            &gt;&lt;span
                              style="
                                top: -2.5500000000000003em;
                                margin-left: -0.03588em;
                                margin-right: 0.05em;
                              "
                              &gt;&lt;span class="pstrut" style="height: 2.7em"&gt;&lt;/span
                              &gt;&lt;span class="sizing reset-size6 size3 mtight"
                                &gt;&lt;span class="mord mathdefault mtight"
                                  &gt;t&lt;/span
                                &gt;&lt;/span
                              &gt;&lt;/span
                            &gt;&lt;/span
                          &gt;&lt;span class="vlist-s"&gt;​&lt;/span&gt;&lt;/span
                        &gt;&lt;span class="vlist-r"
                          &gt;&lt;span class="vlist" style="height: 0.15em"
                            &gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span
                  &gt;&lt;span class="mopen"&gt;⟨&lt;/span&gt;&lt;span class="mord"&gt;∇&lt;/span
                  &gt;&lt;span class="mord"&gt;ℓ&lt;/span&gt;&lt;span class="mopen"&gt;(&lt;/span
                  &gt;&lt;span
                    class="mord mathdefault"
                    style="margin-right: 0.04398em"
                    &gt;z&lt;/span
                  &gt;&lt;span class="mpunct"&gt;;&lt;/span
                  &gt;&lt;span
                    class="mspace"
                    style="margin-right: 0.16666666666666666em"
                  &gt;&lt;/span
                  &gt;&lt;span class="mord"
                    &gt;&lt;span
                      class="mord mathdefault"
                      style="margin-right: 0.02778em"
                      &gt;θ&lt;/span
                    &gt;&lt;span class="msupsub"
                      &gt;&lt;span class="vlist-t vlist-t2"
                        &gt;&lt;span class="vlist-r"
                          &gt;&lt;span
                            class="vlist"
                            style="height: 0.2805559999999999em"
                            &gt;&lt;span
                              style="
                                top: -2.5500000000000003em;
                                margin-left: -0.02778em;
                                margin-right: 0.05em;
                              "
                              &gt;&lt;span class="pstrut" style="height: 2.7em"&gt;&lt;/span
                              &gt;&lt;span class="sizing reset-size6 size3 mtight"
                                &gt;&lt;span class="mord mathdefault mtight"
                                  &gt;t&lt;/span
                                &gt;&lt;/span
                              &gt;&lt;/span
                            &gt;&lt;/span
                          &gt;&lt;span class="vlist-s"&gt;​&lt;/span&gt;&lt;/span
                        &gt;&lt;span class="vlist-r"
                          &gt;&lt;span class="vlist" style="height: 0.15em"
                            &gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span
                  &gt;&lt;span class="mclose"&gt;)&lt;/span&gt;&lt;span class="mpunct"&gt;,&lt;/span
                  &gt;&lt;span
                    class="mspace"
                    style="margin-right: 0.16666666666666666em"
                  &gt;&lt;/span
                  &gt;&lt;span class="mord"&gt;∇&lt;/span&gt;&lt;span class="mord"&gt;ℓ&lt;/span
                  &gt;&lt;span class="mopen"&gt;(&lt;/span
                  &gt;&lt;span class="mord mathdefault"&gt;x&lt;/span
                  &gt;&lt;span class="mpunct"&gt;;&lt;/span
                  &gt;&lt;span
                    class="mspace"
                    style="margin-right: 0.16666666666666666em"
                  &gt;&lt;/span
                  &gt;&lt;span class="mord"
                    &gt;&lt;span
                      class="mord mathdefault"
                      style="margin-right: 0.02778em"
                      &gt;θ&lt;/span
                    &gt;&lt;span class="msupsub"
                      &gt;&lt;span class="vlist-t vlist-t2"
                        &gt;&lt;span class="vlist-r"
                          &gt;&lt;span
                            class="vlist"
                            style="height: 0.2805559999999999em"
                            &gt;&lt;span
                              style="
                                top: -2.5500000000000003em;
                                margin-left: -0.02778em;
                                margin-right: 0.05em;
                              "
                              &gt;&lt;span class="pstrut" style="height: 2.7em"&gt;&lt;/span
                              &gt;&lt;span class="sizing reset-size6 size3 mtight"
                                &gt;&lt;span class="mord mathdefault mtight"
                                  &gt;t&lt;/span
                                &gt;&lt;/span
                              &gt;&lt;/span
                            &gt;&lt;/span
                          &gt;&lt;span class="vlist-s"&gt;​&lt;/span&gt;&lt;/span
                        &gt;&lt;span class="vlist-r"
                          &gt;&lt;span class="vlist" style="height: 0.15em"
                            &gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span
                  &gt;&lt;span class="mclose"&gt;)&lt;/span
                  &gt;&lt;span class="mclose"&gt;⟩&lt;/span&gt;&lt;/span
                &gt;&lt;/span
              &gt;&lt;/span
            &gt;&lt;/span
          &gt;
        &lt;/p&gt;

        &lt;p&gt;
          We can see that selecting $x$ to maximize $\langle \nabla \ell
          (z;\theta_t), \nabla \ell (x;\theta_t)\rangle$ will maximally reduce
          the validation loss on $z$. The method was initially proposed and
          employed in TracIn (&lt;a href="https://arxiv.org/abs/2002.08484"
            &gt;Pruthi et al., 2020&lt;/a
          &gt;) to gain insights into how training examples influence the model’s
          predictions. The formulation is also very suitable for transfer
          learning, because there is no need to assume any relationship between
          $x$ and $z$. But, there are a couple of modifications required to make
          it work for our setting (see
          &lt;a href="https://arxiv.org/abs/2402.04333" target="_blank"
            &gt;our paper&lt;/a
          &gt;
          for details):
        &lt;/p&gt;

        &lt;ol&gt;
          &lt;li&gt;
            &lt;strong&gt;Adam:&lt;/strong&gt; LLMs are generally tuned using Adam, which
            has a more complicated update formula involving the moving averages
            of the gradient moments. We can plug in that update, which we denote
            as $\Gamma(x;\theta_t)$ instead of $\nabla\ell(x;\theta_t)$. See the
            paper for details on the Adam formulation and the resulting
            technical complications.
          &lt;/li&gt;
          &lt;li&gt;
            &lt;strong&gt;Multi-Epoch:&lt;/strong&gt; We usually train our models over
            several epochs, which means each candidate $x\in\mathcal{D}$ would
            be seen several times over the course of training. We would want our
            estimate of the influence of $x$ to take this into account, but we
            want to avoid the computational cost of training on the entire
            candidate dataset for several epochs. Instead, we approximate how
            the model would adapt to seeing the same data by performing a warmup
            training period on a randomly selected 5% of the data and
            aggregating the influence over the whole run. See the paper for how
            we aggregate influences over multiple epochs.
          &lt;/li&gt;
          &lt;li&gt;
            &lt;strong&gt;Variable-Length Instruction Data&lt;/strong&gt;: Instruction
            tuning sequences have differing lengths. Experiments in our paper
            showed that shorter sequences exhibit much larger norms, so the
            inner product $\langle \nabla \ell (z;\theta_t), \nabla \ell
            (x;\theta_t)\rangle$ is much larger for shorter $x$. We therefore
            decide to instead use the cosine similarity instead of the inner
            product when measuring the influence. This is not a failure of the
            formulation described above (which indeed is quite simple); instead,
            it indicates that we ought to perform data selection for
            &lt;em&gt;individual tokens&lt;/em&gt; instead of &lt;em&gt;entire sequences&lt;/em&gt;.
            However, measuring the gradient of every token in the sequence is
            prohibitively expensive with today’s methods, so we stick to
            sequence selection and leave a token-level formulation to future
            work.
          &lt;/li&gt;
        &lt;/ol&gt;

        &lt;p&gt;
          Altogether, once we make the necessary modifications, we arrive at the
          following formula for computing the influence of a training datapoint
          $x$ on a validation point $z$.
        &lt;/p&gt;

        &lt;p&gt;
          &lt;span class="katex-display"
            &gt;&lt;span class="katex"
              &gt;&lt;span class="katex-mathml"
                &gt;&lt;math xmlns="http://www.w3.org/1998/Math/MathML"
                  &gt;&lt;semantics
                    &gt;&lt;mrow
                      &gt;&lt;mtext&gt;Influence&lt;/mtext&gt;&lt;mo stretchy="false"&gt;(&lt;/mo
                      &gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mo separator="true"&gt;,&lt;/mo&gt;&lt;mi&gt;z&lt;/mi
                      &gt;&lt;mo stretchy="false"&gt;)&lt;/mo&gt;&lt;mo&gt;=&lt;/mo
                      &gt;&lt;munderover
                        &gt;&lt;mo&gt;∑&lt;/mo&gt;&lt;mrow&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/mrow
                        &gt;&lt;mi&gt;N&lt;/mi&gt;&lt;/munderover
                      &gt;&lt;msub
                        &gt;&lt;mover accent="true"&gt;&lt;mi&gt;η&lt;/mi&gt;&lt;mo&gt;ˉ&lt;/mo&gt;&lt;/mover
                        &gt;&lt;mi&gt;i&lt;/mi&gt;&lt;/msub
                      &gt;&lt;mi&gt;cos&lt;/mi&gt;&lt;mo&gt;⁡&lt;/mo&gt;&lt;mo stretchy="false"&gt;(&lt;/mo
                      &gt;&lt;mi mathvariant="normal"&gt;∇&lt;/mi
                      &gt;&lt;mi mathvariant="normal"&gt;ℓ&lt;/mi&gt;&lt;mo stretchy="false"&gt;(&lt;/mo
                      &gt;&lt;mi&gt;z&lt;/mi&gt;&lt;mo separator="true"&gt;;&lt;/mo
                      &gt;&lt;msub&gt;&lt;mi&gt;θ&lt;/mi&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;/msub
                      &gt;&lt;mo stretchy="false"&gt;)&lt;/mo&gt;&lt;mo separator="true"&gt;,&lt;/mo
                      &gt;&lt;mi mathvariant="normal"&gt;Γ&lt;/mi&gt;&lt;mo stretchy="false"&gt;(&lt;/mo
                      &gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mo separator="true"&gt;;&lt;/mo
                      &gt;&lt;msub&gt;&lt;mi&gt;θ&lt;/mi&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;/msub
                      &gt;&lt;mo stretchy="false"&gt;)&lt;/mo
                      &gt;&lt;mo stretchy="false"&gt;)&lt;/mo&gt;&lt;/mrow
                    &gt;&lt;annotation encoding="application/x-tex"&gt;
                      \textrm{Influence}(x, z) = \sum_{i=1}^N \bar{\eta}_i \cos
                      (\nabla \ell(z;\theta_i), \Gamma(x; \theta_i))
                    &lt;/annotation&gt;&lt;/semantics
                  &gt;&lt;/math
                &gt;&lt;/span
              &gt;&lt;span class="katex-html" aria-hidden="true"
                &gt;&lt;span class="base"
                  &gt;&lt;span
                    class="strut"
                    style="height: 1em; vertical-align: -0.25em"
                  &gt;&lt;/span
                  &gt;&lt;span class="mord text"
                    &gt;&lt;span class="mord textrm"&gt;Influence&lt;/span&gt;&lt;/span
                  &gt;&lt;span class="mopen"&gt;(&lt;/span
                  &gt;&lt;span class="mord mathdefault"&gt;x&lt;/span
                  &gt;&lt;span class="mpunct"&gt;,&lt;/span
                  &gt;&lt;span
                    class="mspace"
                    style="margin-right: 0.16666666666666666em"
                  &gt;&lt;/span
                  &gt;&lt;span
                    class="mord mathdefault"
                    style="margin-right: 0.04398em"
                    &gt;z&lt;/span
                  &gt;&lt;span class="mclose"&gt;)&lt;/span
                  &gt;&lt;span
                    class="mspace"
                    style="margin-right: 0.2777777777777778em"
                  &gt;&lt;/span
                  &gt;&lt;span class="mrel"&gt;=&lt;/span
                  &gt;&lt;span
                    class="mspace"
                    style="margin-right: 0.2777777777777778em"
                  &gt;&lt;/span&gt;&lt;/span
                &gt;&lt;span class="base"
                  &gt;&lt;span
                    class="strut"
                    style="height: 3.106005em; vertical-align: -1.277669em"
                  &gt;&lt;/span
                  &gt;&lt;span class="mop op-limits"
                    &gt;&lt;span class="vlist-t vlist-t2"
                      &gt;&lt;span class="vlist-r"
                        &gt;&lt;span
                          class="vlist"
                          style="height: 1.8283360000000002em"
                          &gt;&lt;span style="top: -1.872331em; margin-left: 0em"
                            &gt;&lt;span class="pstrut" style="height: 3.05em"&gt;&lt;/span
                            &gt;&lt;span class="sizing reset-size6 size3 mtight"
                              &gt;&lt;span class="mord mtight"
                                &gt;&lt;span class="mord mathdefault mtight"&gt;i&lt;/span
                                &gt;&lt;span class="mrel mtight"&gt;=&lt;/span
                                &gt;&lt;span class="mord mtight"&gt;1&lt;/span&gt;&lt;/span
                              &gt;&lt;/span
                            &gt;&lt;/span
                          &gt;&lt;span style="top: -3.050005em"
                            &gt;&lt;span class="pstrut" style="height: 3.05em"&gt;&lt;/span
                            &gt;&lt;span
                              &gt;&lt;span class="mop op-symbol large-op"
                                &gt;∑&lt;/span
                              &gt;&lt;/span
                            &gt;&lt;/span
                          &gt;&lt;span
                            style="top: -4.3000050000000005em; margin-left: 0em"
                            &gt;&lt;span class="pstrut" style="height: 3.05em"&gt;&lt;/span
                            &gt;&lt;span class="sizing reset-size6 size3 mtight"
                              &gt;&lt;span
                                class="mord mathdefault mtight"
                                style="margin-right: 0.10903em"
                                &gt;N&lt;/span
                              &gt;&lt;/span
                            &gt;&lt;/span
                          &gt;&lt;/span
                        &gt;&lt;span class="vlist-s"&gt;​&lt;/span&gt;&lt;/span
                      &gt;&lt;span class="vlist-r"
                        &gt;&lt;span class="vlist" style="height: 1.277669em"
                          &gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span
                  &gt;&lt;span
                    class="mspace"
                    style="margin-right: 0.16666666666666666em"
                  &gt;&lt;/span
                  &gt;&lt;span class="mord"
                    &gt;&lt;span class="mord accent"
                      &gt;&lt;span class="vlist-t vlist-t2"
                        &gt;&lt;span class="vlist-r"
                          &gt;&lt;span class="vlist" style="height: 0.56778em"
                            &gt;&lt;span style="top: -3em"
                              &gt;&lt;span class="pstrut" style="height: 3em"&gt;&lt;/span
                              &gt;&lt;span class="mord"
                                &gt;&lt;span
                                  class="mord mathdefault"
                                  style="margin-right: 0.03588em"
                                  &gt;η&lt;/span
                                &gt;&lt;/span
                              &gt;&lt;/span
                            &gt;&lt;span style="top: -3em"
                              &gt;&lt;span class="pstrut" style="height: 3em"&gt;&lt;/span
                              &gt;&lt;span
                                class="accent-body"
                                style="left: -0.19444em"
                                &gt;&lt;span class="mord"&gt;ˉ&lt;/span&gt;&lt;/span
                              &gt;&lt;/span
                            &gt;&lt;/span
                          &gt;&lt;span class="vlist-s"&gt;​&lt;/span&gt;&lt;/span
                        &gt;&lt;span class="vlist-r"
                          &gt;&lt;span class="vlist" style="height: 0.19444em"
                            &gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span
                    &gt;&lt;span class="msupsub"
                      &gt;&lt;span class="vlist-t vlist-t2"
                        &gt;&lt;span class="vlist-r"
                          &gt;&lt;span
                            class="vlist"
                            style="height: 0.31166399999999994em"
                            &gt;&lt;span
                              style="
                                top: -2.5500000000000003em;
                                margin-right: 0.05em;
                              "
                              &gt;&lt;span class="pstrut" style="height: 2.7em"&gt;&lt;/span
                              &gt;&lt;span class="sizing reset-size6 size3 mtight"
                                &gt;&lt;span class="mord mathdefault mtight"
                                  &gt;i&lt;/span
                                &gt;&lt;/span
                              &gt;&lt;/span
                            &gt;&lt;/span
                          &gt;&lt;span class="vlist-s"&gt;​&lt;/span&gt;&lt;/span
                        &gt;&lt;span class="vlist-r"
                          &gt;&lt;span class="vlist" style="height: 0.15em"
                            &gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span
                  &gt;&lt;span
                    class="mspace"
                    style="margin-right: 0.16666666666666666em"
                  &gt;&lt;/span
                  &gt;&lt;span class="mop"&gt;cos&lt;/span&gt;&lt;span class="mopen"&gt;(&lt;/span
                  &gt;&lt;span class="mord"&gt;∇&lt;/span&gt;&lt;span class="mord"&gt;ℓ&lt;/span
                  &gt;&lt;span class="mopen"&gt;(&lt;/span
                  &gt;&lt;span
                    class="mord mathdefault"
                    style="margin-right: 0.04398em"
                    &gt;z&lt;/span
                  &gt;&lt;span class="mpunct"&gt;;&lt;/span
                  &gt;&lt;span
                    class="mspace"
                    style="margin-right: 0.16666666666666666em"
                  &gt;&lt;/span
                  &gt;&lt;span class="mord"
                    &gt;&lt;span
                      class="mord mathdefault"
                      style="margin-right: 0.02778em"
                      &gt;θ&lt;/span
                    &gt;&lt;span class="msupsub"
                      &gt;&lt;span class="vlist-t vlist-t2"
                        &gt;&lt;span class="vlist-r"
                          &gt;&lt;span
                            class="vlist"
                            style="height: 0.31166399999999994em"
                            &gt;&lt;span
                              style="
                                top: -2.5500000000000003em;
                                margin-left: -0.02778em;
                                margin-right: 0.05em;
                              "
                              &gt;&lt;span class="pstrut" style="height: 2.7em"&gt;&lt;/span
                              &gt;&lt;span class="sizing reset-size6 size3 mtight"
                                &gt;&lt;span class="mord mathdefault mtight"
                                  &gt;i&lt;/span
                                &gt;&lt;/span
                              &gt;&lt;/span
                            &gt;&lt;/span
                          &gt;&lt;span class="vlist-s"&gt;​&lt;/span&gt;&lt;/span
                        &gt;&lt;span class="vlist-r"
                          &gt;&lt;span class="vlist" style="height: 0.15em"
                            &gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span
                  &gt;&lt;span class="mclose"&gt;)&lt;/span&gt;&lt;span class="mpunct"&gt;,&lt;/span
                  &gt;&lt;span
                    class="mspace"
                    style="margin-right: 0.16666666666666666em"
                  &gt;&lt;/span
                  &gt;&lt;span class="mord"&gt;Γ&lt;/span&gt;&lt;span class="mopen"&gt;(&lt;/span
                  &gt;&lt;span class="mord mathdefault"&gt;x&lt;/span
                  &gt;&lt;span class="mpunct"&gt;;&lt;/span
                  &gt;&lt;span
                    class="mspace"
                    style="margin-right: 0.16666666666666666em"
                  &gt;&lt;/span
                  &gt;&lt;span class="mord"
                    &gt;&lt;span
                      class="mord mathdefault"
                      style="margin-right: 0.02778em"
                      &gt;θ&lt;/span
                    &gt;&lt;span class="msupsub"
                      &gt;&lt;span class="vlist-t vlist-t2"
                        &gt;&lt;span class="vlist-r"
                          &gt;&lt;span
                            class="vlist"
                            style="height: 0.31166399999999994em"
                            &gt;&lt;span
                              style="
                                top: -2.5500000000000003em;
                                margin-left: -0.02778em;
                                margin-right: 0.05em;
                              "
                              &gt;&lt;span class="pstrut" style="height: 2.7em"&gt;&lt;/span
                              &gt;&lt;span class="sizing reset-size6 size3 mtight"
                                &gt;&lt;span class="mord mathdefault mtight"
                                  &gt;i&lt;/span
                                &gt;&lt;/span
                              &gt;&lt;/span
                            &gt;&lt;/span
                          &gt;&lt;span class="vlist-s"&gt;​&lt;/span&gt;&lt;/span
                        &gt;&lt;span class="vlist-r"
                          &gt;&lt;span class="vlist" style="height: 0.15em"
                            &gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span
                  &gt;&lt;span class="mclose"&gt;)&lt;/span
                  &gt;&lt;span class="mclose"&gt;)&lt;/span&gt;&lt;/span
                &gt;&lt;/span
              &gt;&lt;/span
            &gt;&lt;/span
          &gt;
        &lt;/p&gt;

        &lt;p&gt;
          where $\Gamma(x; \theta_i)$ is the Adam update mentioned above, $N$ is
          the number of epochs during warmup training (see next section),
          $\bar\eta_i$ is the average learning rate in the $i$th epoch, and
          $\theta_i$ is the model after the $i$th epoch. The above formula makes
          it clear that we need to handle model gradient vectors, which can be
          very large, and we need to aggregate influences over several model
          checkpoints. In the next section, we describe how we compute this
          formula efficiently.
        &lt;/p&gt;

        &lt;h2 id="selecting-less-data"&gt;Selecting LESS Data&lt;/h2&gt;

        &lt;center&gt;
          &lt;img
            src="/assets/dataselection_img/method2-cropped.svg"
            alt="The four major steps of LESS"
            style="margin-left: auto; margin-right: auto; width: 80%"
          /&gt;
          &lt;br /&gt;&lt;br /&gt;&lt;br /&gt;
          &lt;div class="caption"&gt;
            &lt;span class="caption-label"
              &gt;An overview of the four steps of our algorithm,
              &lt;a href="https://arxiv.org/abs/2402.04333" target="_blank"&gt;LESS&lt;/a
              &gt;.
            &lt;/span&gt;
          &lt;/div&gt;
        &lt;/center&gt;

        &lt;p&gt;
          &lt;strong&gt;LESS&lt;/strong&gt; consists of four major steps to make influence
          estimation feasible and scalable:
        &lt;/p&gt;

        &lt;ol&gt;
          &lt;li&gt;
            &lt;strong&gt;Warmup LoRA training:&lt;/strong&gt; We train the model with a
            warmup phase with a random subset of data and checkpoint the model
            $\theta_i$ and optimizer update $\Gamma$ over $N$ epochs. We choose
            to train with LoRA to operate on gradients in a much smaller space
            (i.e., ~100M parameters for a 7B model).
          &lt;/li&gt;
          &lt;li&gt;
            &lt;strong&gt;Compute gradient features:&lt;/strong&gt; We acquire low-rank
            &lt;strong&gt;Adam&lt;/strong&gt; gradients by further projecting the LoRA
            gradients down to a smaller dimension $d$ (i.e., 8192 in our
            experiments).&lt;sup id="fnref:2" role="doc-noteref"
              &gt;&lt;a href="#fn:2" class="footnote" rel="footnote"&gt;2&lt;/a&gt;&lt;/sup
            &gt;
          &lt;/li&gt;
          &lt;li&gt;
            &lt;strong&gt;Select data:&lt;/strong&gt; Given a few instances from target
            tasks, we first acquire their compressed gradients. We then
            calculate the influence $\mathrm{Inf}$ and pick the examples with
            the highest scores.&lt;sup id="fnref:3" role="doc-noteref"
              &gt;&lt;a href="#fn:3" class="footnote" rel="footnote"&gt;3&lt;/a&gt;&lt;/sup
            &gt;
          &lt;/li&gt;
          &lt;li&gt;
            &lt;strong&gt;Train models:&lt;/strong&gt; We train models on the selected data,
            using either full-parameter fine-tuning or efficient fine-tuning
            approaches like LoRA.
          &lt;/li&gt;
        &lt;/ol&gt;

        &lt;p&gt;
          Note that the first and second steps are computed once per candidate
          training set and can be stored as a
          &lt;strong&gt;gradient datastore&lt;/strong&gt;. The datastore can be reused to
          quickly select data for different validation tasks. Moreover, the
          model used to select data in steps 1-3 can be different from the model
          trained in step 4, and we call this setting LESS-T, where “T” stands
          for transfer.
        &lt;/p&gt;

        &lt;h2 id="results"&gt;Results&lt;/h2&gt;

        &lt;p&gt;
          &lt;strong
            &gt;Training on 5% selected data often outperforms training on the full
            dataset&lt;/strong
          &gt;
        &lt;/p&gt;

        &lt;p&gt;
          We construct our dataset pool to be a combination of subsets
          &lt;a href="https://huggingface.co/datasets/SirNeural/flan_v2"&gt;FLAN V2&lt;/a
          &gt;, COT,
          &lt;a href="https://huggingface.co/datasets/OpenAssistant/oasst1"
            &gt;Open Assistant 1&lt;/a
          &gt;, and
          &lt;a
            href="https://www.notion.so/Using-LESS-data-to-tune-models-data-selection-in-the-era-of-LLMs-f85c2d946baa45afa974b5f021200b38?pvs=21"
            &gt;Dolly&lt;/a
          &gt;
          datasets. On Mistral-7B and Llama-2-13B, we find that the selected 5%
          of the data outperforms using the full dataset. Additionally, we find
          that the data selected with Llama-2-7B is also effective for
          instruction-tuning Llama-2-13B and Mistral-7B (LESS-T).
        &lt;/p&gt;

        &lt;center&gt;
          &lt;div class="figure"&gt;
            &lt;img
              src="/assets/dataselection_img/llama_bar.svg"
              alt="Bar chart of results"
              style="margin-left: auto; margin-right: auto; width: 48%"
            /&gt;
            &lt;img
              src="/assets/dataselection_img/mistral_bar.svg"
              alt="Bar chart of results"
              style="margin-left: auto; margin-right: auto; width: 48%"
            /&gt;
            &lt;br /&gt;

            &lt;div class="caption"&gt;
              &lt;span class="caption-label"
                &gt;Training on just 5% of the data, selected by our algorithm
                &lt;a href="https://arxiv.org/abs/2402.04333" target="_blank"
                  &gt;LESS&lt;/a
                &gt;, outperforms training on the full dataset.
              &lt;/span&gt;
            &lt;/div&gt;
          &lt;/div&gt;
        &lt;/center&gt;

        &lt;center&gt;
          &lt;div class="figure"&gt;
            &lt;img
              src="/assets/dataselection_img/table.png"
              alt="Table of results"
              style="margin-left: auto; margin-right: auto; width: 100%"
            /&gt;
            &lt;br /&gt;

            &lt;div class="caption"&gt;
              &lt;span class="caption-label"
                &gt;In-depth results of using
                &lt;a href="https://arxiv.org/abs/2402.04333" target="_blank"
                  &gt;LESS&lt;/a
                &gt;
                on three benchmarks.&lt;/span
              &gt;
              &lt;br /&gt;LESS-T indicates that we used LLaMA-2-7B to select the data
              for training the model.
            &lt;/div&gt;
          &lt;/div&gt;
        &lt;/center&gt;

        &lt;p&gt;&lt;strong&gt;LESS outperforms the baselines&lt;/strong&gt;&lt;/p&gt;

        &lt;p&gt;
          We compare our approach to several baselines with different
          &lt;strong&gt;relevancy&lt;/strong&gt; criteria.
        &lt;/p&gt;

        &lt;ul&gt;
          &lt;li&gt;Random: random selection&lt;/li&gt;
          &lt;li&gt;
            BM25: computes lexical overlap as a relevancy score and selects the
            examples with the highest scores
          &lt;/li&gt;
          &lt;li&gt;
            DSIR (Xie et al., 2023): weight candidate training data with n-gram
            features, and resample data based on the weights
          &lt;/li&gt;
          &lt;li&gt;
            RDS: uses model’s hidden representations as features and selects the
            examples with the highest similarity scores
          &lt;/li&gt;
        &lt;/ul&gt;

        &lt;p&gt;
          We surprisingly find that LESS is the only approach that consistently
          outperforms random selection. Other approaches, unfortunately, either
          provide minimal improvement over random selection (BM25), or
          underperform random selection. In the next section, we provide
          qualitative examples to have an in-depth understanding of why other
          approaches fail.
        &lt;/p&gt;

        &lt;center&gt;
          &lt;div class="figure"&gt;
            &lt;img
              src="/assets/dataselection_img/baselines.jpeg"
              alt="LESS outperforms baselines (bar chart)"
              style="margin-left: auto; margin-right: auto; width: 60%"
            /&gt;
            &lt;br /&gt;

            &lt;div class="caption"&gt;
              &lt;span class="caption-label"&gt;LESS outperforms all baselines.&lt;/span&gt;
            &lt;/div&gt;
          &lt;/div&gt;
        &lt;/center&gt;

        &lt;p&gt;
          &lt;strong
            &gt;LESS selects examples with similar underlying task
            structures&lt;/strong
          &gt;
        &lt;/p&gt;

        &lt;p&gt;
          We provide top selected examples by BM25, RDS and LESS for a TydiQA
          example. The TydiQA example presents a paragraph and a question in
          Bengali, and the goal is to locate an answer to the question within
          the given paragraph. BM25 and RDS select examples in Bengali, but
          these examples are related to a different task. In contrast, LESS
          selects a question-answering example written in English, which is more
          relevant to the target task.
        &lt;/p&gt;

        &lt;p&gt;
          This pattern holds true for the top examples selected for other
          questions. Upon further investigation, we discovered that the top
          examples selected by BM25 and RDS are predominantly in Bengali,
          whereas LESS consistently chooses examples in English that are
          specifically related to question answering. This observation suggests
          that LESS prioritizes examples that share a similar reasoning process,
          while the other approaches place too much emphasis on superficial cues
          such as language or topic, rather than the underlying task structure.
        &lt;/p&gt;

        &lt;center&gt;
          &lt;div class="figure"&gt;
            &lt;img
              src="/assets/dataselection_img/qualitative.png"
              alt="Qualitative analysis of examples chosen by LESS."
              style="margin-left: auto; margin-right: auto; width: 100%"
            /&gt;
            &lt;br /&gt;

            &lt;div class="caption"&gt;
              &lt;span class="caption-label"
                &gt;Qualitative analysis of selected data.&lt;/span
              &gt;
              &lt;br /&gt;
              LESS circumvents surface-form cues to instead select examples with
              similar reasoning types as the validation data.
            &lt;/div&gt;
          &lt;/div&gt;
        &lt;/center&gt;

        &lt;p&gt;
          &lt;strong
            &gt;More computation on data selection enhances performance&lt;/strong
          &gt;
        &lt;/p&gt;

        &lt;p&gt;
          Our ablations in the
          &lt;a href="https://arxiv.org/abs/2402.04333" target="_blank"&gt;paper&lt;/a&gt;
          show that spending more computation in any of the steps of LESS can
          improve the performance of the method at the cost of additional
          runtime. For example, using a longer warmup phase in step 1,
          increasing the projected dimension in step 2, and aggregating the
          influence estimate over more model checkpoints can all improve
          performance at the cost of runtime and/or memory. We report results in
          a setting where the data selection cost is reasonable: selecting and
          training on the data requires less time than training on all available
          data. However, our results show that training on LESS data can even
          &lt;strong&gt;improve&lt;/strong&gt; performance, so it may be worthwhile to
          invest more resources into the selection stage.
        &lt;/p&gt;

        &lt;h1 id="prior-work"&gt;Past Works on Data Selection&lt;/h1&gt;

        &lt;p&gt;
          In this section, we discuss many related works in-domain data
          selection and transfer data selection.&lt;sup
            id="fnref:4"
            role="doc-noteref"
            &gt;&lt;a href="#fn:4" class="footnote" rel="footnote"&gt;4&lt;/a&gt;&lt;/sup
          &gt;
          Our goal is to briefly cover the many broad intuitions that can inform
          different data selection algorithms.
        &lt;/p&gt;

        &lt;h2 id="in-domain-data-selection"&gt;In-Domain Data Selection&lt;/h2&gt;

        &lt;p&gt;
          As we mentioned before, the goal of data selection depends heavily on
          the setting. When ample labeled in-domain data is available (e.g.,
          image classification in vision is an obvious case, pre-training data
          selection is also one), the goal is to identify the most
          &lt;strong&gt;representative&lt;/strong&gt; subset of data, so much of the past
          work has sought to define the notion of “representative” datapoints.
          &lt;a href="https://arxiv.org/abs/1708.00489"
            &gt;Sener &amp;amp; Severese 2018&lt;/a
          &gt;
          explicitly defines this problem to be a core-set problem (&lt;a
            href="https://arxiv.org/abs/1703.06476"
            &gt;Bachem et al., 2017&lt;/a
          &gt;,
          &lt;a href="https://www.jmlr.org/papers/v20/18-167.html"
            &gt;Tremblay et al., 2019&lt;/a
          &gt;), which aims to choose a set of datapoints such that each data point
          is close to at least one selected data point. Subsequent research has
          linked specific attributes of datapoints, such as being “easy to
          forget” (&lt;a href="https://arxiv.org/abs/1812.05159"
            &gt;Toneva et al., 2019&lt;/a
          &gt;), exhibiting “large prediction error” (&lt;a
            href="https://arxiv.org/abs/2107.07075"
            &gt;Paul et al., 2021&lt;/a
          &gt;), being “hard to memorize” (&lt;a
            href="https://arxiv.org/pdf/2008.03703.pdf"
            &gt;Feldman et al., 2020&lt;/a
          &gt;), being “redundant in clustering” (&lt;a
            href="https://arxiv.org/abs/1901.11409"
            &gt;Birodkar et al., 2019&lt;/a
          &gt;) and having “uncertain predictions” (&lt;a
            href="https://arxiv.org/abs/1905.12737"
            &gt;Chitta et al., 2019&lt;/a
          &gt;) with their representativeness and importance.
          &lt;a href="https://arxiv.org/abs/2206.14486"&gt;Sorscher et al., 2022&lt;/a&gt;
          consolidate these attributes under the category of hard examples,
          proposing that these instances are the most critical for effective
          training. One can also use gradient-based features (&lt;a
            href="https://arxiv.org/abs/1906.01827"
            &gt;Mirzasoleiman et al., 2020&lt;/a
          &gt;,
          &lt;a href="https://proceedings.mlr.press/v119/wang20p.html"
            &gt;Wang et al., 2020&lt;/a
          &gt;,
          &lt;a href="https://arxiv.org/abs/2103.00123"&gt;Killamsetty et al., 2021&lt;/a
          &gt;) for data selection. This naturally lends itself to a more general
          meta-learning formulation, which we discuss more below.
        &lt;/p&gt;

        &lt;h3 id="pre-training"&gt;Pre-Training&lt;/h3&gt;

        &lt;p&gt;
          In recent years, the pre-training and fine-tuning paradigm has proven
          to be an effective way to build large-scale foundation models. Unlike
          models trained with curated task-specific datasets like CIFAR or
          ImageNet, these models are trained with massive Internet-scraped data
          consisting of trillions of tokens or images. These two settings differ
          in two ways: (1) quality: web-scale datasets are generally
          heterogeneous in quality; and (2) scale: data selection has strong
          implications on efficiency in the foundational model era. We consider
          general pre-training data selection to be a form of in-domain data
          selection, as it aims to choose data that covers all potential use
          cases.
        &lt;/p&gt;

        &lt;p&gt;
          &lt;strong&gt;Algorithmic Filtering:&lt;/strong&gt; Early in 2019, researchers
          have found that excluding high-perplexity examples from CommonCrawl
          could significantly boost training efficiency, since such examples are
          often nonsensical or of low quality (&lt;a
            href="https://arxiv.org/abs/1911.00359"
            &gt;Wenzek et al., 2019&lt;/a
          &gt;,
          &lt;a href="https://arxiv.org/abs/2309.04564"&gt;Marion et al., 2023&lt;/a&gt;).
          More sophisticated processing identified high-quality data by
          filtering domains and URLs, ensuring diversity, and removing
          redundancy (e.g., via
          &lt;a
            href="https://www.notion.so/Using-LESS-data-to-tune-models-data-selection-in-the-era-of-LLMs-f85c2d946baa45afa974b5f021200b38?pvs=21"
            &gt;MinHashLSH&lt;/a
          &gt;
          or semantic deduplication (&lt;a href="https://arxiv.org/abs/2303.09540"
            &gt;Abbas et al., 2023&lt;/a
          &gt;)). Many popular datasets were constructed in this way, including C4
          (&lt;a href="https://arxiv.org/abs/1910.10683"&gt;Raffel et al., 2021&lt;/a&gt;),
          RefinedWeb (&lt;a href="https://arxiv.org/abs/2306.01116"
            &gt;Penedo et al., 2023&lt;/a
          &gt;), SlimPajama (&lt;a href="https://arxiv.org/abs/2309.10818"
            &gt;Shen et al., 2023&lt;/a
          &gt;), and more. Similar efforts have been explored for pre-training VIT
          models (&lt;a href="https://arxiv.org/abs/2401.04578"
            &gt;Abbas et al., 2024&lt;/a
          &gt;).
        &lt;/p&gt;

        &lt;p&gt;
          &lt;strong&gt;LLM-Aided Filtering&lt;/strong&gt;: An extreme form of algorithmic
          filtering is to explicitly prompt LLMs to generate data satisfying
          certain properties. The approach’s efficacy is clearest in the
          Phi-series models (&lt;a href="https://arxiv.org/abs/2306.11644"
            &gt;Gunasekar et al., 2023&lt;/a
          &gt;,
          &lt;a
            href="https://www.microsoft.com/en-us/research/blog/phi-2-the-surprising-power-of-small-language-models/"
            &gt;Javaheripi et al., 2024&lt;/a
          &gt;), which achieved strong performance on math and coding benchmarks
          with only 1.3B model parameters. While the Phi-series models focus on
          generating textbook-style data, other recent work has shown that
          rephrasing pre-training data to mimic the stylistic and informational
          density of Wikipedia articles markedly enhances the data’s
          cost-effectiveness (&lt;a href="https://arxiv.org/abs/2401.16380"
            &gt;Maini et al., 2024&lt;/a
          &gt;). Aside from generating data, LLMs can also be used to judge the
          quality of data. Recent works such as QuRating (&lt;a
            href="https://arxiv.org/abs/2402.09739"
            &gt;Wettig et al., 2024&lt;/a
          &gt;) and Ask-LLM (&lt;a href="https://arxiv.org/abs/2402.09668"
            &gt;Sachdeva et al., 2024&lt;/a
          &gt;) aim to exploit capable language models to directly provide quality
          scores to data instances, offering a metric for evaluating their
          potential impact on model training.
        &lt;/p&gt;

        &lt;p&gt;
          &lt;strong&gt;Meta-Learning Formulation&lt;/strong&gt;: Instead of relying on
          human notions of quality, one can also phrase in-domain and transfer
          data selection as meta-learning problems (&lt;a
            href="https://arxiv.org/abs/2011.00050"
            &gt;Nguyen et al., 2020&lt;/a
          &gt;). The outer loop selects data for models in the inner loop to train
          on. Meta-learning approaches have traditionally been very
          computationally expensive. Recent work, dubbed datamodels (&lt;a
            href="https://arxiv.org/abs/2202.00622"
            &gt;Ilyas et al., 2022&lt;/a
          &gt;), seeks to make this bi-level optimization problem more tractable by
          directly training a model to predict the test performance that would
          result from excluding or including a particular datapoint. Subsequent
          work (&lt;a href="https://arxiv.org/abs/2303.14186"&gt;Park et al., 2023&lt;/a
          &gt;) made this approach computationally efficient, and recently,
          &lt;a href="https://arxiv.org/abs/2401.12926"&gt;Engstrom et al., 2024&lt;/a&gt;
          used this approach to score pre-training examples based on how they
          affect performance on a target set of examples. Our work, LESS, can be
          interpreted as one such meta-learning formulation for selecting data
          in the instruction tuning setting.
        &lt;/p&gt;

        &lt;h3 id="instruction-tuning"&gt;Instruction Tuning&lt;/h3&gt;

        &lt;p&gt;
          Instruction tuning stands as a pivotal process in unlocking
          capabilities of pre-trained base models by further training the models
          to make them follow human instructions. Many works have assembled
          massive instruction tuning datasets. Some early datasets are human
          annotated (e.g., Open Assistant and Dolly), though recent trends
          mostly use completions from GPT models (e.g., Orca, ShardGPT,
          UltraChat etc.). The queries in these datasets cover a broad spectrum
          of topics, and could be as diverse as pre-training datasets. They are
          mostly used to build general-purpose chatbots.
        &lt;/p&gt;

        &lt;p&gt;
          Recently, a lot of works have shown that high-quality data is
          essential for instruction tuning. The pioneering work LIMA (&lt;a
            href="https://arxiv.org/abs/2305.11206"
            &gt;Zhou et al., 2023&lt;/a
          &gt;) illustrates that a mere 1,000 meticulously selected high-quality
          human-curated examples could lead to marked performance improvements.
          Numerous studies have thus endeavored to automate the data selection
          pipeline, including strategies for choosing examples based on their
          naturalness (&lt;a href="https://arxiv.org/abs/2307.06290"
            &gt;Cao et al., 2023&lt;/a
          &gt;), employing GPT-4 for quality scoring (&lt;a
            href="https://arxiv.org/abs/2307.08701"
            &gt;Chen et al., 2023&lt;/a
          &gt;), enhancing data diversity (&lt;a
            href="https://arxiv.org/abs/2311.14736"
            &gt;Bukharin et al., 2023&lt;/a
          &gt;, &lt;a href="https://arxiv.org/abs/2312.15685"&gt;Liu et al., 2023&lt;/a&gt;),
          and ensuring broad coverage (&lt;a
            href="https://arxiv.org/abs/2311.15653"
            &gt;Du et al., 2023&lt;/a
          &gt;). Additionally, some research has explored the benefits of
          prioritizing longer examples (&lt;a
            href="https://arxiv.org/abs/2402.04833"
            &gt;Zhao et al., 2024&lt;/a
          &gt;), but it remains uncertain whether if this simply aligns on a
          surface level with the tendency of GPT to favor longer outputs (&lt;a
            href="https://arxiv.org/abs/2306.04751"
            &gt;Wang et al., 2023&lt;/a
          &gt;).
        &lt;/p&gt;

        &lt;h2 id="transfer-data-selection"&gt;Transfer Data Selection&lt;/h2&gt;

        &lt;p&gt;
          Transfer data selection and in-domain data selection differ in
          purpose. While in-domain data selection aims to cover the properties
          of the entire dataset, transfer data selection focuses on enhancing
          the performance of a specific subdistribution of data. This shift in
          focus leads to a change in the selection criterion from
          representativeness to &lt;strong&gt;relevance&lt;/strong&gt;.
        &lt;/p&gt;

        &lt;p&gt;
          In transfer data selection, the goal is to identify the most relevant
          data points from a large pool of available data. This approach is
          particularly useful for building domain-specific models or improving
          the performance of specific tasks or queries. To achieve this, a
          subset of target data is typically required to serve as an anchor for
          the data selection process. Previous works in this area include the
          study by
          &lt;a href="https://aclanthology.org/2020.acl-main.740/"
            &gt;Gururangan et al. (2020)&lt;/a
          &gt;, which demonstrates the effectiveness of continued training on
          topic-specific pre-training data to improve performance on
          domain-specific downstream tasks. Another notable work is by
          &lt;a href="https://arxiv.org/abs/2302.03169"&gt;Xie et al. (2023)&lt;/a&gt;,
          which introduces a data reweighting approach based on the n-gram
          similarity between the source data and the target distribution. While
          these two works focus on aligning data with surface form cues (i.e.,
          topic and ngram matching), LESS selects data that matches the
          underlying task or reasoning type.
        &lt;/p&gt;

        &lt;h1 id="conclusion-and-future-directions"&gt;
          Conclusion and Future Directions
        &lt;/h1&gt;

        &lt;p&gt;
          Data plays a crucial role in determining the capabilities of trained
          deep models. In the in-domain setting, data selection aims to select a
          small yet representative dataset. Reducing the dataset size directly
          reduces the time required for training, which can be extremely useful
          when pre-training LLMs. On other hand, in the transfer setting, one
          seeks to solve tasks that do not have much data associated with them,
          and this requires &lt;em&gt;filtering&lt;/em&gt; the dataset to identify a
          &lt;strong&gt;relevant&lt;/strong&gt; subset to train on. Our method, LESS,
          selects data in the transfer setting to perform targeted instruction
          tuning, and it identifies the most relevant 5% of the dataset that can
          induce better performance than training on the full dataset. More
          broadly, we see data as being an important area of study for improving
          LLMs. Spending more compute on data selection (or, more broadly, data
          generation) in many different settings has proven to be very valuable,
          though one has to ensure that the cost of selection does not
          skyrocket. Data selection can also go beyond improving or preserving
          performance to attribute particular model behaviors to the training
          data. For example, a recent follow-up to LESS (&lt;a
            href="https://arxiv.org/abs/2404.01099"
            &gt;He et al., 2024&lt;/a
          &gt;) identifies seemingly benign data that somehow breaks the safety of
          models during fine-tuning. We’re excited to see where data selection
          goes next!
        &lt;/p&gt;

        &lt;p&gt;
          &lt;strong&gt;Acknowledgements&lt;/strong&gt;: LESS is co-authored with Suchin
          Gururangan, Sanjeev Arora, and Danqi Chen. We thank (in alphabetical
          order) Dan Friedman, Tianyu Gao, Lucy He, Austin Wang, Alex Wettig,
          and Howard Yen for their helpful feedback on this post!
        &lt;/p&gt;

        &lt;hr /&gt;
        &lt;p&gt;&lt;strong&gt;Footnotes&lt;/strong&gt;:&lt;/p&gt;

        &lt;div class="footnotes" role="doc-endnotes"&gt;
          &lt;ol&gt;
            &lt;li id="fn:1" role="doc-endnote"&gt;
              &lt;p&gt;
                In some cases, the distribution of the training data may not
                match that of the evaluation data. However, the training data
                could still serve as a good approximation and correlates with
                the performance on the evaluation data. For instance, the
                pre-training loss measured on a held-out dataset, typically
                provides a reliable indication of the model’s performance on
                downstream tasks. Therefore, we still consider pre-training data
                selection as in-domain data selection. &lt;a
                  href="#fnref:1"
                  class="reversefootnote"
                  role="doc-backlink"
                  &gt;&amp;#8617;&lt;/a
                &gt;
              &lt;/p&gt;
            &lt;/li&gt;
            &lt;li id="fn:2" role="doc-endnote"&gt;
              &lt;p&gt;
                We use the efficient random projection implementation used in
                TRAK (&lt;a href="https://arxiv.org/abs/2303.14186"
                  &gt;Park et al., 2023&lt;/a
                &gt;). See their amazing codebase
                &lt;a href="https://github.com/MadryLab/trak"&gt;here&lt;/a&gt;! &lt;a
                  href="#fnref:2"
                  class="reversefootnote"
                  role="doc-backlink"
                  &gt;&amp;#8617;&lt;/a
                &gt;
              &lt;/p&gt;
            &lt;/li&gt;
            &lt;li id="fn:3" role="doc-endnote"&gt;
              &lt;p&gt;
                According to the derivation, the training gradients are
                normalized using the Adam update rule, and the validation
                gradients for the target instances are used as-is. Both are
                compressed via random projection. &lt;a
                  href="#fnref:3"
                  class="reversefootnote"
                  role="doc-backlink"
                  &gt;&amp;#8617;&lt;/a
                &gt;
              &lt;/p&gt;
            &lt;/li&gt;
            &lt;li id="fn:4" role="doc-endnote"&gt;
              &lt;p&gt;
                A
                &lt;a href="https://arxiv.org/abs/2402.16827"
                  &gt;recent survey paper&lt;/a
                &gt;
                provides a more comprehensive and formal treatment of data
                selection methods in the context of language models. &lt;a
                  href="#fnref:4"
                  class="reversefootnote"
                  role="doc-backlink"
                  &gt;&amp;#8617;&lt;/a
                &gt;
              &lt;/p&gt;
            &lt;/li&gt;
          &lt;/ol&gt;
        &lt;/div&gt;
      </content></entry><entry><title>How to Scale Hyperparameters as Batch Size Increases</title><id>https://sadhikamalladi.github.io/blog/2024/01/22/SDEs-ScalingRules/</id><link href="https://sadhikamalladi.github.io/blog/2024/01/22/SDEs-ScalingRules/" rel="alternate" type="text/html" /><published>2024-01-22T00:00:00-08:00</published><updated>2024-01-22T00:00:00-08:00</updated><author><name>Sadhika Malladi</name></author><summary>Understanding Optimization using Stochastic Differential Equations</summary><content type="html">
        &lt;p&gt;
          &lt;strong&gt;TL;DR:&lt;/strong&gt; Stochastic differential equations (SDEs)
          provide rigorous, empirically validated
          &lt;strong&gt;scaling rules&lt;/strong&gt; that prescribe how to adjust
          hyperparameters when scaling training runs (Adam or SGD, language or
          vision) to highly distributed settings without sacrificing
          performance. This post focuses on empirically useful insights, and a
          subsequent post will describe the theoretical toolbox of SDEs in more
          detail.
        &lt;/p&gt;

        &lt;h1 id="scaling-rules-increase-batch-size-without-hurting-performance"&gt;
          Scaling Rules: Increase Batch Size without Hurting Performance
        &lt;/h1&gt;

        &lt;p&gt;
          Even as early as 2011, researchers recognized that scaling training
          runs across GPUs can yield huge efficiency gains (&lt;a
            href="https://arxiv.org/abs/1106.5730"
            &gt;HOGWILD! by Niu et al.&lt;/a
          &gt;). For example, loading a large minibatch using data parallel with 8x
          as many GPUs allows you to finish training nearly 8x as fast (modulo
          communication and latency). However, using a very large batch size
          naively hurts SGD performance and puts a limit on how much you can
          scale (&lt;a href="https://arxiv.org/abs/1609.04836"
            &gt;Keskar et al., 2017&lt;/a
          &gt;). A few papers subsequently suggested that one needs a
          &lt;strong&gt;scaling rule&lt;/strong&gt; to adjust the hyperparameters when
          increasing the batch size. There are different scaling rules for
          different optimizers.
        &lt;/p&gt;

        &lt;hr /&gt;

        &lt;p&gt;&lt;strong&gt;Linear Scaling Rule (for SGD)&lt;/strong&gt;&lt;/p&gt;

        &lt;p&gt;
          When scaling the batch size by $\kappa$, scale the learning rate also
          by $\kappa$.
        &lt;/p&gt;

        &lt;p&gt;&lt;strong&gt;Square Root Scaling Rule (for Adam)&lt;/strong&gt;&lt;/p&gt;

        &lt;p&gt;
          When scaling the batch size by $\kappa$, scale the learning rate by
          $\sqrt{\kappa}$. Also, change the other hyperparameters, setting
          $\beta_1 = 1 - \kappa(1-\beta_1)$, $\beta_2 = 1-\kappa(1-\beta_2)$,
          and $\epsilon = \epsilon / \sqrt{\kappa}$.
        &lt;/p&gt;

        &lt;hr /&gt;

        &lt;p&gt;
          Even so, without rigorous theory, it was unknown what the maximal
          performant batch size was, so scaling runs up was an expensive
          trial-and-error game. In 2014,
          &lt;a href="https://arxiv.org/abs/1404.5997"
            &gt;Krizhevsky heuristically derived&lt;/a
          &gt;
          a &lt;strong&gt;&lt;em&gt;square root scaling rule&lt;/em&gt; for SGD&lt;/strong&gt;, stating
          that the learning rate should be scaled by $\sqrt{\kappa}$ when
          scaling the batch size by $\kappa$ (see the bottom of page 5).
          &lt;strong&gt;This ends up being incorrect!&lt;/strong&gt; Even Krizhevsky notes
          that the
          &lt;strong
            &gt;linear scaling rule yields empirically stronger
            performance.&lt;/strong
          &gt;
          But later work in (&lt;a
            href="https://proceedings.neurips.cc/paper/2017/hash/a5e0ff62be0b08456fc7f1e88812af3d-Abstract.html"
            &gt;Hoffer et al., 2017&lt;/a
          &gt;) agreed with the square root scaling rule.
        &lt;/p&gt;

        &lt;p&gt;
          At the same time, an empirical work trying to train a ResNet-50 on
          ImageNet in 1 hour (i.e., in a highly distributed setting),
          &lt;strong
            &gt;derived the linear scaling rule under the assumption that the
            gradient doesn’t change much during training&lt;/strong
          &gt;
          (&lt;a href="https://arxiv.org/abs/1706.02677"&gt;Goyal et al., 2017&lt;/a&gt;).
          Despite a few other optimization tricks,
          &lt;strong&gt;the test accuracy still degraded at large batch size.&lt;/strong&gt;
          (&lt;a href="https://arxiv.org/abs/2006.15081"&gt;Smith et al., 2020&lt;/a&gt;)
          also derived the linear scaling rule using reasoning analogous to a
          central limit theorem but still noted the
          &lt;strong&gt;batch size needs to be small&lt;/strong&gt; for the rule to hold.
        &lt;/p&gt;

        &lt;p&gt;
          To clear up which scaling rule is correct and when it will break, we
          can turn to SDEs!
          &lt;strong
            &gt;&lt;a href="https://openreview.net/forum?id=goEdyJ_nVQI"
              &gt;Our work in NeurIPS 2021&lt;/a
            &gt;
            used SDEs to design a simple test (requiring just one baseline run!)
            for the largest batch size you can parallelize to via the linear
            scaling rule without sacrificing performance.&lt;/strong
          &gt;
          See the figure below. We also designed an efficient simulation of the
          SDE that provided evidence that
          &lt;strong
            &gt;the SDE is the correct way to model many realistic SGD training
            settings.&lt;/strong
          &gt;
          (The next section explains why SDEs are so useful in analyzing SGD.)
        &lt;/p&gt;

        &lt;center&gt;
          &lt;div class="figure"&gt;
            &lt;img
              src="/assets/sde_img/cifar10_lsr.png"
              alt="Using SDE theory to predict the failure of the linear scaling rule using just one baseline run."
              style="margin-left: auto; margin-right: auto; width: 90%"
            /&gt;
            &lt;br /&gt;

            &lt;div class="caption"&gt;
              &lt;span class="caption-label"
                &gt;Figure from
                &lt;a href="https://openreview.net/forum?id=goEdyJ_nVQI"
                  &gt;our work in NeurIPS 2021&lt;/a
                &gt;.&lt;/span
              &gt;
              Predicting the failure of the linear scaling rule with just one
              baseline run, using insights from SDEs! The red dashed line
              indicates where the necessary condition for the scaling rule
              fails, and the shaded red area covers the batch sizes at which the
              test error has increased by 20% or more from its baseline value.
              These experiments are on CIFAR-10 but there are plenty more
              settings in the paper!
            &lt;/div&gt;
          &lt;/div&gt;
        &lt;/center&gt;

        &lt;p&gt;
          In the meantime, language models started to grow popular. (&lt;a
            href="https://arxiv.org/abs/1904.00962"
            &gt;You et al., 2020&lt;/a
          &gt;) designed a scheme to train a BERT model in 76 minutes with Adam,
          and they empirically discovered that
          &lt;strong
            &gt;scaling the learning rate by $\sqrt{\kappa}$ works well&lt;/strong
          &gt;.
          &lt;a href="https://openreview.net/forum?id=F2mhzjHkQP"
            &gt;Our work in NeurIPS 2022&lt;/a
          &gt;
          approached the question theoretically and established new SDE
          approximations for Adam and RMSprop. These yielded
          &lt;strong
            &gt;a square root scaling rule for Adam, which requires scaling the
            other optimization hyperparameters in addition to the learning
            rate.&lt;/strong
          &gt;
          Experiments showed that
          &lt;strong
            &gt;the square root scaling rule preserved test accuracy, perplexity,
            and even post-fine-tuning performance.&lt;/strong
          &gt;
          Naturally, these scaling rules will break at some large batch size,
          but we couldn’t find any empirical setting at the time where this
          happened.
        &lt;/p&gt;

        &lt;p&gt;
          Various works have since used the SDE approximation to derive useful
          distributed training protocols when using EMA (&lt;a
            href="https://openreview.net/forum?id=DkeeXVdQyu"
            &gt;Busbridge et al., 2023&lt;/a
          &gt;) and Local SGD (&lt;a href="https://arxiv.org/abs/2310.14423"
            &gt;Gu et al., 2023&lt;/a
          &gt;).
        &lt;/p&gt;

        &lt;h1 id="into-the-weeds-with-sdes"&gt;Into the Weeds with SDEs&lt;/h1&gt;

        &lt;p&gt;
          Now that I’ve described the importance of SDEs, I’ll describe what
          they are exactly and how they can be used. SDEs describe a continuous
          trajectory of the model parameters over the course of training. The
          hypothesis is that this continuous trajectory is a reasonable and
          easily analyzable approximation of the discrete parameter trajectory
          that SGD prescribes. So let’s see how it shakes out.
        &lt;/p&gt;

        &lt;p&gt;
          At each step, SGD samples a minibatch $B_t\subset \mathcal{D}$ and
          updates the parameters according to some loss function $\ell$:
          $\theta_{t+1} = \theta_t - \eta\nabla \ell(B_t;\theta_t)$. Everyone
          knows the crucial hyperparameter $\eta$, the learning rate, but people
          often overlook the batch size hyperparameter, $B$, which is usually
          set to be the largest allowable value on the GPU (for pre-training)
          and carefully grid-searched (in fine-tuning). These two
          hyperparameters together control how much noise there is in the
          gradient estimate and how it affects the trajectory. The
          &lt;strong&gt;gradient noise&lt;/strong&gt; is what makes SGD different from GD,
          so we will try to study this property carefully to understand why SGD
          results in better generalization than GD.
        &lt;/p&gt;

        &lt;p&gt;
          One approach to studying SGD is to use a continuous approximation of
          the discrete trajectory. This admits the derivation of results about
          the exploration and convergence of the optimization. One way to get a
          continuous approximation is to take the learning rate to be very small
          (i.e., $\eta\to 0$), which gets us to the famous
          &lt;strong&gt;gradient flow&lt;/strong&gt; (GF) limit: $dX_t =
          -\nabla\ell(\mathcal{D};X_t)dt$. I’ll start to use $X$ to denote the
          continuous trajectory of the parameters and $\theta$ to denote the
          discrete one. GF used to be and still somewhat is the tool of choice
          to analyze how optimization proceeds. But if we look closely, we can
          see that GF is agnostic to the batch size! So, even if I was doing
          full-batch GD, I would still end up with GF as the limit. This means
          that GF can’t tell us much about the benefit of SGD over GD.
        &lt;/p&gt;

        &lt;p&gt;
          OK, so we need to find a continuous process that can actually model
          the &lt;strong&gt;gradient noise&lt;/strong&gt;. This is where SDEs come in handy!
          Here’s the SDE used to approximate SGD (&lt;a
            href="https://jmlr.org/papers/v20/17-526.html"
            &gt;Li et al., 2019&lt;/a
          &gt;):
        &lt;/p&gt;

        &lt;p&gt;
          &lt;span class="katex-display"
            &gt;&lt;span class="katex"
              &gt;&lt;span class="katex-mathml"
                &gt;&lt;math xmlns="http://www.w3.org/1998/Math/MathML"
                  &gt;&lt;semantics
                    &gt;&lt;mrow
                      &gt;&lt;mi&gt;d&lt;/mi&gt;&lt;msub&gt;&lt;mi&gt;X&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/msub
                      &gt;&lt;mo&gt;=&lt;/mo
                      &gt;&lt;munder
                        &gt;&lt;munder
                          &gt;&lt;mrow
                            &gt;&lt;mo&gt;−&lt;/mo&gt;&lt;mi mathvariant="normal"&gt;∇&lt;/mi
                            &gt;&lt;mi mathvariant="normal"&gt;ℓ&lt;/mi
                            &gt;&lt;mo stretchy="false"&gt;(&lt;/mo
                            &gt;&lt;mi mathvariant="script"&gt;D&lt;/mi
                            &gt;&lt;mo separator="true"&gt;;&lt;/mo
                            &gt;&lt;msub&gt;&lt;mi&gt;X&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/msub
                            &gt;&lt;mo stretchy="false"&gt;)&lt;/mo&gt;&lt;mi&gt;d&lt;/mi
                            &gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow
                          &gt;&lt;mo stretchy="true"&gt;⏟&lt;/mo&gt;&lt;/munder
                        &gt;&lt;mtext&gt;Drift: Gradient Flow&lt;/mtext&gt;&lt;/munder
                      &gt;&lt;mo&gt;+&lt;/mo
                      &gt;&lt;munder
                        &gt;&lt;munder
                          &gt;&lt;mrow
                            &gt;&lt;mo stretchy="false"&gt;(&lt;/mo&gt;&lt;mi&gt;η&lt;/mi
                            &gt;&lt;mi mathvariant="normal"&gt;Σ&lt;/mi
                            &gt;&lt;mo stretchy="false"&gt;(&lt;/mo
                            &gt;&lt;msub&gt;&lt;mi&gt;X&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/msub
                            &gt;&lt;mo stretchy="false"&gt;)&lt;/mo
                            &gt;&lt;msup
                              &gt;&lt;mo stretchy="false"&gt;)&lt;/mo
                              &gt;&lt;mrow
                                &gt;&lt;mn&gt;1&lt;/mn&gt;&lt;mi mathvariant="normal"&gt;/&lt;/mi
                                &gt;&lt;mn&gt;2&lt;/mn&gt;&lt;/mrow
                              &gt;&lt;/msup
                            &gt;&lt;mi&gt;d&lt;/mi&gt;&lt;msub&gt;&lt;mi&gt;W&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/msub&gt;&lt;/mrow
                          &gt;&lt;mo stretchy="true"&gt;⏟&lt;/mo&gt;&lt;/munder
                        &gt;&lt;mtext&gt;Diffusion: Brownian Motion&lt;/mtext&gt;&lt;/munder
                      &gt;&lt;/mrow
                    &gt;&lt;annotation encoding="application/x-tex"&gt;
                      dX_t = \underbrace{-\nabla
                      \ell(\mathcal{D};X_t)dt}_{\text{Drift: Gradient Flow}} +
                      \underbrace{(\eta\Sigma(X_t))^{1/2}
                      dW_t}_{\text{Diffusion: Brownian Motion}}
                    &lt;/annotation&gt;&lt;/semantics
                  &gt;&lt;/math
                &gt;&lt;/span
              &gt;&lt;span class="katex-html" aria-hidden="true"
                &gt;&lt;span class="base"
                  &gt;&lt;span
                    class="strut"
                    style="height: 0.84444em; vertical-align: -0.15em"
                  &gt;&lt;/span
                  &gt;&lt;span class="mord mathdefault"&gt;d&lt;/span
                  &gt;&lt;span class="mord"
                    &gt;&lt;span
                      class="mord mathdefault"
                      style="margin-right: 0.07847em"
                      &gt;X&lt;/span
                    &gt;&lt;span class="msupsub"
                      &gt;&lt;span class="vlist-t vlist-t2"
                        &gt;&lt;span class="vlist-r"
                          &gt;&lt;span
                            class="vlist"
                            style="height: 0.2805559999999999em"
                            &gt;&lt;span
                              style="
                                top: -2.5500000000000003em;
                                margin-left: -0.07847em;
                                margin-right: 0.05em;
                              "
                              &gt;&lt;span class="pstrut" style="height: 2.7em"&gt;&lt;/span
                              &gt;&lt;span class="sizing reset-size6 size3 mtight"
                                &gt;&lt;span class="mord mathdefault mtight"
                                  &gt;t&lt;/span
                                &gt;&lt;/span
                              &gt;&lt;/span
                            &gt;&lt;/span
                          &gt;&lt;span class="vlist-s"&gt;​&lt;/span&gt;&lt;/span
                        &gt;&lt;span class="vlist-r"
                          &gt;&lt;span class="vlist" style="height: 0.15em"
                            &gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span
                  &gt;&lt;span
                    class="mspace"
                    style="margin-right: 0.2777777777777778em"
                  &gt;&lt;/span
                  &gt;&lt;span class="mrel"&gt;=&lt;/span
                  &gt;&lt;span
                    class="mspace"
                    style="margin-right: 0.2777777777777778em"
                  &gt;&lt;/span&gt;&lt;/span
                &gt;&lt;span class="base"
                  &gt;&lt;span
                    class="strut"
                    style="height: 2.334108em; vertical-align: -1.584108em"
                  &gt;&lt;/span
                  &gt;&lt;span class="mord munder"
                    &gt;&lt;span class="vlist-t vlist-t2"
                      &gt;&lt;span class="vlist-r"
                        &gt;&lt;span class="vlist" style="height: 0.75em"
                          &gt;&lt;span style="top: -1.415892em"
                            &gt;&lt;span class="pstrut" style="height: 3em"&gt;&lt;/span
                            &gt;&lt;span class="sizing reset-size6 size3 mtight"
                              &gt;&lt;span class="mord mtight"
                                &gt;&lt;span class="mord text mtight"
                                  &gt;&lt;span class="mord mtight"
                                    &gt;Drift: Gradient Flow&lt;/span
                                  &gt;&lt;/span
                                &gt;&lt;/span
                              &gt;&lt;/span
                            &gt;&lt;/span
                          &gt;&lt;span style="top: -3em"
                            &gt;&lt;span class="pstrut" style="height: 3em"&gt;&lt;/span
                            &gt;&lt;span class="mord munder"
                              &gt;&lt;span class="vlist-t vlist-t2"
                                &gt;&lt;span class="vlist-r"
                                  &gt;&lt;span class="vlist" style="height: 0.75em"
                                    &gt;&lt;span
                                      class="svg-align"
                                      style="top: -2.102em"
                                      &gt;&lt;span
                                        class="pstrut"
                                        style="height: 3em"
                                      &gt;&lt;/span
                                      &gt;&lt;span
                                        class="stretchy"
                                        style="
                                          height: 0.548em;
                                          min-width: 1.6em;
                                        "
                                        &gt;&lt;span
                                          class="brace-left"
                                          style="height: 0.548em"
                                          &gt;&lt;svg
                                            width="400em"
                                            height="0.548em"
                                            viewBox="0 0 400000 548"
                                            preserveAspectRatio="xMinYMin slice"
                                          &gt;
                                            &lt;path
                                              d="M0 6l6-6h17c12.688 0 19.313.3 20 1 4 4 7.313 8.3 10 13  35.313 51.3 80.813 93.8 136.5 127.5 55.688 33.7 117.188 55.8 184.5 66.5.688  0 2 .3 4 1 18.688 2.7 76 4.3 172 5h399450v120H429l-6-1c-124.688-8-235-61.7 -331-161C60.687 138.7 32.312 99.3 7 54L0 41V6z"
                                            &gt;&lt;/path&gt;&lt;/svg&gt;&lt;/span
                                        &gt;&lt;span
                                          class="brace-center"
                                          style="height: 0.548em"
                                          &gt;&lt;svg
                                            width="400em"
                                            height="0.548em"
                                            viewBox="0 0 400000 548"
                                            preserveAspectRatio="xMidYMin slice"
                                          &gt;
                                            &lt;path
                                              d="M199572 214 c100.7 8.3 195.3 44 280 108 55.3 42 101.7 93 139 153l9 14c2.7-4 5.7-8.7 9-14  53.3-86.7 123.7-153 211-199 66.7-36 137.3-56.3 212-62h199568v120H200432c-178.3  11.7-311.7 78.3-403 201-6 8-9.7 12-11 12-.7.7-6.7 1-18 1s-17.3-.3-18-1c-1.3 0 -5-4-11-12-44.7-59.3-101.3-106.3-170-141s-145.3-54.3-229-60H0V214z"
                                            &gt;&lt;/path&gt;&lt;/svg&gt;&lt;/span
                                        &gt;&lt;span
                                          class="brace-right"
                                          style="height: 0.548em"
                                          &gt;&lt;svg
                                            width="400em"
                                            height="0.548em"
                                            viewBox="0 0 400000 548"
                                            preserveAspectRatio="xMaxYMin slice"
                                          &gt;
                                            &lt;path
                                              d="M399994 0l6 6v35l-6 11c-56 104-135.3 181.3-238 232-57.3  28.7-117 45-179 50H-300V214h399897c43.3-7 81-15 113-26 100.7-33 179.7-91 237 -174 2.7-5 6-9 10-13 .7-1 7.3-1 20-1h17z"
                                            &gt;&lt;/path&gt;&lt;/svg&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span
                                    &gt;&lt;span style="top: -3em"
                                      &gt;&lt;span
                                        class="pstrut"
                                        style="height: 3em"
                                      &gt;&lt;/span
                                      &gt;&lt;span class="mord"
                                        &gt;&lt;span class="mord"&gt;−&lt;/span
                                        &gt;&lt;span class="mord"&gt;∇&lt;/span
                                        &gt;&lt;span class="mord"&gt;ℓ&lt;/span
                                        &gt;&lt;span class="mopen"&gt;(&lt;/span
                                        &gt;&lt;span class="mord"
                                          &gt;&lt;span
                                            class="mord mathcal"
                                            style="margin-right: 0.02778em"
                                            &gt;D&lt;/span
                                          &gt;&lt;/span
                                        &gt;&lt;span class="mpunct"&gt;;&lt;/span
                                        &gt;&lt;span
                                          class="mspace"
                                          style="
                                            margin-right: 0.16666666666666666em;
                                          "
                                        &gt;&lt;/span
                                        &gt;&lt;span class="mord"
                                          &gt;&lt;span
                                            class="mord mathdefault"
                                            style="margin-right: 0.07847em"
                                            &gt;X&lt;/span
                                          &gt;&lt;span class="msupsub"
                                            &gt;&lt;span class="vlist-t vlist-t2"
                                              &gt;&lt;span class="vlist-r"
                                                &gt;&lt;span
                                                  class="vlist"
                                                  style="
                                                    height: 0.2805559999999999em;
                                                  "
                                                  &gt;&lt;span
                                                    style="
                                                      top: -2.5500000000000003em;
                                                      margin-left: -0.07847em;
                                                      margin-right: 0.05em;
                                                    "
                                                    &gt;&lt;span
                                                      class="pstrut"
                                                      style="height: 2.7em"
                                                    &gt;&lt;/span
                                                    &gt;&lt;span
                                                      class="sizing reset-size6 size3 mtight"
                                                      &gt;&lt;span
                                                        class="mord mathdefault mtight"
                                                        &gt;t&lt;/span
                                                      &gt;&lt;/span
                                                    &gt;&lt;/span
                                                  &gt;&lt;/span
                                                &gt;&lt;span class="vlist-s"
                                                  &gt;​&lt;/span
                                                &gt;&lt;/span
                                              &gt;&lt;span class="vlist-r"
                                                &gt;&lt;span
                                                  class="vlist"
                                                  style="height: 0.15em"
                                                  &gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span
                                        &gt;&lt;span class="mclose"&gt;)&lt;/span
                                        &gt;&lt;span class="mord mathdefault"&gt;d&lt;/span
                                        &gt;&lt;span class="mord mathdefault"
                                          &gt;t&lt;/span
                                        &gt;&lt;/span
                                      &gt;&lt;/span
                                    &gt;&lt;/span
                                  &gt;&lt;span class="vlist-s"&gt;​&lt;/span&gt;&lt;/span
                                &gt;&lt;span class="vlist-r"
                                  &gt;&lt;span class="vlist" style="height: 0.898em"
                                    &gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span
                        &gt;&lt;span class="vlist-s"&gt;​&lt;/span&gt;&lt;/span
                      &gt;&lt;span class="vlist-r"
                        &gt;&lt;span class="vlist" style="height: 1.584108em"
                          &gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span
                  &gt;&lt;span
                    class="mspace"
                    style="margin-right: 0.2222222222222222em"
                  &gt;&lt;/span
                  &gt;&lt;span class="mbin"&gt;+&lt;/span
                  &gt;&lt;span
                    class="mspace"
                    style="margin-right: 0.2222222222222222em"
                  &gt;&lt;/span&gt;&lt;/span
                &gt;&lt;span class="base"
                  &gt;&lt;span
                    class="strut"
                    style="height: 2.522108em; vertical-align: -1.584108em"
                  &gt;&lt;/span
                  &gt;&lt;span class="mord munder"
                    &gt;&lt;span class="vlist-t vlist-t2"
                      &gt;&lt;span class="vlist-r"
                        &gt;&lt;span
                          class="vlist"
                          style="height: 0.9379999999999997em"
                          &gt;&lt;span style="top: -1.415892em"
                            &gt;&lt;span class="pstrut" style="height: 3em"&gt;&lt;/span
                            &gt;&lt;span class="sizing reset-size6 size3 mtight"
                              &gt;&lt;span class="mord mtight"
                                &gt;&lt;span class="mord text mtight"
                                  &gt;&lt;span class="mord mtight"
                                    &gt;Diffusion: Brownian Motion&lt;/span
                                  &gt;&lt;/span
                                &gt;&lt;/span
                              &gt;&lt;/span
                            &gt;&lt;/span
                          &gt;&lt;span style="top: -3em"
                            &gt;&lt;span class="pstrut" style="height: 3em"&gt;&lt;/span
                            &gt;&lt;span class="mord munder"
                              &gt;&lt;span class="vlist-t vlist-t2"
                                &gt;&lt;span class="vlist-r"
                                  &gt;&lt;span class="vlist" style="height: 0.938em"
                                    &gt;&lt;span
                                      class="svg-align"
                                      style="top: -2.102em"
                                      &gt;&lt;span
                                        class="pstrut"
                                        style="height: 3em"
                                      &gt;&lt;/span
                                      &gt;&lt;span
                                        class="stretchy"
                                        style="
                                          height: 0.548em;
                                          min-width: 1.6em;
                                        "
                                        &gt;&lt;span
                                          class="brace-left"
                                          style="height: 0.548em"
                                          &gt;&lt;svg
                                            width="400em"
                                            height="0.548em"
                                            viewBox="0 0 400000 548"
                                            preserveAspectRatio="xMinYMin slice"
                                          &gt;
                                            &lt;path
                                              d="M0 6l6-6h17c12.688 0 19.313.3 20 1 4 4 7.313 8.3 10 13  35.313 51.3 80.813 93.8 136.5 127.5 55.688 33.7 117.188 55.8 184.5 66.5.688  0 2 .3 4 1 18.688 2.7 76 4.3 172 5h399450v120H429l-6-1c-124.688-8-235-61.7 -331-161C60.687 138.7 32.312 99.3 7 54L0 41V6z"
                                            &gt;&lt;/path&gt;&lt;/svg&gt;&lt;/span
                                        &gt;&lt;span
                                          class="brace-center"
                                          style="height: 0.548em"
                                          &gt;&lt;svg
                                            width="400em"
                                            height="0.548em"
                                            viewBox="0 0 400000 548"
                                            preserveAspectRatio="xMidYMin slice"
                                          &gt;
                                            &lt;path
                                              d="M199572 214 c100.7 8.3 195.3 44 280 108 55.3 42 101.7 93 139 153l9 14c2.7-4 5.7-8.7 9-14  53.3-86.7 123.7-153 211-199 66.7-36 137.3-56.3 212-62h199568v120H200432c-178.3  11.7-311.7 78.3-403 201-6 8-9.7 12-11 12-.7.7-6.7 1-18 1s-17.3-.3-18-1c-1.3 0 -5-4-11-12-44.7-59.3-101.3-106.3-170-141s-145.3-54.3-229-60H0V214z"
                                            &gt;&lt;/path&gt;&lt;/svg&gt;&lt;/span
                                        &gt;&lt;span
                                          class="brace-right"
                                          style="height: 0.548em"
                                          &gt;&lt;svg
                                            width="400em"
                                            height="0.548em"
                                            viewBox="0 0 400000 548"
                                            preserveAspectRatio="xMaxYMin slice"
                                          &gt;
                                            &lt;path
                                              d="M399994 0l6 6v35l-6 11c-56 104-135.3 181.3-238 232-57.3  28.7-117 45-179 50H-300V214h399897c43.3-7 81-15 113-26 100.7-33 179.7-91 237 -174 2.7-5 6-9 10-13 .7-1 7.3-1 20-1h17z"
                                            &gt;&lt;/path&gt;&lt;/svg&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span
                                    &gt;&lt;span style="top: -3em"
                                      &gt;&lt;span
                                        class="pstrut"
                                        style="height: 3em"
                                      &gt;&lt;/span
                                      &gt;&lt;span class="mord"
                                        &gt;&lt;span class="mopen"&gt;(&lt;/span
                                        &gt;&lt;span
                                          class="mord mathdefault"
                                          style="margin-right: 0.03588em"
                                          &gt;η&lt;/span
                                        &gt;&lt;span class="mord"&gt;Σ&lt;/span
                                        &gt;&lt;span class="mopen"&gt;(&lt;/span
                                        &gt;&lt;span class="mord"
                                          &gt;&lt;span
                                            class="mord mathdefault"
                                            style="margin-right: 0.07847em"
                                            &gt;X&lt;/span
                                          &gt;&lt;span class="msupsub"
                                            &gt;&lt;span class="vlist-t vlist-t2"
                                              &gt;&lt;span class="vlist-r"
                                                &gt;&lt;span
                                                  class="vlist"
                                                  style="
                                                    height: 0.2805559999999999em;
                                                  "
                                                  &gt;&lt;span
                                                    style="
                                                      top: -2.5500000000000003em;
                                                      margin-left: -0.07847em;
                                                      margin-right: 0.05em;
                                                    "
                                                    &gt;&lt;span
                                                      class="pstrut"
                                                      style="height: 2.7em"
                                                    &gt;&lt;/span
                                                    &gt;&lt;span
                                                      class="sizing reset-size6 size3 mtight"
                                                      &gt;&lt;span
                                                        class="mord mathdefault mtight"
                                                        &gt;t&lt;/span
                                                      &gt;&lt;/span
                                                    &gt;&lt;/span
                                                  &gt;&lt;/span
                                                &gt;&lt;span class="vlist-s"
                                                  &gt;​&lt;/span
                                                &gt;&lt;/span
                                              &gt;&lt;span class="vlist-r"
                                                &gt;&lt;span
                                                  class="vlist"
                                                  style="height: 0.15em"
                                                  &gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span
                                        &gt;&lt;span class="mclose"&gt;)&lt;/span
                                        &gt;&lt;span class="mclose"
                                          &gt;&lt;span class="mclose"&gt;)&lt;/span
                                          &gt;&lt;span class="msupsub"
                                            &gt;&lt;span class="vlist-t"
                                              &gt;&lt;span class="vlist-r"
                                                &gt;&lt;span
                                                  class="vlist"
                                                  style="height: 0.938em"
                                                  &gt;&lt;span
                                                    style="
                                                      top: -3.113em;
                                                      margin-right: 0.05em;
                                                    "
                                                    &gt;&lt;span
                                                      class="pstrut"
                                                      style="height: 2.7em"
                                                    &gt;&lt;/span
                                                    &gt;&lt;span
                                                      class="sizing reset-size6 size3 mtight"
                                                      &gt;&lt;span class="mord mtight"
                                                        &gt;&lt;span
                                                          class="mord mtight"
                                                          &gt;1&lt;/span
                                                        &gt;&lt;span
                                                          class="mord mtight"
                                                          &gt;/&lt;/span
                                                        &gt;&lt;span
                                                          class="mord mtight"
                                                          &gt;2&lt;/span
                                                        &gt;&lt;/span
                                                      &gt;&lt;/span
                                                    &gt;&lt;/span
                                                  &gt;&lt;/span
                                                &gt;&lt;/span
                                              &gt;&lt;/span
                                            &gt;&lt;/span
                                          &gt;&lt;/span
                                        &gt;&lt;span class="mord mathdefault"&gt;d&lt;/span
                                        &gt;&lt;span class="mord"
                                          &gt;&lt;span
                                            class="mord mathdefault"
                                            style="margin-right: 0.13889em"
                                            &gt;W&lt;/span
                                          &gt;&lt;span class="msupsub"
                                            &gt;&lt;span class="vlist-t vlist-t2"
                                              &gt;&lt;span class="vlist-r"
                                                &gt;&lt;span
                                                  class="vlist"
                                                  style="
                                                    height: 0.2805559999999999em;
                                                  "
                                                  &gt;&lt;span
                                                    style="
                                                      top: -2.5500000000000003em;
                                                      margin-left: -0.13889em;
                                                      margin-right: 0.05em;
                                                    "
                                                    &gt;&lt;span
                                                      class="pstrut"
                                                      style="height: 2.7em"
                                                    &gt;&lt;/span
                                                    &gt;&lt;span
                                                      class="sizing reset-size6 size3 mtight"
                                                      &gt;&lt;span
                                                        class="mord mathdefault mtight"
                                                        &gt;t&lt;/span
                                                      &gt;&lt;/span
                                                    &gt;&lt;/span
                                                  &gt;&lt;/span
                                                &gt;&lt;span class="vlist-s"
                                                  &gt;​&lt;/span
                                                &gt;&lt;/span
                                              &gt;&lt;span class="vlist-r"
                                                &gt;&lt;span
                                                  class="vlist"
                                                  style="height: 0.15em"
                                                  &gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span
                                  &gt;&lt;span class="vlist-s"&gt;​&lt;/span&gt;&lt;/span
                                &gt;&lt;span class="vlist-r"
                                  &gt;&lt;span class="vlist" style="height: 0.898em"
                                    &gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span
                        &gt;&lt;span class="vlist-s"&gt;​&lt;/span&gt;&lt;/span
                      &gt;&lt;span class="vlist-r"
                        &gt;&lt;span class="vlist" style="height: 1.584108em"
                          &gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span
          &gt;&lt;/span&gt;
        &lt;/p&gt;

        &lt;p&gt;
          The first term of our SDE is the &lt;strong&gt;drift&lt;/strong&gt;, the
          deterministic term in the process, which is just the full-batch
          gradient in this case. The second term is the
          &lt;strong&gt;diffusion&lt;/strong&gt;, which we traditionally model using
          Brownian motion $W_t$. So, the SDE is actually very intuitive: we
          believe that SGD is doing something like GD + noise and we can clearly
          see that intuition in these two terms. Another key feature is that the
          learning rate $\eta$ actually shows up here! As we increase the
          learning rate, the noise in the SDE increases, which also makes sense:
          when we take larger steps with a mini-batch gradient, the trajectory
          increasingly diverges from full-batch GD. The last remaining mystery
          in this equation is $\Sigma(X_t)$, which is the
          &lt;strong&gt;gradient noise covariance&lt;/strong&gt;.
        &lt;/p&gt;

        &lt;p&gt;
          &lt;span class="katex-display"
            &gt;&lt;span class="katex"
              &gt;&lt;span class="katex-mathml"
                &gt;&lt;math xmlns="http://www.w3.org/1998/Math/MathML"
                  &gt;&lt;semantics
                    &gt;&lt;mrow
                      &gt;&lt;msup
                        &gt;&lt;mi mathvariant="normal"&gt;Σ&lt;/mi
                        &gt;&lt;mrow
                          &gt;&lt;mo stretchy="false"&gt;(&lt;/mo&gt;&lt;mi&gt;B&lt;/mi
                          &gt;&lt;mo stretchy="false"&gt;)&lt;/mo&gt;&lt;/mrow
                        &gt;&lt;/msup
                      &gt;&lt;mo stretchy="false"&gt;(&lt;/mo
                      &gt;&lt;msub&gt;&lt;mi&gt;X&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/msub
                      &gt;&lt;mo stretchy="false"&gt;)&lt;/mo&gt;&lt;mo&gt;=&lt;/mo
                      &gt;&lt;msub
                        &gt;&lt;mi mathvariant="double-struck"&gt;E&lt;/mi&gt;&lt;mi&gt;γ&lt;/mi&gt;&lt;/msub
                      &gt;&lt;mo stretchy="false"&gt;[&lt;/mo&gt;&lt;mo stretchy="false"&gt;(&lt;/mo
                      &gt;&lt;mi mathvariant="normal"&gt;∇&lt;/mi
                      &gt;&lt;mi mathvariant="normal"&gt;ℓ&lt;/mi&gt;&lt;mo stretchy="false"&gt;(&lt;/mo
                      &gt;&lt;mi mathvariant="script"&gt;D&lt;/mi&gt;&lt;mo separator="true"&gt;;&lt;/mo
                      &gt;&lt;msub&gt;&lt;mi&gt;X&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/msub
                      &gt;&lt;mo stretchy="false"&gt;)&lt;/mo&gt;&lt;mo&gt;−&lt;/mo
                      &gt;&lt;mi mathvariant="normal"&gt;∇&lt;/mi
                      &gt;&lt;mi mathvariant="normal"&gt;ℓ&lt;/mi&gt;&lt;mo stretchy="false"&gt;(&lt;/mo
                      &gt;&lt;msub&gt;&lt;mi&gt;B&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/msub
                      &gt;&lt;mo separator="true"&gt;;&lt;/mo
                      &gt;&lt;msub&gt;&lt;mi&gt;X&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/msub
                      &gt;&lt;mo stretchy="false"&gt;)&lt;/mo&gt;&lt;mo stretchy="false"&gt;(&lt;/mo
                      &gt;&lt;mi mathvariant="normal"&gt;∇&lt;/mi
                      &gt;&lt;mi mathvariant="normal"&gt;ℓ&lt;/mi&gt;&lt;mo stretchy="false"&gt;(&lt;/mo
                      &gt;&lt;mi mathvariant="script"&gt;D&lt;/mi&gt;&lt;mo separator="true"&gt;;&lt;/mo
                      &gt;&lt;msub&gt;&lt;mi&gt;X&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/msub
                      &gt;&lt;mo stretchy="false"&gt;)&lt;/mo&gt;&lt;mo&gt;−&lt;/mo
                      &gt;&lt;mi mathvariant="normal"&gt;∇&lt;/mi
                      &gt;&lt;mi mathvariant="normal"&gt;ℓ&lt;/mi&gt;&lt;mo stretchy="false"&gt;(&lt;/mo
                      &gt;&lt;msub&gt;&lt;mi&gt;B&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/msub
                      &gt;&lt;mo separator="true"&gt;;&lt;/mo
                      &gt;&lt;msub&gt;&lt;mi&gt;X&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/msub
                      &gt;&lt;mo stretchy="false"&gt;)&lt;/mo
                      &gt;&lt;msup
                        &gt;&lt;mo stretchy="false"&gt;)&lt;/mo
                        &gt;&lt;mi mathvariant="normal"&gt;⊤&lt;/mi&gt;&lt;/msup
                      &gt;&lt;mo stretchy="false"&gt;]&lt;/mo&gt;&lt;/mrow
                    &gt;&lt;annotation encoding="application/x-tex"&gt;
                      \Sigma^{(B)}(X_t) = \mathbb{E}_\gamma [(\nabla
                      \ell(\mathcal{D}; X_t) - \nabla \ell(B_t; X_t)(\nabla
                      \ell(\mathcal{D}; X_t) - \nabla \ell(B_t;X_t))^\top]
                    &lt;/annotation&gt;&lt;/semantics
                  &gt;&lt;/math
                &gt;&lt;/span
              &gt;&lt;span class="katex-html" aria-hidden="true"
                &gt;&lt;span class="base"
                  &gt;&lt;span
                    class="strut"
                    style="height: 1.188em; vertical-align: -0.25em"
                  &gt;&lt;/span
                  &gt;&lt;span class="mord"
                    &gt;&lt;span class="mord"&gt;Σ&lt;/span
                    &gt;&lt;span class="msupsub"
                      &gt;&lt;span class="vlist-t"
                        &gt;&lt;span class="vlist-r"
                          &gt;&lt;span class="vlist" style="height: 0.938em"
                            &gt;&lt;span style="top: -3.113em; margin-right: 0.05em"
                              &gt;&lt;span class="pstrut" style="height: 2.7em"&gt;&lt;/span
                              &gt;&lt;span class="sizing reset-size6 size3 mtight"
                                &gt;&lt;span class="mord mtight"
                                  &gt;&lt;span class="mopen mtight"&gt;(&lt;/span
                                  &gt;&lt;span
                                    class="mord mathdefault mtight"
                                    style="margin-right: 0.05017em"
                                    &gt;B&lt;/span
                                  &gt;&lt;span class="mclose mtight"&gt;)&lt;/span&gt;&lt;/span
                                &gt;&lt;/span
                              &gt;&lt;/span
                            &gt;&lt;/span
                          &gt;&lt;/span
                        &gt;&lt;/span
                      &gt;&lt;/span
                    &gt;&lt;/span
                  &gt;&lt;span class="mopen"&gt;(&lt;/span
                  &gt;&lt;span class="mord"
                    &gt;&lt;span
                      class="mord mathdefault"
                      style="margin-right: 0.07847em"
                      &gt;X&lt;/span
                    &gt;&lt;span class="msupsub"
                      &gt;&lt;span class="vlist-t vlist-t2"
                        &gt;&lt;span class="vlist-r"
                          &gt;&lt;span
                            class="vlist"
                            style="height: 0.2805559999999999em"
                            &gt;&lt;span
                              style="
                                top: -2.5500000000000003em;
                                margin-left: -0.07847em;
                                margin-right: 0.05em;
                              "
                              &gt;&lt;span class="pstrut" style="height: 2.7em"&gt;&lt;/span
                              &gt;&lt;span class="sizing reset-size6 size3 mtight"
                                &gt;&lt;span class="mord mathdefault mtight"
                                  &gt;t&lt;/span
                                &gt;&lt;/span
                              &gt;&lt;/span
                            &gt;&lt;/span
                          &gt;&lt;span class="vlist-s"&gt;​&lt;/span&gt;&lt;/span
                        &gt;&lt;span class="vlist-r"
                          &gt;&lt;span class="vlist" style="height: 0.15em"
                            &gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span
                  &gt;&lt;span class="mclose"&gt;)&lt;/span
                  &gt;&lt;span
                    class="mspace"
                    style="margin-right: 0.2777777777777778em"
                  &gt;&lt;/span
                  &gt;&lt;span class="mrel"&gt;=&lt;/span
                  &gt;&lt;span
                    class="mspace"
                    style="margin-right: 0.2777777777777778em"
                  &gt;&lt;/span&gt;&lt;/span
                &gt;&lt;span class="base"
                  &gt;&lt;span
                    class="strut"
                    style="height: 1.036108em; vertical-align: -0.286108em"
                  &gt;&lt;/span
                  &gt;&lt;span class="mord"
                    &gt;&lt;span class="mord"&gt;&lt;span class="mord mathbb"&gt;E&lt;/span&gt;&lt;/span
                    &gt;&lt;span class="msupsub"
                      &gt;&lt;span class="vlist-t vlist-t2"
                        &gt;&lt;span class="vlist-r"
                          &gt;&lt;span
                            class="vlist"
                            style="height: 0.15139200000000003em"
                            &gt;&lt;span
                              style="
                                top: -2.5500000000000003em;
                                margin-right: 0.05em;
                              "
                              &gt;&lt;span class="pstrut" style="height: 2.7em"&gt;&lt;/span
                              &gt;&lt;span class="sizing reset-size6 size3 mtight"
                                &gt;&lt;span
                                  class="mord mathdefault mtight"
                                  style="margin-right: 0.05556em"
                                  &gt;γ&lt;/span
                                &gt;&lt;/span
                              &gt;&lt;/span
                            &gt;&lt;/span
                          &gt;&lt;span class="vlist-s"&gt;​&lt;/span&gt;&lt;/span
                        &gt;&lt;span class="vlist-r"
                          &gt;&lt;span class="vlist" style="height: 0.286108em"
                            &gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span
                  &gt;&lt;span class="mopen"&gt;[&lt;/span&gt;&lt;span class="mopen"&gt;(&lt;/span
                  &gt;&lt;span class="mord"&gt;∇&lt;/span&gt;&lt;span class="mord"&gt;ℓ&lt;/span
                  &gt;&lt;span class="mopen"&gt;(&lt;/span
                  &gt;&lt;span class="mord"
                    &gt;&lt;span class="mord mathcal" style="margin-right: 0.02778em"
                      &gt;D&lt;/span
                    &gt;&lt;/span
                  &gt;&lt;span class="mpunct"&gt;;&lt;/span
                  &gt;&lt;span
                    class="mspace"
                    style="margin-right: 0.16666666666666666em"
                  &gt;&lt;/span
                  &gt;&lt;span class="mord"
                    &gt;&lt;span
                      class="mord mathdefault"
                      style="margin-right: 0.07847em"
                      &gt;X&lt;/span
                    &gt;&lt;span class="msupsub"
                      &gt;&lt;span class="vlist-t vlist-t2"
                        &gt;&lt;span class="vlist-r"
                          &gt;&lt;span
                            class="vlist"
                            style="height: 0.2805559999999999em"
                            &gt;&lt;span
                              style="
                                top: -2.5500000000000003em;
                                margin-left: -0.07847em;
                                margin-right: 0.05em;
                              "
                              &gt;&lt;span class="pstrut" style="height: 2.7em"&gt;&lt;/span
                              &gt;&lt;span class="sizing reset-size6 size3 mtight"
                                &gt;&lt;span class="mord mathdefault mtight"
                                  &gt;t&lt;/span
                                &gt;&lt;/span
                              &gt;&lt;/span
                            &gt;&lt;/span
                          &gt;&lt;span class="vlist-s"&gt;​&lt;/span&gt;&lt;/span
                        &gt;&lt;span class="vlist-r"
                          &gt;&lt;span class="vlist" style="height: 0.15em"
                            &gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span
                  &gt;&lt;span class="mclose"&gt;)&lt;/span
                  &gt;&lt;span
                    class="mspace"
                    style="margin-right: 0.2222222222222222em"
                  &gt;&lt;/span
                  &gt;&lt;span class="mbin"&gt;−&lt;/span
                  &gt;&lt;span
                    class="mspace"
                    style="margin-right: 0.2222222222222222em"
                  &gt;&lt;/span&gt;&lt;/span
                &gt;&lt;span class="base"
                  &gt;&lt;span
                    class="strut"
                    style="height: 1em; vertical-align: -0.25em"
                  &gt;&lt;/span
                  &gt;&lt;span class="mord"&gt;∇&lt;/span&gt;&lt;span class="mord"&gt;ℓ&lt;/span
                  &gt;&lt;span class="mopen"&gt;(&lt;/span
                  &gt;&lt;span class="mord"
                    &gt;&lt;span
                      class="mord mathdefault"
                      style="margin-right: 0.05017em"
                      &gt;B&lt;/span
                    &gt;&lt;span class="msupsub"
                      &gt;&lt;span class="vlist-t vlist-t2"
                        &gt;&lt;span class="vlist-r"
                          &gt;&lt;span
                            class="vlist"
                            style="height: 0.2805559999999999em"
                            &gt;&lt;span
                              style="
                                top: -2.5500000000000003em;
                                margin-left: -0.05017em;
                                margin-right: 0.05em;
                              "
                              &gt;&lt;span class="pstrut" style="height: 2.7em"&gt;&lt;/span
                              &gt;&lt;span class="sizing reset-size6 size3 mtight"
                                &gt;&lt;span class="mord mathdefault mtight"
                                  &gt;t&lt;/span
                                &gt;&lt;/span
                              &gt;&lt;/span
                            &gt;&lt;/span
                          &gt;&lt;span class="vlist-s"&gt;​&lt;/span&gt;&lt;/span
                        &gt;&lt;span class="vlist-r"
                          &gt;&lt;span class="vlist" style="height: 0.15em"
                            &gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span
                  &gt;&lt;span class="mpunct"&gt;;&lt;/span
                  &gt;&lt;span
                    class="mspace"
                    style="margin-right: 0.16666666666666666em"
                  &gt;&lt;/span
                  &gt;&lt;span class="mord"
                    &gt;&lt;span
                      class="mord mathdefault"
                      style="margin-right: 0.07847em"
                      &gt;X&lt;/span
                    &gt;&lt;span class="msupsub"
                      &gt;&lt;span class="vlist-t vlist-t2"
                        &gt;&lt;span class="vlist-r"
                          &gt;&lt;span
                            class="vlist"
                            style="height: 0.2805559999999999em"
                            &gt;&lt;span
                              style="
                                top: -2.5500000000000003em;
                                margin-left: -0.07847em;
                                margin-right: 0.05em;
                              "
                              &gt;&lt;span class="pstrut" style="height: 2.7em"&gt;&lt;/span
                              &gt;&lt;span class="sizing reset-size6 size3 mtight"
                                &gt;&lt;span class="mord mathdefault mtight"
                                  &gt;t&lt;/span
                                &gt;&lt;/span
                              &gt;&lt;/span
                            &gt;&lt;/span
                          &gt;&lt;span class="vlist-s"&gt;​&lt;/span&gt;&lt;/span
                        &gt;&lt;span class="vlist-r"
                          &gt;&lt;span class="vlist" style="height: 0.15em"
                            &gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span
                  &gt;&lt;span class="mclose"&gt;)&lt;/span&gt;&lt;span class="mopen"&gt;(&lt;/span
                  &gt;&lt;span class="mord"&gt;∇&lt;/span&gt;&lt;span class="mord"&gt;ℓ&lt;/span
                  &gt;&lt;span class="mopen"&gt;(&lt;/span
                  &gt;&lt;span class="mord"
                    &gt;&lt;span class="mord mathcal" style="margin-right: 0.02778em"
                      &gt;D&lt;/span
                    &gt;&lt;/span
                  &gt;&lt;span class="mpunct"&gt;;&lt;/span
                  &gt;&lt;span
                    class="mspace"
                    style="margin-right: 0.16666666666666666em"
                  &gt;&lt;/span
                  &gt;&lt;span class="mord"
                    &gt;&lt;span
                      class="mord mathdefault"
                      style="margin-right: 0.07847em"
                      &gt;X&lt;/span
                    &gt;&lt;span class="msupsub"
                      &gt;&lt;span class="vlist-t vlist-t2"
                        &gt;&lt;span class="vlist-r"
                          &gt;&lt;span
                            class="vlist"
                            style="height: 0.2805559999999999em"
                            &gt;&lt;span
                              style="
                                top: -2.5500000000000003em;
                                margin-left: -0.07847em;
                                margin-right: 0.05em;
                              "
                              &gt;&lt;span class="pstrut" style="height: 2.7em"&gt;&lt;/span
                              &gt;&lt;span class="sizing reset-size6 size3 mtight"
                                &gt;&lt;span class="mord mathdefault mtight"
                                  &gt;t&lt;/span
                                &gt;&lt;/span
                              &gt;&lt;/span
                            &gt;&lt;/span
                          &gt;&lt;span class="vlist-s"&gt;​&lt;/span&gt;&lt;/span
                        &gt;&lt;span class="vlist-r"
                          &gt;&lt;span class="vlist" style="height: 0.15em"
                            &gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span
                  &gt;&lt;span class="mclose"&gt;)&lt;/span
                  &gt;&lt;span
                    class="mspace"
                    style="margin-right: 0.2222222222222222em"
                  &gt;&lt;/span
                  &gt;&lt;span class="mbin"&gt;−&lt;/span
                  &gt;&lt;span
                    class="mspace"
                    style="margin-right: 0.2222222222222222em"
                  &gt;&lt;/span&gt;&lt;/span
                &gt;&lt;span class="base"
                  &gt;&lt;span
                    class="strut"
                    style="height: 1.149108em; vertical-align: -0.25em"
                  &gt;&lt;/span
                  &gt;&lt;span class="mord"&gt;∇&lt;/span&gt;&lt;span class="mord"&gt;ℓ&lt;/span
                  &gt;&lt;span class="mopen"&gt;(&lt;/span
                  &gt;&lt;span class="mord"
                    &gt;&lt;span
                      class="mord mathdefault"
                      style="margin-right: 0.05017em"
                      &gt;B&lt;/span
                    &gt;&lt;span class="msupsub"
                      &gt;&lt;span class="vlist-t vlist-t2"
                        &gt;&lt;span class="vlist-r"
                          &gt;&lt;span
                            class="vlist"
                            style="height: 0.2805559999999999em"
                            &gt;&lt;span
                              style="
                                top: -2.5500000000000003em;
                                margin-left: -0.05017em;
                                margin-right: 0.05em;
                              "
                              &gt;&lt;span class="pstrut" style="height: 2.7em"&gt;&lt;/span
                              &gt;&lt;span class="sizing reset-size6 size3 mtight"
                                &gt;&lt;span class="mord mathdefault mtight"
                                  &gt;t&lt;/span
                                &gt;&lt;/span
                              &gt;&lt;/span
                            &gt;&lt;/span
                          &gt;&lt;span class="vlist-s"&gt;​&lt;/span&gt;&lt;/span
                        &gt;&lt;span class="vlist-r"
                          &gt;&lt;span class="vlist" style="height: 0.15em"
                            &gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span
                  &gt;&lt;span class="mpunct"&gt;;&lt;/span
                  &gt;&lt;span
                    class="mspace"
                    style="margin-right: 0.16666666666666666em"
                  &gt;&lt;/span
                  &gt;&lt;span class="mord"
                    &gt;&lt;span
                      class="mord mathdefault"
                      style="margin-right: 0.07847em"
                      &gt;X&lt;/span
                    &gt;&lt;span class="msupsub"
                      &gt;&lt;span class="vlist-t vlist-t2"
                        &gt;&lt;span class="vlist-r"
                          &gt;&lt;span
                            class="vlist"
                            style="height: 0.2805559999999999em"
                            &gt;&lt;span
                              style="
                                top: -2.5500000000000003em;
                                margin-left: -0.07847em;
                                margin-right: 0.05em;
                              "
                              &gt;&lt;span class="pstrut" style="height: 2.7em"&gt;&lt;/span
                              &gt;&lt;span class="sizing reset-size6 size3 mtight"
                                &gt;&lt;span class="mord mathdefault mtight"
                                  &gt;t&lt;/span
                                &gt;&lt;/span
                              &gt;&lt;/span
                            &gt;&lt;/span
                          &gt;&lt;span class="vlist-s"&gt;​&lt;/span&gt;&lt;/span
                        &gt;&lt;span class="vlist-r"
                          &gt;&lt;span class="vlist" style="height: 0.15em"
                            &gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span
                  &gt;&lt;span class="mclose"&gt;)&lt;/span
                  &gt;&lt;span class="mclose"
                    &gt;&lt;span class="mclose"&gt;)&lt;/span
                    &gt;&lt;span class="msupsub"
                      &gt;&lt;span class="vlist-t"
                        &gt;&lt;span class="vlist-r"
                          &gt;&lt;span
                            class="vlist"
                            style="height: 0.8991079999999999em"
                            &gt;&lt;span style="top: -3.113em; margin-right: 0.05em"
                              &gt;&lt;span class="pstrut" style="height: 2.7em"&gt;&lt;/span
                              &gt;&lt;span class="sizing reset-size6 size3 mtight"
                                &gt;&lt;span class="mord mtight"&gt;⊤&lt;/span&gt;&lt;/span
                              &gt;&lt;/span
                            &gt;&lt;/span
                          &gt;&lt;/span
                        &gt;&lt;/span
                      &gt;&lt;/span
                    &gt;&lt;/span
                  &gt;&lt;span class="mclose"&gt;]&lt;/span&gt;&lt;/span
                &gt;&lt;/span
              &gt;&lt;/span
            &gt;&lt;/span
          &gt;
        &lt;/p&gt;

        &lt;p&gt;
          where $\gamma$ denotes the optimization seed (which determines the
          minibatch $B_t$). We can assume the gradient for a single datapoint is
          the full batch gradient plus some Gaussian noise (discussed more at
          the end of the post). Then, if we take a minibatch $B_t$ with size
          $B$, the minibatch gradient is:
        &lt;/p&gt;

        &lt;p&gt;
          &lt;span class="katex-display"
            &gt;&lt;span class="katex"
              &gt;&lt;span class="katex-mathml"
                &gt;&lt;math xmlns="http://www.w3.org/1998/Math/MathML"
                  &gt;&lt;semantics
                    &gt;&lt;mrow
                      &gt;&lt;mi mathvariant="normal"&gt;∇&lt;/mi
                      &gt;&lt;mi mathvariant="normal"&gt;ℓ&lt;/mi&gt;&lt;mo stretchy="false"&gt;(&lt;/mo
                      &gt;&lt;msub&gt;&lt;mi&gt;B&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/msub
                      &gt;&lt;mo separator="true"&gt;;&lt;/mo
                      &gt;&lt;msub&gt;&lt;mi&gt;X&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/msub
                      &gt;&lt;mo stretchy="false"&gt;)&lt;/mo&gt;&lt;mo&gt;=&lt;/mo
                      &gt;&lt;mi mathvariant="normal"&gt;∇&lt;/mi
                      &gt;&lt;mi mathvariant="normal"&gt;ℓ&lt;/mi&gt;&lt;mo stretchy="false"&gt;(&lt;/mo
                      &gt;&lt;mi mathvariant="script"&gt;D&lt;/mi&gt;&lt;mo separator="true"&gt;;&lt;/mo
                      &gt;&lt;msub&gt;&lt;mi&gt;X&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/msub
                      &gt;&lt;mo stretchy="false"&gt;)&lt;/mo&gt;&lt;mo&gt;+&lt;/mo
                      &gt;&lt;mfrac&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;mi&gt;B&lt;/mi&gt;&lt;/mfrac
                      &gt;&lt;munderover
                        &gt;&lt;mo&gt;∑&lt;/mo&gt;&lt;mrow&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/mrow
                        &gt;&lt;mi&gt;B&lt;/mi&gt;&lt;/munderover
                      &gt;&lt;msub&gt;&lt;mi&gt;z&lt;/mi&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;/msub&gt;&lt;/mrow
                    &gt;&lt;annotation encoding="application/x-tex"&gt;
                      \nabla \ell(B_t;X_t) = \nabla \ell(\mathcal{D}; X_t) +
                      \frac{1}{B} \sum_{i=1}^B z_i
                    &lt;/annotation&gt;&lt;/semantics
                  &gt;&lt;/math
                &gt;&lt;/span
              &gt;&lt;span class="katex-html" aria-hidden="true"
                &gt;&lt;span class="base"
                  &gt;&lt;span
                    class="strut"
                    style="height: 1em; vertical-align: -0.25em"
                  &gt;&lt;/span
                  &gt;&lt;span class="mord"&gt;∇&lt;/span&gt;&lt;span class="mord"&gt;ℓ&lt;/span
                  &gt;&lt;span class="mopen"&gt;(&lt;/span
                  &gt;&lt;span class="mord"
                    &gt;&lt;span
                      class="mord mathdefault"
                      style="margin-right: 0.05017em"
                      &gt;B&lt;/span
                    &gt;&lt;span class="msupsub"
                      &gt;&lt;span class="vlist-t vlist-t2"
                        &gt;&lt;span class="vlist-r"
                          &gt;&lt;span
                            class="vlist"
                            style="height: 0.2805559999999999em"
                            &gt;&lt;span
                              style="
                                top: -2.5500000000000003em;
                                margin-left: -0.05017em;
                                margin-right: 0.05em;
                              "
                              &gt;&lt;span class="pstrut" style="height: 2.7em"&gt;&lt;/span
                              &gt;&lt;span class="sizing reset-size6 size3 mtight"
                                &gt;&lt;span class="mord mathdefault mtight"
                                  &gt;t&lt;/span
                                &gt;&lt;/span
                              &gt;&lt;/span
                            &gt;&lt;/span
                          &gt;&lt;span class="vlist-s"&gt;​&lt;/span&gt;&lt;/span
                        &gt;&lt;span class="vlist-r"
                          &gt;&lt;span class="vlist" style="height: 0.15em"
                            &gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span
                  &gt;&lt;span class="mpunct"&gt;;&lt;/span
                  &gt;&lt;span
                    class="mspace"
                    style="margin-right: 0.16666666666666666em"
                  &gt;&lt;/span
                  &gt;&lt;span class="mord"
                    &gt;&lt;span
                      class="mord mathdefault"
                      style="margin-right: 0.07847em"
                      &gt;X&lt;/span
                    &gt;&lt;span class="msupsub"
                      &gt;&lt;span class="vlist-t vlist-t2"
                        &gt;&lt;span class="vlist-r"
                          &gt;&lt;span
                            class="vlist"
                            style="height: 0.2805559999999999em"
                            &gt;&lt;span
                              style="
                                top: -2.5500000000000003em;
                                margin-left: -0.07847em;
                                margin-right: 0.05em;
                              "
                              &gt;&lt;span class="pstrut" style="height: 2.7em"&gt;&lt;/span
                              &gt;&lt;span class="sizing reset-size6 size3 mtight"
                                &gt;&lt;span class="mord mathdefault mtight"
                                  &gt;t&lt;/span
                                &gt;&lt;/span
                              &gt;&lt;/span
                            &gt;&lt;/span
                          &gt;&lt;span class="vlist-s"&gt;​&lt;/span&gt;&lt;/span
                        &gt;&lt;span class="vlist-r"
                          &gt;&lt;span class="vlist" style="height: 0.15em"
                            &gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span
                  &gt;&lt;span class="mclose"&gt;)&lt;/span
                  &gt;&lt;span
                    class="mspace"
                    style="margin-right: 0.2777777777777778em"
                  &gt;&lt;/span
                  &gt;&lt;span class="mrel"&gt;=&lt;/span
                  &gt;&lt;span
                    class="mspace"
                    style="margin-right: 0.2777777777777778em"
                  &gt;&lt;/span&gt;&lt;/span
                &gt;&lt;span class="base"
                  &gt;&lt;span
                    class="strut"
                    style="height: 1em; vertical-align: -0.25em"
                  &gt;&lt;/span
                  &gt;&lt;span class="mord"&gt;∇&lt;/span&gt;&lt;span class="mord"&gt;ℓ&lt;/span
                  &gt;&lt;span class="mopen"&gt;(&lt;/span
                  &gt;&lt;span class="mord"
                    &gt;&lt;span class="mord mathcal" style="margin-right: 0.02778em"
                      &gt;D&lt;/span
                    &gt;&lt;/span
                  &gt;&lt;span class="mpunct"&gt;;&lt;/span
                  &gt;&lt;span
                    class="mspace"
                    style="margin-right: 0.16666666666666666em"
                  &gt;&lt;/span
                  &gt;&lt;span class="mord"
                    &gt;&lt;span
                      class="mord mathdefault"
                      style="margin-right: 0.07847em"
                      &gt;X&lt;/span
                    &gt;&lt;span class="msupsub"
                      &gt;&lt;span class="vlist-t vlist-t2"
                        &gt;&lt;span class="vlist-r"
                          &gt;&lt;span
                            class="vlist"
                            style="height: 0.2805559999999999em"
                            &gt;&lt;span
                              style="
                                top: -2.5500000000000003em;
                                margin-left: -0.07847em;
                                margin-right: 0.05em;
                              "
                              &gt;&lt;span class="pstrut" style="height: 2.7em"&gt;&lt;/span
                              &gt;&lt;span class="sizing reset-size6 size3 mtight"
                                &gt;&lt;span class="mord mathdefault mtight"
                                  &gt;t&lt;/span
                                &gt;&lt;/span
                              &gt;&lt;/span
                            &gt;&lt;/span
                          &gt;&lt;span class="vlist-s"&gt;​&lt;/span&gt;&lt;/span
                        &gt;&lt;span class="vlist-r"
                          &gt;&lt;span class="vlist" style="height: 0.15em"
                            &gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span
                  &gt;&lt;span class="mclose"&gt;)&lt;/span
                  &gt;&lt;span
                    class="mspace"
                    style="margin-right: 0.2222222222222222em"
                  &gt;&lt;/span
                  &gt;&lt;span class="mbin"&gt;+&lt;/span
                  &gt;&lt;span
                    class="mspace"
                    style="margin-right: 0.2222222222222222em"
                  &gt;&lt;/span&gt;&lt;/span
                &gt;&lt;span class="base"
                  &gt;&lt;span
                    class="strut"
                    style="height: 3.106005em; vertical-align: -1.277669em"
                  &gt;&lt;/span
                  &gt;&lt;span class="mord"
                    &gt;&lt;span class="mopen nulldelimiter"&gt;&lt;/span
                    &gt;&lt;span class="mfrac"
                      &gt;&lt;span class="vlist-t vlist-t2"
                        &gt;&lt;span class="vlist-r"
                          &gt;&lt;span class="vlist" style="height: 1.32144em"
                            &gt;&lt;span style="top: -2.314em"
                              &gt;&lt;span class="pstrut" style="height: 3em"&gt;&lt;/span
                              &gt;&lt;span class="mord"
                                &gt;&lt;span
                                  class="mord mathdefault"
                                  style="margin-right: 0.05017em"
                                  &gt;B&lt;/span
                                &gt;&lt;/span
                              &gt;&lt;/span
                            &gt;&lt;span style="top: -3.23em"
                              &gt;&lt;span class="pstrut" style="height: 3em"&gt;&lt;/span
                              &gt;&lt;span
                                class="frac-line"
                                style="border-bottom-width: 0.04em"
                              &gt;&lt;/span&gt;&lt;/span
                            &gt;&lt;span style="top: -3.677em"
                              &gt;&lt;span class="pstrut" style="height: 3em"&gt;&lt;/span
                              &gt;&lt;span class="mord"
                                &gt;&lt;span class="mord"&gt;1&lt;/span&gt;&lt;/span
                              &gt;&lt;/span
                            &gt;&lt;/span
                          &gt;&lt;span class="vlist-s"&gt;​&lt;/span&gt;&lt;/span
                        &gt;&lt;span class="vlist-r"
                          &gt;&lt;span class="vlist" style="height: 0.686em"
                            &gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span
                    &gt;&lt;span class="mclose nulldelimiter"&gt;&lt;/span&gt;&lt;/span
                  &gt;&lt;span
                    class="mspace"
                    style="margin-right: 0.16666666666666666em"
                  &gt;&lt;/span
                  &gt;&lt;span class="mop op-limits"
                    &gt;&lt;span class="vlist-t vlist-t2"
                      &gt;&lt;span class="vlist-r"
                        &gt;&lt;span
                          class="vlist"
                          style="height: 1.8283360000000002em"
                          &gt;&lt;span style="top: -1.872331em; margin-left: 0em"
                            &gt;&lt;span class="pstrut" style="height: 3.05em"&gt;&lt;/span
                            &gt;&lt;span class="sizing reset-size6 size3 mtight"
                              &gt;&lt;span class="mord mtight"
                                &gt;&lt;span class="mord mathdefault mtight"&gt;i&lt;/span
                                &gt;&lt;span class="mrel mtight"&gt;=&lt;/span
                                &gt;&lt;span class="mord mtight"&gt;1&lt;/span&gt;&lt;/span
                              &gt;&lt;/span
                            &gt;&lt;/span
                          &gt;&lt;span style="top: -3.050005em"
                            &gt;&lt;span class="pstrut" style="height: 3.05em"&gt;&lt;/span
                            &gt;&lt;span
                              &gt;&lt;span class="mop op-symbol large-op"
                                &gt;∑&lt;/span
                              &gt;&lt;/span
                            &gt;&lt;/span
                          &gt;&lt;span
                            style="top: -4.3000050000000005em; margin-left: 0em"
                            &gt;&lt;span class="pstrut" style="height: 3.05em"&gt;&lt;/span
                            &gt;&lt;span class="sizing reset-size6 size3 mtight"
                              &gt;&lt;span
                                class="mord mathdefault mtight"
                                style="margin-right: 0.05017em"
                                &gt;B&lt;/span
                              &gt;&lt;/span
                            &gt;&lt;/span
                          &gt;&lt;/span
                        &gt;&lt;span class="vlist-s"&gt;​&lt;/span&gt;&lt;/span
                      &gt;&lt;span class="vlist-r"
                        &gt;&lt;span class="vlist" style="height: 1.277669em"
                          &gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span
                  &gt;&lt;span
                    class="mspace"
                    style="margin-right: 0.16666666666666666em"
                  &gt;&lt;/span
                  &gt;&lt;span class="mord"
                    &gt;&lt;span
                      class="mord mathdefault"
                      style="margin-right: 0.04398em"
                      &gt;z&lt;/span
                    &gt;&lt;span class="msupsub"
                      &gt;&lt;span class="vlist-t vlist-t2"
                        &gt;&lt;span class="vlist-r"
                          &gt;&lt;span
                            class="vlist"
                            style="height: 0.31166399999999994em"
                            &gt;&lt;span
                              style="
                                top: -2.5500000000000003em;
                                margin-left: -0.04398em;
                                margin-right: 0.05em;
                              "
                              &gt;&lt;span class="pstrut" style="height: 2.7em"&gt;&lt;/span
                              &gt;&lt;span class="sizing reset-size6 size3 mtight"
                                &gt;&lt;span class="mord mathdefault mtight"
                                  &gt;i&lt;/span
                                &gt;&lt;/span
                              &gt;&lt;/span
                            &gt;&lt;/span
                          &gt;&lt;span class="vlist-s"&gt;​&lt;/span&gt;&lt;/span
                        &gt;&lt;span class="vlist-r"
                          &gt;&lt;span class="vlist" style="height: 0.15em"
                            &gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span
          &gt;&lt;/span&gt;
        &lt;/p&gt;

        &lt;p&gt;
          where $z_i\sim \mathcal{N}(0, \Sigma(X_t))$ is drawn i.i.d. per
          datapoint. Then, when we scale the batch size, we can see how the
          gradient noise, and thus the diffusion term in the SDE, changes.
        &lt;/p&gt;

        &lt;p&gt;
          &lt;span class="katex-display"
            &gt;&lt;span class="katex"
              &gt;&lt;span class="katex-mathml"
                &gt;&lt;math xmlns="http://www.w3.org/1998/Math/MathML"
                  &gt;&lt;semantics
                    &gt;&lt;mrow
                      &gt;&lt;msup
                        &gt;&lt;mi mathvariant="normal"&gt;Σ&lt;/mi
                        &gt;&lt;mrow
                          &gt;&lt;mo stretchy="false"&gt;(&lt;/mo&gt;&lt;mi&gt;κ&lt;/mi&gt;&lt;mi&gt;B&lt;/mi
                          &gt;&lt;mo stretchy="false"&gt;)&lt;/mo&gt;&lt;/mrow
                        &gt;&lt;/msup
                      &gt;&lt;mo stretchy="false"&gt;(&lt;/mo
                      &gt;&lt;msub&gt;&lt;mi&gt;X&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/msub
                      &gt;&lt;mo stretchy="false"&gt;)&lt;/mo&gt;&lt;mo&gt;=&lt;/mo
                      &gt;&lt;mfrac&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;mi&gt;κ&lt;/mi&gt;&lt;/mfrac
                      &gt;&lt;msup
                        &gt;&lt;mi mathvariant="normal"&gt;Σ&lt;/mi
                        &gt;&lt;mrow
                          &gt;&lt;mo stretchy="false"&gt;(&lt;/mo&gt;&lt;mi&gt;B&lt;/mi
                          &gt;&lt;mo stretchy="false"&gt;)&lt;/mo&gt;&lt;/mrow
                        &gt;&lt;/msup
                      &gt;&lt;mo stretchy="false"&gt;(&lt;/mo
                      &gt;&lt;msub&gt;&lt;mi&gt;X&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/msub
                      &gt;&lt;mo stretchy="false"&gt;)&lt;/mo&gt;&lt;/mrow
                    &gt;&lt;annotation encoding="application/x-tex"&gt;
                      \Sigma^{(\kappa B)}(X_t) = \frac 1\kappa \Sigma^{(B)}(X_t)
                    &lt;/annotation&gt;&lt;/semantics
                  &gt;&lt;/math
                &gt;&lt;/span
              &gt;&lt;span class="katex-html" aria-hidden="true"
                &gt;&lt;span class="base"
                  &gt;&lt;span
                    class="strut"
                    style="height: 1.188em; vertical-align: -0.25em"
                  &gt;&lt;/span
                  &gt;&lt;span class="mord"
                    &gt;&lt;span class="mord"&gt;Σ&lt;/span
                    &gt;&lt;span class="msupsub"
                      &gt;&lt;span class="vlist-t"
                        &gt;&lt;span class="vlist-r"
                          &gt;&lt;span class="vlist" style="height: 0.938em"
                            &gt;&lt;span style="top: -3.113em; margin-right: 0.05em"
                              &gt;&lt;span class="pstrut" style="height: 2.7em"&gt;&lt;/span
                              &gt;&lt;span class="sizing reset-size6 size3 mtight"
                                &gt;&lt;span class="mord mtight"
                                  &gt;&lt;span class="mopen mtight"&gt;(&lt;/span
                                  &gt;&lt;span class="mord mathdefault mtight"&gt;κ&lt;/span
                                  &gt;&lt;span
                                    class="mord mathdefault mtight"
                                    style="margin-right: 0.05017em"
                                    &gt;B&lt;/span
                                  &gt;&lt;span class="mclose mtight"&gt;)&lt;/span&gt;&lt;/span
                                &gt;&lt;/span
                              &gt;&lt;/span
                            &gt;&lt;/span
                          &gt;&lt;/span
                        &gt;&lt;/span
                      &gt;&lt;/span
                    &gt;&lt;/span
                  &gt;&lt;span class="mopen"&gt;(&lt;/span
                  &gt;&lt;span class="mord"
                    &gt;&lt;span
                      class="mord mathdefault"
                      style="margin-right: 0.07847em"
                      &gt;X&lt;/span
                    &gt;&lt;span class="msupsub"
                      &gt;&lt;span class="vlist-t vlist-t2"
                        &gt;&lt;span class="vlist-r"
                          &gt;&lt;span
                            class="vlist"
                            style="height: 0.2805559999999999em"
                            &gt;&lt;span
                              style="
                                top: -2.5500000000000003em;
                                margin-left: -0.07847em;
                                margin-right: 0.05em;
                              "
                              &gt;&lt;span class="pstrut" style="height: 2.7em"&gt;&lt;/span
                              &gt;&lt;span class="sizing reset-size6 size3 mtight"
                                &gt;&lt;span class="mord mathdefault mtight"
                                  &gt;t&lt;/span
                                &gt;&lt;/span
                              &gt;&lt;/span
                            &gt;&lt;/span
                          &gt;&lt;span class="vlist-s"&gt;​&lt;/span&gt;&lt;/span
                        &gt;&lt;span class="vlist-r"
                          &gt;&lt;span class="vlist" style="height: 0.15em"
                            &gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span
                  &gt;&lt;span class="mclose"&gt;)&lt;/span
                  &gt;&lt;span
                    class="mspace"
                    style="margin-right: 0.2777777777777778em"
                  &gt;&lt;/span
                  &gt;&lt;span class="mrel"&gt;=&lt;/span
                  &gt;&lt;span
                    class="mspace"
                    style="margin-right: 0.2777777777777778em"
                  &gt;&lt;/span&gt;&lt;/span
                &gt;&lt;span class="base"
                  &gt;&lt;span
                    class="strut"
                    style="height: 2.00744em; vertical-align: -0.686em"
                  &gt;&lt;/span
                  &gt;&lt;span class="mord"
                    &gt;&lt;span class="mopen nulldelimiter"&gt;&lt;/span
                    &gt;&lt;span class="mfrac"
                      &gt;&lt;span class="vlist-t vlist-t2"
                        &gt;&lt;span class="vlist-r"
                          &gt;&lt;span class="vlist" style="height: 1.32144em"
                            &gt;&lt;span style="top: -2.314em"
                              &gt;&lt;span class="pstrut" style="height: 3em"&gt;&lt;/span
                              &gt;&lt;span class="mord mathdefault"&gt;κ&lt;/span&gt;&lt;/span
                            &gt;&lt;span style="top: -3.23em"
                              &gt;&lt;span class="pstrut" style="height: 3em"&gt;&lt;/span
                              &gt;&lt;span
                                class="frac-line"
                                style="border-bottom-width: 0.04em"
                              &gt;&lt;/span&gt;&lt;/span
                            &gt;&lt;span style="top: -3.677em"
                              &gt;&lt;span class="pstrut" style="height: 3em"&gt;&lt;/span
                              &gt;&lt;span class="mord"&gt;1&lt;/span&gt;&lt;/span
                            &gt;&lt;/span
                          &gt;&lt;span class="vlist-s"&gt;​&lt;/span&gt;&lt;/span
                        &gt;&lt;span class="vlist-r"
                          &gt;&lt;span class="vlist" style="height: 0.686em"
                            &gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span
                    &gt;&lt;span class="mclose nulldelimiter"&gt;&lt;/span&gt;&lt;/span
                  &gt;&lt;span class="mord"
                    &gt;&lt;span class="mord"&gt;Σ&lt;/span
                    &gt;&lt;span class="msupsub"
                      &gt;&lt;span class="vlist-t"
                        &gt;&lt;span class="vlist-r"
                          &gt;&lt;span class="vlist" style="height: 0.938em"
                            &gt;&lt;span style="top: -3.113em; margin-right: 0.05em"
                              &gt;&lt;span class="pstrut" style="height: 2.7em"&gt;&lt;/span
                              &gt;&lt;span class="sizing reset-size6 size3 mtight"
                                &gt;&lt;span class="mord mtight"
                                  &gt;&lt;span class="mopen mtight"&gt;(&lt;/span
                                  &gt;&lt;span
                                    class="mord mathdefault mtight"
                                    style="margin-right: 0.05017em"
                                    &gt;B&lt;/span
                                  &gt;&lt;span class="mclose mtight"&gt;)&lt;/span&gt;&lt;/span
                                &gt;&lt;/span
                              &gt;&lt;/span
                            &gt;&lt;/span
                          &gt;&lt;/span
                        &gt;&lt;/span
                      &gt;&lt;/span
                    &gt;&lt;/span
                  &gt;&lt;span class="mopen"&gt;(&lt;/span
                  &gt;&lt;span class="mord"
                    &gt;&lt;span
                      class="mord mathdefault"
                      style="margin-right: 0.07847em"
                      &gt;X&lt;/span
                    &gt;&lt;span class="msupsub"
                      &gt;&lt;span class="vlist-t vlist-t2"
                        &gt;&lt;span class="vlist-r"
                          &gt;&lt;span
                            class="vlist"
                            style="height: 0.2805559999999999em"
                            &gt;&lt;span
                              style="
                                top: -2.5500000000000003em;
                                margin-left: -0.07847em;
                                margin-right: 0.05em;
                              "
                              &gt;&lt;span class="pstrut" style="height: 2.7em"&gt;&lt;/span
                              &gt;&lt;span class="sizing reset-size6 size3 mtight"
                                &gt;&lt;span class="mord mathdefault mtight"
                                  &gt;t&lt;/span
                                &gt;&lt;/span
                              &gt;&lt;/span
                            &gt;&lt;/span
                          &gt;&lt;span class="vlist-s"&gt;​&lt;/span&gt;&lt;/span
                        &gt;&lt;span class="vlist-r"
                          &gt;&lt;span class="vlist" style="height: 0.15em"
                            &gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span
                  &gt;&lt;span class="mclose"&gt;)&lt;/span&gt;&lt;/span
                &gt;&lt;/span
              &gt;&lt;/span
            &gt;&lt;/span
          &gt;
        &lt;/p&gt;

        &lt;p&gt;
          This equation also matches our intuitions: when we increase the batch
          size, the noise (i.e., diffusion) in the process should get smaller,
          since the trajectory is closer to full-batch GD. Now we have basically
          everything we need in order to derive some scaling rules!
        &lt;/p&gt;

        &lt;h2 id="deriving-scaling-rules-from-sdes"&gt;
          Deriving Scaling Rules from SDEs
        &lt;/h2&gt;

        &lt;p&gt;
          Now that we have the basics of SDEs down, we can go back to the
          scaling rules at the start of the post.
        &lt;/p&gt;

        &lt;hr /&gt;

        &lt;p&gt;&lt;strong&gt;Linear Scaling Rule (for SGD)&lt;/strong&gt;&lt;/p&gt;

        &lt;p&gt;
          When scaling the batch size by $\kappa$, scale the learning rate also
          by $\kappa$.
        &lt;/p&gt;

        &lt;p&gt;&lt;strong&gt;Square Root Scaling Rule (for Adam)&lt;/strong&gt;&lt;/p&gt;

        &lt;p&gt;
          When scaling the batch size by $\kappa$, scale the learning rate by
          $\sqrt{\kappa}$. Also, change the other hyperparameters, setting
          $\beta_1 = 1 - \kappa(1-\beta_1)$, $\beta_2 = 1-\kappa(1-\beta_2)$,
          and $\epsilon = \epsilon / \sqrt{\kappa}$.
        &lt;/p&gt;

        &lt;hr /&gt;

        &lt;p&gt;
          The reasoning to derive scaling rules from the SDE goes like this.
          Changing the batch size scales $\Sigma(X_t)$ by $1/\kappa$ so one
          should scaling the learning rate proportionally to
          &lt;strong&gt;preserve the scale of the diffusion term&lt;/strong&gt;. Language
          models are usually trained with the Adam optimizer, which requires a
          different SDE approximation (see our NeurIPS 2022 paper
          &lt;a href="https://openreview.net/forum?id=F2mhzjHkQP"&gt;here&lt;/a&gt;). That
          SDE, through similar reasoning, yields the square root scaling rule.
        &lt;/p&gt;

        &lt;p&gt;
          Keep in mind that the accuracy of the SDE approximation is a
          sufficient condition for the scaling rule: if the SDE is a good
          approximation, then the scaling rule should hold. So when does the SDE
          approximation break?
          &lt;a href="https://openreview.net/forum?id=F2mhzjHkQP"
            &gt;Section 4.1 of our paper&lt;/a
          &gt;
          illustrates a good warmup setting to build intuition on when the
          gradient noise gets to be too small and the SDE approximation is no
          longer good. If you find your hyperparameters taken to extremes by
          these scaling rules, consider shifting your baseline run to a larger
          batch size.
        &lt;/p&gt;

        &lt;h1 id="what-do-scaling-rules-preserve"&gt;
          What do Scaling Rules Preserve?
        &lt;/h1&gt;

        &lt;p&gt;
          Traditionally, when training vision models with SGD, the empirical
          goal of using a scaling rule is to preserve the test accuracy of the
          model. But for language models trained with Adam,
          &lt;strong&gt;what metrics are we interested in?&lt;/strong&gt; Perplexity?
          In-context learning ability? Or even more broadly, factuality?
          Truthfulness? Fortunately, when we use the formal language of SDEs, we
          can show results that conceptually encapsulate all of these
          possibilities (and ones we haven’t dreamed up yet).
        &lt;/p&gt;

        &lt;p&gt;
          To recap, the way that one gets from SDEs to scaling rules is as
          follows. Changing the batch size changes the equation of the SDE.
          Adjusting the optimization hyperparameters according to the scaling
          rule changes the equation of the SDE back to what it was originally
          with the baseline run. If the SDE is indeed a
          &lt;em&gt;faithful approximation&lt;/em&gt; of SGD/Adam, then preserving the SDE
          equation under changing batch sizes will also preserve the
          optimization trajectory of the baseline run of SGD/Adam.
        &lt;/p&gt;

        &lt;p&gt;
          So now we enter the question of what a “faithful approximation” is. As
          I hinted before, we might only care about
          &lt;strong&gt;preserving certain features of our trained model&lt;/strong&gt;
          (e.g., test accuracy or truthfulness) under changing batch sizes. And
          we have
          &lt;strong
            &gt;no idea how these features change with the actual
            parameters&lt;/strong
          &gt;
          — for example, two models may have parameters that are very close to
          each other but have different test accuracies. Yet another
          complicating factor is that
          &lt;strong&gt;SGD and Adam are stochastic&lt;/strong&gt; (i.e., depend on the
          optimization seed), so we have to figure out what it means for two
          stochastic trajectories to approximate each other.
        &lt;/p&gt;

        &lt;p&gt;
          Altogether, we need a more sophisticated notion of approximation than
          just saying the two trajectories are close to each other in $\ell_2$
          norm. SDEs handle this issue using &lt;strong&gt;test functions&lt;/strong&gt;,
          which compute something as a function of the model parameters. This
          notion of approximation, called a
          &lt;strong&gt;weak approximation&lt;/strong&gt; (&lt;a
            href="https://jmlr.org/papers/v20/17-526.html"
            &gt;Li et al., 2019&lt;/a
          &gt;), gives a guarantee that the SDEs and SGD/Adam produce models that
          have close values of all possible test functions. This is a very broad
          class of functions, and it likely includes all of the ones we would
          care about. The tricky thing is, to make the theory work out, we have
          to assume something fairly general about these test functions — that
          they grow at most polynomially with the model parameters — and we
          can’t be confident that test accuracy, truthfulness, etc. satisfy this
          assumption. But, fortunately,
          &lt;strong
            &gt;extensive experiments in our papers (&lt;a
              href="https://openreview.net/forum?id=goEdyJ_nVQI"
              &gt;on SGD&lt;/a
            &gt;
            and
            &lt;a
              href="https://proceedings.neurips.cc/paper_files/paper/2022/hash/32ac710102f0620d0f28d5d05a44fe08-Abstract-Conference.html"
              &gt;on Adam&lt;/a
            &gt;) show that scaling rules preserve various test functions,
            including test accuracy, perplexity, and even post-fine-tuning
            performance&lt;/strong
          &gt;
          (on GLUE tasks).
        &lt;/p&gt;

        &lt;center&gt;
          &lt;div class="figure"&gt;
            &lt;img
              src="/assets/sde_img/roberta.png"
              alt="Training RoBERTa models on Wiki+Books with different batch sizes using the square root scaling rule."
              style="margin-left: auto; margin-right: auto; width: 40%"
            /&gt;
            &lt;img
              src="/assets/sde_img/gpt.png"
              alt="Training RoBERTa models on Wiki+Books with different batch sizes using the square root scaling rule."
              style="margin-left: auto; margin-right: auto; width: 40%"
            /&gt;
            &lt;br /&gt;

            &lt;div class="caption"&gt;
              &lt;span class="caption-label"
                &gt;Figures from
                &lt;a href="https://openreview.net/forum?id=F2mhzjHkQP"
                  &gt;our NeurIPS 2022 paper.&lt;/a
                &gt;&lt;/span
              &gt;
              Training RoBERTa models on Wiki+Books (left) and GPT-2 on
              WikiText-103 (right) with different batch sizes. Adjusting the
              Adam optimizer hyperparameters per the square root scaling rule
              ensures performance is preserved!
            &lt;/div&gt;
          &lt;/div&gt;
        &lt;/center&gt;

        &lt;h1 id="discussions-and-extensions"&gt;Discussions and Extensions&lt;/h1&gt;

        &lt;p&gt;
          Here we discuss some of the more nuanced points of SDEs for those who
          are interested. I will dive more into these ideas in a subsequent
          post, which will contain more of the theoretical underpinnings of
          SDEs.
        &lt;/p&gt;

        &lt;ol&gt;
          &lt;li&gt;
            &lt;p&gt;
              &lt;strong&gt;Is SGD gradient noise actually Gaussian?&lt;/strong&gt; In the
              SDE section, I assumed that the gradient of one datapoint is the
              full-batch gradient plus some Gaussian noise. It is generally
              agreed that gradient noise is additive, but a big topic of
              discussion in the SDE community is whether the gradient noise is
              actually Gaussian or if it’s “heavy-tailed” (i.e.,
              third-and-higher moments are non-negligible). The SDE derived
              above, called an Ito SDE, uses
              &lt;a href="https://en.wikipedia.org/wiki/Brownian_motion"
                &gt;Brownian motion&lt;/a
              &gt;
              for the diffusion term, but some argue that we should instead use
              a the more general
              &lt;a href="https://en.wikipedia.org/wiki/Lévy_process"
                &gt;Levy process&lt;/a
              &gt;
              for the diffusion, which yields a Levy SDE (&lt;a
                href="http://proceedings.mlr.press/v97/simsekli19a/simsekli19a.pdf"
                &gt;Simsekli et al., 2019&lt;/a
              &gt;). These are usually harder to analyze, and much of the empirical
              evidence motivating the usage of a Levy SDE has been disproven (&lt;a
                href="https://openreview.net/forum?id=wXgk_iCiYGo"
                &gt;Xie et al., 2021&lt;/a
              &gt;). Moreover,
              &lt;a href="https://openreview.net/forum?id=goEdyJ_nVQI"
                &gt;our work in NeurIPS 2021&lt;/a
              &gt;
              showed empirically that using Gaussian noise with the first and
              second moments matching the naturally occurring noise in SGD is
              sufficient to preserve the generalization performance of SGD in
              vision. For these reasons, and also because of the ease of using
              an Ito SDE for analysis, the Ito SDE remains the common choice for
              approximating discrete optimization trajectories. More recently,
              however, the SGD noise has been considered to be heavy-tailed when
              nodes in a distributed setting fail unexpectedly (&lt;a
                href="https://openreview.net/forum?id=C6PiH9Fkjd"
                &gt;Schaipp et al., 2023&lt;/a
              &gt;). Due to the computational resources to make these measurements,
              it’s hard to collect evidence of which setting is more faithful to
              language modeling training, but our insights derived using Ito
              SDEs (&lt;a href="https://openreview.net/forum?id=F2mhzjHkQP"
                &gt;in our NeurIPS 2022 paper&lt;/a
              &gt;) generally preserve the performance of language models when
              performing pre-training and fine-tuning. Other analyses using
              similar noise assumptions, including our
              &lt;a href="https://arxiv.org/abs/2307.15196"
                &gt;upcoming ICLR 2024 paper&lt;/a
              &gt;
              showing that using momentum does not change the SGD trajectory
              when using small learning rates, also empirically hold when
              training language models.
            &lt;/p&gt;
          &lt;/li&gt;
          &lt;li&gt;
            &lt;p&gt;
              &lt;strong&gt;Implicit Bias of SGD&lt;/strong&gt;: The SDE that I described
              doesn’t directly give much information about the implicit bias
              (i.e., the ability of SGD to choose generalizing solutions out of
              many possible empirical risk minimizers). Recent works have
              studied using the SDE to describe the late phase of training,
              where the training loss is nearly zero, and the primary term
              driving the trajectory is the diffusion (&lt;a
                href="https://openreview.net/forum?id=siCt4xZn5Ve"
                &gt;Li et al., 2022&lt;/a
              &gt;). Analyzing this SDE reveals that SGD tends towards flatter
              minima. It is not totally clear how flatness and generalization
              are related to each other (&lt;a
                href="https://openreview.net/forum?id=SJgIPJBFvH"
                &gt;Jiang et al., 2020&lt;/a
              &gt;;
              &lt;a href="https://openreview.net/forum?id=VZp9X410D3"
                &gt;Andriushchenko et al., 2023&lt;/a
              &gt;;
              &lt;a href="https://openreview.net/forum?id=Dkmpa6wCIx"
                &gt;Wen et al., 2023&lt;/a
              &gt;), but a careful analysis of the SDE may yield a more nuanced
              understanding than classical generalization measures.
            &lt;/p&gt;
          &lt;/li&gt;
        &lt;/ol&gt;

        &lt;h1 id="conclusion"&gt;Conclusion&lt;/h1&gt;

        &lt;p&gt;
          SDEs are a powerful, generalizable tool for studying stochastic
          optimization. The results are mostly agnostic to the model
          architecture and dataset, so just a little bit of understanding goes a
          long way and provides useful empirical insights. In a separate post,
          I’ll talk a bit more about the proof techniques and rigorous
          considerations involved in using SDEs to approximate discrete
          optimization.
        &lt;/p&gt;

        &lt;p&gt;
          &lt;strong&gt;Acknowledgements&lt;/strong&gt;: Thanks (in alphabetical order) to
          Tianyu Gao, Surbhi Goel, Bingbin Liu, Kaifeng Lyu, Abhi Venigalla,
          Mengzhou Xia, and Howard Yen for their feedback on this post! I’m
          deeply indebted to Zhiyuan Li for patiently teaching me about SDEs and
          their technical features. Feedback from Twitter prompted me to add a
          link to the section in our paper that provides an easy setting to work
          out the scaling rule in. The works I mention in this paper were
          co-authored with (in alphabetical order) Sanjeev Arora, Zhiyuan Li,
          Kaifeng Lyu, Abhishek Panigrahi, Runzhe Wang, and Tianhao Wang.
        &lt;/p&gt;
      </content></entry></feed>