# Rethinking Learning Rate and Batch Size (Part 1): The Status Quo

Source: <https://kexue.fm/archives/11260>

Machine translated. Refer to the original post to verify accuracy.

---
author: Su Jianlin
date: September 1, 2025
lang: en
markdown_url: 11260-Rethinking-Learning-Rate-and-Batch-Size-I-Current-Status.md
next_url: 11250-Cool-Papers-Update-Simple-Adaptation-for-Zotero-Connector.html
post_id: 11260
prev_url: 11267-Why-is-Adam-s-Update-RMS-0.2.html
site_title: English (unofficial) translations of posts at kexue.fm
source_url: "https://kexue.fm/archives/11260"
title: "Rethinking Learning Rate and Batch Size (Part 1): The Status Quo"
---

In the previous articles [*When the Batch Size Increases, How Should the Learning Rate Change Accordingly?*](./10542-How-Should-the-Learning-Rate-Change-as-the-Batch-Size-Increases.html) and [*How Does Adam’s Epsilon Affect the Scaling Law of the Learning Rate?*](./10563-How-Does-Adam-s-Epsilon-Affect-the-Scaling-Law-of-Learning-Rate.html), we theoretically discussed how the learning rate varies with the batch size, where the classic part is the second-order expansion analysis proposed by OpenAI. However, when we want to handle non-SGD optimizers, the computation involved in this analytical approach often becomes rather complicated, giving one a feeling of not knowing where to start.

In the next few articles, the author will reorganize and rethink the relevant details in the aforementioned articles, attempt to simplify some of the derivation steps, provide a more general and lighter derivation path, and explore the possibility of extending it to the Muon optimizer.

# Overview of the Method

First, let us review the previous analytical method. In [*When the Batch Size Increases, How Should the Learning Rate Change Accordingly?*](./10542-How-Should-the-Learning-Rate-Change-as-the-Batch-Size-Increases.html), we introduced several lines of thought for analyzing the relationship between the learning rate and the batch size, among which the second-order approximate analysis proposed by OpenAI in [*An Empirical Model of Large-Batch Training*](https://papers.cool/arxiv/1812.06162) took up most of the space; this article also follows the same idea.

Next, we need to introduce some notation. Let the loss function be $`\mathcal{L}(\boldsymbol{w})`$, where $`\boldsymbol{w}\in\mathbb{R}^N`$ is the parameter vector and $`\boldsymbol{g}`$ is its gradient. Note that the ideal loss function is the expectation computed over all training samples, but in practice we can only sample a batch to compute it, which makes the gradient stochastic as well. We denote the gradient of a single sample by $`\tilde{\boldsymbol{g}}`$, whose mean is $`\boldsymbol{g}`$ and whose covariance matrix is $`\boldsymbol{\Sigma}`$; when the batch size is $`B`$, the gradient is denoted by $`\tilde{\boldsymbol{g}}_B`$, whose mean is still $`\boldsymbol{g}`$ but whose covariance matrix becomes $`\boldsymbol{\Sigma}/B`$.

Furthermore, let the current learning rate be $`\eta`$ and the update vector be $`\tilde{\boldsymbol{\varphi}}_B`$; then the loss function after the update will be
``` math
\begin{equation}\begin{aligned}
\mathcal{L}(\boldsymbol{w} - \eta\tilde{\boldsymbol{\varphi}}_B) \approx&\, \mathcal{L}(\boldsymbol{w}) - \eta \tilde{\boldsymbol{\varphi}}_B^{\top}\boldsymbol{g} + \frac{1}{2}\eta^2\tilde{\boldsymbol{\varphi}}_B^{\top}\boldsymbol{H}\tilde{\boldsymbol{\varphi}}_B \\
=&\, \mathcal{L}(\boldsymbol{w}) - \eta \tilde{\boldsymbol{\varphi}}_B^{\top}\boldsymbol{g} + \frac{1}{2}\eta^2\mathop{\mathrm{tr}}(\tilde{\boldsymbol{\varphi}}_B\tilde{\boldsymbol{\varphi}}_B^{\top}\boldsymbol{H})
\end{aligned}\end{equation}
```
On the right-hand side we have Taylor-expanded to second order; $`\boldsymbol{H}`$ is the Hessian matrix, $`\mathop{\mathrm{tr}}`$ denotes the trace of a matrix, and the second equality uses the identity $`\mathop{\mathrm{tr}}(\boldsymbol{A}\boldsymbol{B})=\mathop{\mathrm{tr}}(\boldsymbol{B}\boldsymbol{A})`$. In order to obtain a deterministic result, we take expectations on both sides:
``` math
\begin{equation}\mathbb{E}[\mathcal{L}(\boldsymbol{w} - \eta\tilde{\boldsymbol{\varphi}}_B)] \approx \mathcal{L}(\boldsymbol{w}) - \eta\mathbb{E}[\tilde{\boldsymbol{\varphi}}_B]^{\top}\boldsymbol{g} + \frac{1}{2}\eta^2 \mathop{\mathrm{tr}}(\mathbb{E}[\tilde{\boldsymbol{\varphi}}_B\tilde{\boldsymbol{\varphi}}_B^{\top}]\boldsymbol{H})\end{equation}
```
We view the right-hand side as a quadratic function of $`\eta`$ and assume that the quadratic coefficient is positive (a stronger assumption is that the matrix $`\boldsymbol{H}`$ is positive definite); then we can obtain the minimizer
``` math
\begin{equation}\eta^* \approx \frac{\mathbb{E}[\tilde{\boldsymbol{\varphi}}_B]^{\top}\boldsymbol{g}}{\mathop{\mathrm{tr}}(\mathbb{E}[\tilde{\boldsymbol{\varphi}}_B\tilde{\boldsymbol{\varphi}}_B^{\top}]\boldsymbol{H})}\end{equation}
```
This is the learning rate that makes the loss function decrease fastest <u>on average</u>, i.e., the theoretical optimum of the learning rate. What we need to do is to compute $`\mathbb{E}[\tilde{\boldsymbol{\varphi}}_B]`$ and $`\mathbb{E}[\tilde{\boldsymbol{\varphi}}_B\tilde{\boldsymbol{\varphi}}_B^{\top}]`$ for the specific $`\tilde{\boldsymbol{\varphi}}_B`$, and then extract from the above formula its relationship with the batch size (i.e., $`B`$).

# Warm-up Exercise

As the first example, we naturally consider the simplest case, SGD, for which $`\tilde{\boldsymbol{\varphi}}_B=\tilde{\boldsymbol{g}}_B`$. Then we can easily obtain $`\mathbb{E}[\tilde{\boldsymbol{\varphi}}_B]=\boldsymbol{g}`$ and $`\mathbb{E}[\tilde{\boldsymbol{\varphi}}_B\tilde{\boldsymbol{\varphi}}_B^{\top}]=\boldsymbol{g}\boldsymbol{g}^{\top} + \boldsymbol{\Sigma}/B`$, so we have
``` math
\begin{equation}\eta^* \approx \frac{\boldsymbol{g}^{\top}\boldsymbol{g}}{\mathop{\mathrm{tr}}((\boldsymbol{g}\boldsymbol{g}^{\top} + \boldsymbol{\Sigma}/B)\boldsymbol{H})} = \frac{\boldsymbol{g}^{\top}\boldsymbol{g}}{\boldsymbol{g}^{\top}\boldsymbol{H}\boldsymbol{g} + \mathop{\mathrm{tr}}(\boldsymbol{\Sigma}\boldsymbol{H})/B} = \frac{\eta_{\max}}{1 + \mathcal{B}_{\text{noise}}/B}\label{eq:eta-sgd}\end{equation}
```
where
``` math
\begin{equation}\eta_{\max} = \frac{\boldsymbol{g}^{\top}\boldsymbol{g}}{\boldsymbol{g}^{\top}\boldsymbol{H}\boldsymbol{g}},\qquad\mathcal{B}_{\text{noise}} = \frac{\mathop{\mathrm{tr}}(\boldsymbol{\Sigma}\boldsymbol{H})}{\boldsymbol{g}^{\top}\boldsymbol{H}\boldsymbol{g}}\end{equation}
```

There are several ways to interpret the result $`\eqref{eq:eta-sgd}`$. First, it is a monotonically increasing but bounded function, with the upper bound $`\eta_{\max}`$, which indicates that the learning rate cannot increase indefinitely; compared with the simple linear law or the square-root law, this agrees better with our intuitive understanding. When $`B \ll \mathcal{B}_{\text{noise}}`$, we have
``` math
\begin{equation}\eta^* \approx \frac{\eta_{\max}}{1 + \mathcal{B}_{\text{noise}}/B} \approx \frac{\eta_{\max}}{\mathcal{B}_{\text{noise}}/B} = \eta_{\max} B / \mathcal{B}_{\text{noise}}\end{equation}
```
which shows that when the batch size is relatively small, the learning rate of SGD is indeed linear in the batch size; it also suggests that $`\mathcal{B}_{\text{noise}}`$ is a key statistic. However, the definition of $`\mathcal{B}_{\text{noise}}`$ depends on the Hessian matrix $`\boldsymbol{H}`$, which is almost impossible to compute exactly for LLMs, so in practice we usually assume that it is (a multiple of) the identity matrix, obtaining the simplified form
``` math
\begin{equation}\mathcal{B}_{\text{simple}} = \frac{\mathop{\mathrm{tr}}(\boldsymbol{\Sigma})}{\boldsymbol{g}^{\top}\boldsymbol{g}}\end{equation}
```
This result has the form of the noise strength ($`\mathop{\mathrm{tr}}(\boldsymbol{\Sigma})`$) divided by the signal strength ($`\boldsymbol{g}^{\top}\boldsymbol{g}`$); it is essentially the reciprocal of the signal-to-noise ratio, indicating that the smaller the signal-to-noise ratio, the larger the batch size needed in order to use the same $`\eta_{\max}`$, which also agrees with our intuitive understanding. $`\mathop{\mathrm{tr}}(\boldsymbol{\Sigma})`$ depends only on the diagonal elements of $`\boldsymbol{\Sigma}`$, which means that we only need to estimate the mean and the variance of each parameter independently, and this is feasible in practice.

# Data Efficiency

Besides the direct relationship between the learning rate and the batch size, the author believes that the asymptotic relationship between the amount of training data and the number of training steps derived from it is also a brilliant part that must be learned. In particular, this conclusion seems to be more general than the learning rate relation $`\eqref{eq:eta-sgd}`$, because, as we will see later, SignSGD also yields a conclusion of the same form, yet its learning rate rule is not given by Equation $`\eqref{eq:eta-sgd}`$.

The original paper’s discussion of this part is rather involved; the derivation below has been simplified by the author. Specifically, substituting $`\eta^*`$ back into $`\mathcal{L}(\boldsymbol{w} - \eta\tilde{\boldsymbol{g}}_B)`$, we obtain
``` math
\begin{equation}\overline{\Delta\mathcal{L}} = \mathcal{L}(\boldsymbol{w}) - \mathbb{E}[\mathcal{L}(\boldsymbol{w} - \eta^*\tilde{\boldsymbol{g}}_B)] \approx \frac{\Delta\mathcal{L}_{\max}}{1 + \mathcal{B}_{\text{noise}}/B}\end{equation}
```
where $`\Delta\mathcal{L}_{\max} = \frac{(\boldsymbol{g}^{\top}\boldsymbol{g})^2}{2\boldsymbol{g}^{\top}\boldsymbol{H}\boldsymbol{g}}`$. How should we understand this result? First, it is a monotonically increasing function of $`B`$, equal to $`\Delta\mathcal{L}_{\max}`$ when $`B\to\infty`$; in other words, if we could use an infinitely large batch size, then the loss decrease per step would be $`\Delta\mathcal{L}_{\max}`$, and the number of training steps required in that case is minimal, denoted by $`S_{\min}`$.

If the batch size is a finite value, then the average loss decrease per step is only $`\overline{\Delta\mathcal{L}}`$, which means that on average we have to spend $`1 + \mathcal{B}_{\text{noise}}/B`$ steps to achieve the decrease that one step achieves with an infinite batch size; hence, in order to reach the same loss, we have to train for $`S = (1 + \mathcal{B}_{\text{noise}}/B)S_{\min}`$ steps.

Since the batch size is $`B`$, it is easy to derive that the total amount of data consumed in training is $`E = BS = (B + \mathcal{B}_{\text{noise}})S_{\min}`$. From this result we can see that after increasing the batch size, in order to achieve the same result, we still need to increase the amount of data $`E`$ appropriately; when $`B\to 0`$, the amount of data required is minimal, namely $`E_{\min} = \mathcal{B}_{\text{noise}}S_{\min}`$. Using this notation, we can write
``` math
\begin{equation}\left(\frac{S}{S_{\min}} - 1\right)\left(\frac{E}{E_{\min}} - 1\right) = 1\end{equation}
```
This is the classical relationship between the amount of training data and the number of training steps. It has two parameters, $`S_{\min}`$ and $`E_{\min}`$; we can also experimentally search over multiple $`(S,E)`$ pairs to fit this formula, thereby estimating $`S_{\min}`$ and $`E_{\min}`$, and further estimating $`\mathcal{B}_{\text{noise}} = E_{\min} / S_{\min}`$. For more details of the analysis, please refer back to the previous article [*When the Batch Size Increases, How Should the Learning Rate Change Accordingly?*](./10542-How-Should-the-Learning-Rate-Change-as-the-Batch-Size-Increases.html) or OpenAI’s original paper [*An Empirical Model of Large-Batch Training*](https://papers.cool/arxiv/1812.06162).

# Analysis of the Difficulties

Despite everything written above, we are still confined to SGD. From a computational point of view, SGD is trivial; what is truly complicated is the case where $`\tilde{\boldsymbol{\varphi}}_B`$ depends nonlinearly on $`\tilde{\boldsymbol{g}}_B`$. For example, SignSGD corresponds to $`\tilde{\boldsymbol{\varphi}}_B=\mathop{\mathrm{sign}}(\tilde{\boldsymbol{g}}_B)`$, which is often used in theoretical analyses as an approximation to Adam; a more accurate approximation is SoftSignSGD, which takes $`\epsilon`$ into account, and we attempted to analyze it in [*How Does Adam’s Epsilon Affect the Scaling Law of the Learning Rate?*](./10563-How-Does-Adam-s-Epsilon-Affect-the-Scaling-Law-of-Learning-Rate.html).

In these nonlinear scenarios, computing $`\mathbb{E}[\tilde{\boldsymbol{\varphi}}_B]`$ and $`\mathbb{E}[\tilde{\boldsymbol{\varphi}}_B\tilde{\boldsymbol{\varphi}}_B^{\top}]`$ is often rather difficult, even if we assume the distribution of $`\tilde{\boldsymbol{g}}_B`$ to be a simple normal distribution (note that in the analysis of SGD, we did not need to make any normality assumption about its distribution). For instance, in a previous article, for SignSGD with $`\tilde{\boldsymbol{\varphi}}_B=\mathop{\mathrm{sign}}(\tilde{\boldsymbol{g}}_B)`$, in order to compute $`\mathbb{E}[\tilde{\boldsymbol{\varphi}}_B]`$, we went through the following steps:

> 1.  Assume that the components of $`\tilde{\boldsymbol{g}}_B`$ are mutually independent, so that the problem reduces to the expectation of a single component $`\tilde{\varphi}_B=\mathop{\mathrm{sign}}(\tilde{g}_B)`$ (not in bold);
>
> 2.  Assume that $`\tilde{g}_B`$ (now a scalar) follows a normal distribution; then $`\mathbb{E}[\tilde{\varphi}_B]`$ can be computed, with the answer having to be expressed in terms of the $`\mathop{\mathrm{erf}}`$ function;
>
> 3.  Approximate the $`\mathop{\mathrm{erf}}`$ function by a function of the form $`x/\sqrt{x^2+c}`$ to simplify the result.

That is to say, we have to go through a pile of roundabout steps before barely obtaining an approximate result that can still be analyzed further (this process first appeared in the paper by Tencent, [*Surge Phenomenon in Optimal Learning Rate and Batch Size Scaling*](https://papers.cool/arxiv/2405.14578)), and even this counts as the simple case, because SoftSignSGD is even more complicated:

> 1.  Assume that the components of $`\tilde{\boldsymbol{g}}_B`$ are mutually independent, so that the problem reduces to the expectation of a single component $`\tilde{\varphi}_B=\mathop{\mathrm{softsign}}(\tilde{g}_B, \epsilon)`$;
>
> 2.  Approximate the $`\mathop{\mathrm{softsign}}`$ function by a piecewise linear function, so that the integral below can be computed;
>
> 3.  Assume that $`\tilde{g}_B`$ follows a normal distribution; combined with the approximation in step 2, $`\mathbb{E}[\tilde{\varphi}_B]`$ can be computed, and the answer is a complicated function involving $`\mathop{\mathrm{erf}}`$;
>
> 4.  Approximate the complicated function by a function of the form $`x/\sqrt{x^2+c}`$ to simplify the result.

And things do not end there. With so much effort and so many assumptions, we have only barely computed $`\mathbb{E}[\tilde{\boldsymbol{\varphi}}_B]`$; we still have to compute $`\mathbb{E}[\tilde{\boldsymbol{\varphi}}_B\tilde{\boldsymbol{\varphi}}_B^{\top}]`$, which is often even more complicated (SignSGD is an exception, since $`\mathop{\mathrm{sign}}(x)^2`$ is always 1, it actually becomes simpler). However, the computational complexity is only a secondary issue; the main problem is that these steps do not seem to exhibit any pattern that could be generalized—it appears that each problem can only be analyzed case by case, which is truly mentally exhausting.

# To Be Continued

In order to prevent the article from becoming too long, we stop here for now; this article has mainly provided a brief review of the existing analytical results and the computational difficulties. In the next article, the author will introduce some attempts that were made to reduce the mental burden during the derivation process.

<span style="color: FF8800">***When republishing, please include the address of this article:*** <a href="./11260-Rethinking-Learning-Rate-and-Batch-Size-I-Current-Status.html" class="uri">https://kexue.fm/archives/11260</a></span>

<span style="color: FF8800">***For more detailed republishing matters, please refer to:*** [*Scientific Spaces FAQ*](./06508-Scientific-Space-Browsing-Guide-FAQ.html)</span>
