Dust:無需反向傳播的 Transformer 預訓練
Dust: Pretraining Transformers Without Backpropagation
Dust 是一種利用啟用空間零階最佳化的預訓練方法,能在不使用反向傳播的情況下訓練 Transformer。
- 效能:在 1M 至 20M tokens 的實驗中,Dust 在相同計算量下可與或優於傳統反向傳播,且在 10M tokens 時隨 population 增大,效能趨近於 backprop。
Dust 以啟用空間搜尋替代反向傳播,展示在大規模算力下可超越傳統梯度訓練,提示可探索更廣泛的訓練範式。
譯文尚未完整,完整內容請切換至原文。
2026年10月聯絡方式 s@qlabs.sh·程式碼·

TL;DR
- 我們提出第一個零階方法,在預訓練 transformer 語言模型時能與 backprop 競爭。Dust 在每個 token 上獨立擾動啟用(節點擾動),因此每個 token 是一個 虛擬 群體成員,一次前向傳播即可並行評估所有 token。
- Dust 在大族群(即大量計算)下與 backprop 的近似度很高,且在多種設定下甚至 超過 backprop。這暗示在計算資源充足的情況下,我們或能超越 backprop。
- Dust 的效率比權重空間 ES 高出數個數量級。從 1M token 起,Dust 的效率比基於我們的外推,EGGROLL(最先進的 ES 方法)的 transformer 實作高出大約 $10^3$ 到 $10^4$ 倍。
- 零階方法普遍被認為無法擴充套件至大型網路。驚人的是,我們發現較大的模型在族群效率上更佳,而非較差:一個 243M 引數模型在大多數族群大小下,表現優於小 120 倍的模型。
- 隨著族群增大,Dust 的梯度估計與 backprop 更貼近,且在我們測試的每個規模(最高到 1B token)仍保持良好對齊,這對於擴充套件性是令人鼓舞的。
目錄
TL;DR 1 簡介 2 方法 2.1 啟用空間擾動 2.2 信用分配 2.3 幹擾與調整 3 無反向傳播的預訓練 3.1 設定 3.2 主要結果 3.3 Adam 下的塵埃 4 高維空間搜尋 4.1 過引數化 4.2 類似反向傳播梯度的出現 5 結論 6 相關工作 參考文獻 附錄
1 導言
深度學習一直以 backprop 為核心,這是唯一能夠訓練現代神經網路(包括基於 transformer 的語言模型)的責任分配演算法。backprop 需要可微分性並產生一階梯度,深度學習的架構、最佳化器與硬體都在此限制下共同演化。
然而,隨著全球可用計算資源的增加,我們可能更偏好基於搜尋的更通用且蠻力的學習演算法,而非像可微性、反向傳播及高階梯度近似等歸納偏差。苦澀教訓((Sutton, 2019Richard S. Sutton. The bitter lesson. http://www.incompleteideas.net/IncIdeas/BitterLesson.html, 2019. 部落格文章.))指出,隨著計算資源增長而擴充套件的通用方法最終會勝出,而 AlphaGo Zero((Silver et al., 2017David Silver, Julian Schrittwieser, Karen Simonyan, Ioannis Antonoglou, Aja Huang, Arthur Guez, Thomas Hubert, Lucas Baker, Matthew Lai, Adrian Bolton, Yutian Chen, Timothy Lillicrap, Fan Hui, Laurent Sifre, George van den Driessche, Thore Graepel, and Demis Hassabis. 在沒有人工知識的情況下掌握圍棋遊戲. Nature, 550 (7676): 354–359, 2017. doi: 10.1038/nature24270.))是明顯的例子。 在 AlphaGo 上以人類資料進行自我啟動最初能加速網路學習,但隨著大量計算,純粹自我對弈的網路最終超越了它。 同樣地,可微性與反向傳播在低計算環境下可能是良好的歸納偏差,因為它們能提高學習效率,但在高計算環境下卻限制了可行架構的空間。 即使在同一架構內,基於梯度的方法也無法最佳探索損失景觀((Liu et al., 2020Shengchao Liu, Dimitris Papailiopoulos, and Dimitris Achlioptas. 存在壞的全域性最小值且SGD可以到達它們. 在Advances in Neural Information Processing Systems, volume 33, 2020.))。 這也可能解釋為何目前的神經網路需要大量資料才能泛化。 基於搜尋的更靈活信用分配演算法可能是邁向更好泛化的重要一步。
In this paper, we aim to replace backprop with a learning algorithm based much more on brute-force computation and much less on analytic structure. We call it Dust. Dust is a zeroth-order optimization algorithm that perturbs activations, rewards each perturbation by how much it lowers the loss, and averages the reward-weighted perturbations over a population to estimate the gradient. Traditional ES methods that perturb weights(Salimans 等人,2017Tim Salimans、Jonathan Ho、Xi Chen、Szymon Sidor 與 Ilya Sutskever。演化策略作為可擴充套件的強化學習替代方案。arXiv 預印本 arXiv:1703.03864,2017。), like EGGROLL(Sarkar 等人,2025Bidipta Sarkar、Mattie Fellows、Juan Agustin Duque、Alistair Letcher、Antonio León Villares、Anya Sims、Clarisse Wibault、Dmitry Samsonov、Dylan Cope、Jarek Liesen、Kang Li、Lukas Seier、Theo Wolf、Uljad Berdica、Valentin Mohl、Alexander David Goldie、Aaron Courville、Karin Sevegnani、Shimon Whiteson 與 Jakob Nicolaus Foerster。演化策略在超級規模。arXiv 預印本 arXiv:2511.16652,2025。), scale with population, but scaling the population is costly because each member must be materialized and evaluated. We remove both costs with the concept of 虛擬族群, where we avoid materializing every member by bypassing weight space entirely and instead perturb activations, as in node perturbation(Werfel 等人,2003Justin Werfel、Xiaohui Xie 與 H. Sebastian Seung。線性前饋網路中隨機梯度下降的學習曲線。收錄於神經資訊處理系統進展,第 16 卷,2003。; Widrow 與 Lehr,1990Bernard Widrow 與 Michael A. Lehr。30 年的自適應神經網路:感知機、Madaline 與反向傳播。IEEE 會議錄,78(9):1415–1442,1990。doi:10.1109/5.58323。). We do so independently at every token, so each token is a member and one forward pass evaluates them all in parallel.
啟用值是一個比權重更有趣的搜尋空間。機械化可解釋性已證明,推理(無論是否可口頭表述)存在於啟用值中(Gurnee et al., 2026Wes Gurnee, Nicholas Sofroniew, Adam Pearce, Mateusz Piotrowski, Isaac Kauvar, Runjin Chen, Anna Soligo, Paul Bogdan, Euan Ong, Rowan Wang, Ben Thompson, David Abrahams, Subhash Kantamneni, Emmanuel Ameisen, Joshua Batson, and Jack Lindsey. 可口頭表述的表示在語言模型中形成全球工作空間。arXiv 預印本 arXiv:2607.15495, 2026.; Lindsey et al., 2025Jack Lindsey, Wes Gurnee, Emmanuel Ameisen, Brian Chen, Adam Pearce, Nicholas L. Turner, Craig Citro, et al. 大型語言模型的生物學。Transformer Circuits Thread, 2025.),這意味著此方法可以將訓練轉變為對潛在推理的搜尋(Vegesna and Dahal, 2025Akshay Vegesna and Samip Dahal. 在神經網路訓練中分離搜尋與學習。arXiv 預印本 arXiv:2509.10973, 2025.)。我們接著將啟用空間擾動與一條非常通用的信用分配規則配對,該規則為 transformer 塊中的不同層型別分配不同的 token 層級獎勵。這兩個偏差,以及一些實作細節和效率措施,例如避免擾動模組之間的幹擾,構成整個演算法。
我們的貢獻如下。
- 我們提出第一個零階方法,在預訓練 transformer 語言模型時能與反向傳播競爭。在大規模族群下,Dust 在多個設定中超越反向傳播,這暗示在計算資源豐富的情況下,我們或許能超越反向傳播。
- Dust 的效率比權重空間 ES 高出數個量級。從 1M 代幣起,根據我們的推估,Dust 的效率比 EGGROLL 的 transformer 實作高出 $10^3$ 到 $10^4$ 倍。
- 與傳統觀念相反,較大的模型往往更具族群效率,而非較低,並且能利用更大的族群。這為過引數化提供了新的觀點,即它是一個更大的搜尋空間,可能具有更佳的幾何結構。
- 隨著族群增大,Dust 的梯度估計與反向傳播的對齊更好,且在我們測試的每個規模(最高到 1B 代幣)都保持穩定,這對擴充套件性來說頗具鼓舞。
本文的目標是為一種以搜尋為基礎的信用分配演算法奠定基礎,該演算法在我們能想到的最困難任務:預訓練 transformer 時能與反向傳播競爭。我們並未試圖使其在計算效率上足以取代今日的反向傳播,也未訓練其所能開啟的新型神經網路,例如帶有外部程式迴圈的網路,或多步驟迴圈的 transformer,反向傳播時間難以訓練的。這兩項工作留待未來。
2 方法
Dust 的工作原理如下。 我們將高斯噪聲獨立加入每個線性層的輸出,針對每個 token,執行一次前向傳播,並根據該 token 的損失變化為其噪聲賦予獎勵。將這些獎勵加權後的噪聲在多次抽樣後平均,即為該層輸出處的估計誤差,其與層輸入的外積即為權重梯度。注意力內部則使用其變體:它們透過對當前及未來 token 的注意力輸出估計誤差進行加權,而非直接使用 token 的損失。核心直覺是,雖然權重空間 ES 在每次前向傳播中評估一個族群成員,但我們在每個 token 上並行評估一個成員,且成員是通過向隱藏狀態新增噪聲來實現,這是廉價的。在現代 Transformer 中,單次前向傳播即可評估至少 三個量級 大於權重空間 ES 的族群。以下將詳細說明每個元件。
2.1 啟用空間擾動
演化策略的瓶頸是族群大小。每個成員需要自己的擾動權重複本及自己的前向傳播。EGGROLL (Sarkar et al., 2025Bidipta Sarkar, Mattie Fellows, Juan Agustin Duque, Alistair Letcher, Antonio León Villares, Anya Sims, Clarisse Wibault, Dmitry Samsonov, Dylan Cope, Jarek Liesen, Kang Li, Lukas Seier, Theo Wolf, Uljad Berdica, Valentin Mohl, Alexander David Goldie, Aaron Courville, Karin Sevegnani, Shimon Whiteson, and Jakob Nicolaus Foerster. Evolution strategies at the hyperscale. arXiv preprint arXiv:2511.16652, 2025.) 使複本成本低,透過低秩擾動,但每個成員仍然是批次中的一個序列元素,因此族群受限於可負擔的前向傳播數量。我們改為在每個 token 上獨立擾動啟用。於該 token 時,網路表現得好像已對產生啟用層的權重做了低秩擾動,而擾動並未實際體現在權重上。我們稱之為「虛擬族群」。transformer 中的序列有數千個 token,因此一次前向傳播可評估數千個成員,而非一個。模型中的每個權重皆以此方式訓練,唯 $2L$ 殘差混合尺度由普通權重空間 ES 訓練。
在激發值(activations)而非權重上加入雜訊就是節點擾動 (Widrow and Lehr, 1990Bernard Widrow and Michael A. Lehr. 30 years of adaptive neural networks: Perceptron, Madaline, and backpropagation. Proceedings of the IEEE, 78 (9): 1415–1442, 1990. doi: 10.1109/5.58323.),而支援它的常用論點是維度 (Ren et al., 2023Mengye Ren, Simon Kornblith, Renjie Liao, and Geoffrey Hinton. Scaling forward gradient with local losses. In International Conference on Learning Representations, 2023.; Werfel et al., 2003Justin Werfel, Xiaohui Xie, and H. Sebastian Seung. Learning curves for stochastic gradient descent in linear feedforward networks. In Advances in Neural Information Processing Systems, volume 16, 2003.)。某個層的輸出有 $d_{\mathrm{out}}$ 個專案,其權重則有 $d_{\mathrm{out}} \times d_{\mathrm{in}}$,因此激發值雜訊存在於一個小得多的空間中。直觀來看,對於在許多 token 具有大型激發值的 Transformer 來說,維度論點並不成立。單一序列上的雜訊是一個 $T \times d_{\mathrm{out}}$ 張量,一旦 $T \ge d_{\mathrm{in}}$,它所包含的專案數量就會至少和權重矩陣一樣多。然而,透過針對每個 token 獨立的擾動與獎勵,擾動激發值所帶來的沿著 token 軸的新穎且高效率母體,與 EGGROLL 本身已經依賴的批次軸是正交的。
2.2 信用分配
對於線性層 $y_t = W x_t$,我們對所有 token 的輸出進行微調(jitter),$y_t \to y_t + \sigma a_t$(其中 $a_t \sim \mathcal{N}(0, I)$ 且 $\sigma$ 為雜訊規模),並執行前向傳遞。在每個 token $s$ 處,我們計算置中後的損失減少量 $c_s = \tilde{\ell}_s - \ell_s$,其中 $\ell_s$ 是受擾動的損失,$\tilde{\ell}_s$ 是該 token 在同一個批次前向傳遞中一起評估的多個抽樣之間的平均受擾動損失。token $t$ 處微調的獎勵是 $t$ 處的損失減少量,以及在衰減 $\gamma$ 下,微調透過注意力機制也能觸及的其後 token 的損失減少量,
$$r_t = \sum_{s \ge t} \gamma^{\,s-t} c_s .$$
(1)
當 $\gamma = 0$ 時,微調僅由其自身的 token 來獲得獎勵。我們交由第 2.3 節的微調來決定哪些層能夠看到未來的 token。所有 token 的一個獨立微調就是一個抽樣(draw),而一個母體就是 $K$ 個抽樣。經過抽樣平均後,獎勵加權雜訊
$$\hat g_t = -\frac{1}{K\sigma}\sum_{i=1}^{K} r_t^{(i)} a_t^{(i)}$$
(2)
即為該層輸出處的估計誤差,它與層的輸入(前向傳遞已經計算過)的外積在 token 上加總後,即為權重梯度,
$$\widehat{G}_W = \sum_t \hat g_t\, x_t^\top .$$
(3)
反向傳播會使用相同的輸入形成相同的外部積。唯一的區別在於,它是從鏈式法則獲得輸出誤差,而我們則是從母體獲得輸出誤差。對於嵌入層, $x_t$ 是一熱編碼(one-hot),因此外積是將 $\hat g_t$ 散佈加總(scatter-add)到該 token 的列中。
在無限樣本數的情況下,上述估計器即為整個方法,且每一次擾動都可納入一次前向傳播。每一次抽樣的以獎勵加權的噪聲等於梯度加上一個沒有偏好方向的誤差。隨著抽樣次數增加,梯度線性累加,而誤差以平方根方式累加,因而其比例隨樣本數增大而下降,幹擾在極限下消失。
2.3 幹擾與調校
在我們能負擔的樣本數下,主要成本是幹擾,即我們在同一次前向傳播中對許多層和許多 token 進行抖動,因此獎勵單個 token 噪聲的損失變化也會捕捉到該次傳播中其他所有擾動的影響。我們以三種方式減少此類幹擾。首先,將不同層型別分別在獨立的前向傳播中抖動,每個都有自己的噪聲比例,每個區塊也有自己的傳播。這些傳播比完整前向傳播更便宜,因為乾淨的前向已被快取,對區塊 $l$ 的抽樣只重新執行從 $l$ 開始的區塊。其次,注意力內部(查詢、鍵、值、門控、值嵌入)單獨抖動。Token 的損失幾乎不會感知其抖動,因此它們透過注意力輸出獲得獎勵,詳見下文。第三,語言模型頭直接在快取的 logits 上抖動,只重新評估交叉熵以及每次抽樣僅使用一部分詞彙表,這只佔前向傳播的一小部分成本,並允許頭部處理更大的樣本數。
對於注意力內部,我們僅使用抖動重新計算其區塊的注意力輸出,從快取的乾淨啟用值開始。我們使用與估計的注意力輸出梯度對齊的方式來評分抖動,
$$c_s = -\\langle \\hat g_s, \\Delta o_s \\rangle ,$$
(4)
其中 $\\\\Delta o_s$ 是抖動對 token $s$ 的注意力輸出所造成的變化,$\\\\hat g_s$ 則是根據方程式 2 推估的該輸出的誤差。獎勵則是方程式 1,以這些分數取代損失減少,並且按每個 head 計算。
超引數、每層型別的噪音尺度、注意力內部的信用衰減以及每層所分配的人口比例,皆可透過兩種方式調整。其一是使用網格搜尋,在每個設定下以少量 token 進行訓練,並保留能降低損失最大的設定。此方法可靠但成本高昂。其二是使用網格搜尋,最大化我們估計值與單一批次反向梯度之間的餘弦相似度,且不需要任何訓練。單一批次上較大的餘弦相似度並不一定在訓練後降低損失,因此餘弦相似度用於挑選候選者,實際訓練決定結果。無論哪種方式,調參大多為一次性成本,因為它找到的設定在不同 token 預算和人口規模下大多能通用,唯一例外是最大人口 10M 和 20M tokens 時,較慢的信用衰減以及將抽樣偏向注意力側仍有收益(Appendix F)。因此,這個搜尋回復了方法的一般原則,而非單一執行的設定。正如預期,所有層除了鍵、值、閘門和值嵌入之外,都不需要未來 tokens 的信用,後者會被後續 tokens 閱讀並取得接近 1 的 γ。
3 先行訓練不使用反向傳播
3.1 設定
我們在 FineWeb 上使用 4096-token BPE 分詞器訓練 GPT 風格的 transformer,批次為 16k token(8 個 2048 token 的序列),訓練一個 epoch,使用帶動量的 SGD 並保持學習率不變。基礎模型有 8 層、寬度 512。每種方法都使用相同的協定,並為每個 cell 分配三個種子,並在每個 token 預算和 population 上分別調參。Dust 和 backprop 共用一套動量與學習率網格;我們以相同的 transformer 架構實作 EGGROLL(命名為 EGGROLL-Transformer),並在自己的步長、動量、噪音尺度與 fitness shaping 網格上調參。驗證集與測試集各為 544 個序列的留出集。我們報告最佳驗證點的測試損失。
我們在 Dust 的抽樣中計算 population,在 EGGROLL (Sarkar et al., 2025Bidipta Sarkar, Mattie Fellows, Juan Agustin Duque, Alistair Letcher, Antonio León Villares, Anya Sims, Clarisse Wibault, Dmitry Samsonov, Dylan Cope, Jarek Liesen, Kang Li, Lukas Seier, Theo Wolf, Uljad Berdica, Valentin Mohl, Alexander David Goldie, Aaron Courville, Karin Sevegnani, Shimon Whiteson, and Jakob Nicolaus Foerster. Evolution strategies at the hyperscale. arXiv preprint arXiv:2511.16652, 2025.),我們的權重空間基準。一次抽樣是一次在選定層的所有 token 上進行 jitter,並以 token 損失為報酬,population K 是每次更新的抽樣次數。一次抽樣比一次正向傳播略便宜,因為乾淨的正向傳播已被快取,抽樣 jitter block l 只重新執行從 l 開始的 block。K 排除頭部和注意力內部的抽樣,這些只佔更新 FLOPs 的少量比例(Appendix E)。綜合而言,從 256 的 population 起,Dust 使用的計算量比同樣 population 的 EGGROLL 少,因此比較對 EGGROLL 有利。
3.2 主要結果
我們將 token 預算從 100k 掃描至 20M,群體大小從 64 到 16k,並在每個預算的相同網格上調整倒傳遞(backprop)(Table 1、Figure 2)。在 100k 與 1M tokens 時,Dust 的表現低於倒傳遞,在 100k 時差了幾百次抽樣,在 1M 時差了一千次。在 10M 與 20M tokens 時,與倒傳遞的差距隨群體大小而縮小。在 10M 時,階梯已經平緩,其擬合極限剛好落在倒傳遞上方。在 20M 時,階梯在 16k 次抽樣時仍在下降。其冪次律擬合將極限定在 4.431(95% 區間為 3.89 到 4.58),低於倒傳遞的 4.633,但由於階梯仍在下降,該擬合受到較鬆散的約束,因此我們將其解讀為差距隨群體大小持續縮小的證據,而非測得的極限。
權重空間 ES 的效率低得多。在群體大小大 256 倍的情況下,EGGROLL 在 16k 時仍未達到 Dust 在 64 次抽樣的水準。它在 100k tokens 時差距縮小至 0.02 以內,而在 1M、10M 與 20M 時則保持在高出 0.4 到 0.6 的水準(Figure 2)。EGGROLL 的階梯在 16k 時仍急劇下降,因此隨著群體增加它會持續改善,但要延續該階梯以達到 Dust 最小群體所需的水準,大約需要該群體的數千倍到 $10^4$ 倍(Appendix D)。Figure 1 也顯示了訓練過程中的差距;Dust 在 1k 與 16k 次抽樣下,全程緊貼著倒傳遞的曲線,而每個 EGGROLL 曲線都在早期落後,並在 64 次抽樣的 Dust 上方顯著平緩。
3.3 Adam 下的 Dust
雖然在本文的剩餘部分中我們主要聚焦於 SGD,但我們在所有三種方法上重複了使用 Adam 的 1M token 階梯(Figure 3),並在每個群體大小重新調整倒傳遞的學習率、Dust 自身的超引數以及 EGGROLL 的步階大小、動量與適應度塑造(fitness shaping)。有趣的是,EGGROLL 從 Adam 中幾乎沒有獲得任何好處,其調校後的 Adam 階梯在 Table 1 中每個群體大小都落在其 SGD 階梯的 0.01 範圍內。然而,Adam 同時改善了 Dust 與倒傳遞,並使階梯的形狀與之前相似。Dust 在大型群體下逼近倒傳遞,且其極限的區間位於倒傳遞下方(Figure 3)。因此,即使現代最佳化器是針對倒傳遞梯度所最佳化,Dust 的估計值已經適用於現代最佳化器。我們懷疑最佳化器與 Dust 的共同演進(coevolution)可以帶來進一步的收益,並將此留作未來的工作。
4 高維空間中的搜尋
4.1 過度引數化
每列中的測試損失每列最低每列最高
| 群體 | 引數 | 跨尺寸 | |||
|---|---|---|---|---|---|
| 2.0M | 7.3M | 38M | 243M | ||
| 64 | 5.705 | 5.556 | 5.558 | 5.719 | |
| 256 | 5.486 | 5.362 | 5.358 | 5.419 | |
| 1k | 5.265 | 5.161 | 5.158 | 5.214 | |
| 4k | 5.189 | 5.095 | 5.053 | 5.124 | |
| 16k | 5.171 | 5.065 | 5.036 | 5.086 | |
| 倒傳遞 | 5.180 | 5.066 | 5.015 | 5.048 | |
傳統觀點認為零階方法無法訓練大型網路(Lillicrap et al., 2020Timothy P. Lillicrap, Adam Santoro, Luke Marris, Colin J. Akerman, and Geoffrey Hinton. Backpropagation and the brain. Nature Reviews Neuroscience, 21 (6): 335–346, 2020. doi: 10.1038/s41583-020-0277-3.)。一次前向傳播返回單一標量,因而梯度估計的變異隨擾動維度數量增長,進而需要更多族群才能得到有用更新(Nesterov and Spokoiny, 2017Yurii Nesterov and Vladimir Spokoiny. Random gradient-free minimization of convex functions. Foundations of Computational Mathematics, 17 (2): 527–566, 2017. doi: 10.1007/s10208-015-9296-2.; Werfel et al., 2003Justin Werfel, Xiaohui Xie, and H. Sebastian Seung. Learning curves for stochastic gradient descent in linear feedforward networks. In Advances in Neural Information Processing Systems, volume 16, 2003.)。我們直接測試此觀點,訓練四個規模,2M、7M、38M 及 243M 個引數(引數數量相差 120 倍),在固定 10M tokens 的族群掃描中(圖 4;參見附錄 F 以獲得調參細節)。我們得到兩個顯著觀察,挑戰傳統智慧。
- 較大模型往往更具族群效率,而非更差。 在 256 以上的每個族群中,損失隨引數從 2M 下降到 7M,再到 38M,243M 模型僅略差。即使在我們測試的最小族群中,$120\times$ 大小的模型表現相近,且在其他族群中更優。至某一程度,較大的模型從尺寸中獲得的收益超過了變異所帶來的損失。然而,我們觀察到,與相同族群和模型尺寸的反向傳播相比,差距隨模型尺寸增大而增加,但差距很小,尤其在大族群尺寸下。
- 較大模型隨族群增大仍持續改進,而小模型則趨於平坦。 小模型較早飽和,而大模型則隨著大族群尺寸持續改進。超過 1k 次抽樣後,38M 和 243M 模型比 2M 和 7M 模型提升約 30%,且無論族群多大,微型模型都無法競爭。
因此,正確的思考方式是將模型尺寸視為 搜尋空間的尺寸與幾何。較大的模型擁有更大的搜尋空間,這使它能充分利用大族群,並可能具備更良好的損失地形幾何,這或許是搜尋即使在小族群下仍更有效的原因。
4.2 Backprop-Like Gradients 的出現
我們測量 Dust 的估計值與反向傳播梯度在同一批次中的餘弦相似度,按層型別及每層,在覆蓋兩個量級 tokens 的反向傳播訓練檢查點上,亦即從 10M 到 1B tokens,樣本量從 64 增至 128k 前向傳遞 per step (Figure 5)。餘弦相似度隨樣本量在每個層型別和每個訓練階段上升,且符合兩引數法則
$$\cos(K) = \frac{c_{\max}}{\sqrt{1 + c/K}}$$
(5)
符合每個層型別的 RMSE 低於 0.06,$c_{\max}$ 為上限,$c$ 為層達到 $c_{\max}/\sqrt{2}$ 時的樣本量。實用梯度僅由大樣本量產生,沒有內建鏈式法則,僅需輕微調整超引數以最大化餘弦相似度,詳見 Section 2.3。對於 100M tokens 的檢查點,層型別平均餘弦相似度的擬合 RMSE 為 $0.0032$–$0.0367$(詳見附錄 G)。Dust 的梯度在餘弦相似度上接近反向傳播,但其逼近程度因層型別及層而異。EGGROLL 的估計值,以相同方式在相同前向傳遞次數下測量,隨樣本量增長,但在 128k 時仍低於 0.05 的餘弦相似度,僅頭部型別除外,這可能解釋為何在 Section 3.2 中訓練效果不佳。
重要的是,隨著 tokens 數量增加,餘弦相似度在大多數層保持穩定,這對擴張來說是令人鼓舞的。Section 3 顯示匹配並超越反向傳播所需的樣本量隨 tokens 增長。然而,在大樣本量下,餘弦相似度在兩個量級 tokens 上保持平坦,意味著在足夠大樣本量時,這個需求可能不再隨更多 tokens 增長。更重要的是,Dust 的梯度雖逼近反向傳播但並未完全收斂,這其實是一個好特性。估計值指向相似方向但不是反向傳播梯度,導致不同的最佳化軌跡,在我們的實驗中該軌跡甚至可能優於反向傳播(Section 3.2)。
5 結論
自從 Rumelhart 等人 (1986)David E. Rumelhart, Geoffrey E. Hinton, 和 Ronald J. Williams. 透過反向傳播學習表示。Nature,323: 533–536,1986。 之後,backprop 成為訓練神經網路的演算法,而我們所建構的架構、最佳化器與硬體皆圍繞它。隨著計算資源越來越充裕,我們認為有更優秀的替代方案。於是我們推出 Dust,一種在現有 ES 演算法上大幅提升、並在預訓練 transformer 時極為接近 backprop 的演算法,甚至在大量計算下能超越它。
有許多有趣的未解問題。第一個是 Dust 是否能透過 隱式探索損失景觀,找到比 backprop 的一階梯度更佳的方向,捕捉高階曲率以將搜尋拉向平坦區域。我們有提示它能做到,因為在大族群時它有時會表現得比 backprop 更好,但機制尚不明確。第二個是 Dust 擴大了架構的搜尋空間,因為它不需要網路是端到端可微的,並且在 backprop 已知困難的情況下(如使用時間反向傳播訓練的迴圈或迴圈計算)可能表現更佳。第三個是計算效率,這篇論文並未聚焦於此。我們必須在計算效率上提升數個量級,才能使 Dust 成為目前計算量下 backprop 的實用替代方案。
早期對語言模型零階預訓練的研究 ( (Allaire 等人,2025Nathan Allaire, Mahsa Ghazvini Nejad, Sébastien Le Digabel, 和 Vahid Partovi Nia. 零階最佳化用於語言模型預訓練。在 Proceedings of ICPRAM,頁 113–121,2025。doi: 10.5220/0013261100003905. URL https://doi.org/10.5220/0013261100003905。) 探討了從零開始用權重擾動訓練 transformer 的困難。其後的 KronZO 方法 ( (Allaire 等人,2026Nathan Allaire, Sébastien Le Digabel, Dominique Orban, 和 Vahid Partovi Nia. 零階克羅內克最佳化用於語言模型預訓練。SN Computer Science,7,2026。URL https://www.gerad.ca/en/papers/G-2025-44。Article 162。) 透過帶克羅內克結構的緊湊擾動與選擇性方向更新,在降低記憶體使用的同時提升預訓練效果。EGGROLL ( (Sarkar 等人,2025Bidipta Sarkar, Mattie Fellows, Juan Agustin Duque, Alistair Letcher, Antonio León Villares, Anya Sims, Clarisse Wibault, Dmitry Samsonov, Dylan Cope, Jarek Liesen, Kang Li, Lukas Seier, Theo Wolf, Uljad Berdica, Valentin Mohl, Alexander David Goldie, Aaron Courville, Karin Sevegnani, Shimon Whiteson, 和 Jakob Nicolaus Foerster. Evolution strategies at the hyperscale。arXiv preprint arXiv:2511.16652,2025。) 透過低秩結構,使大量權重擾動在 GPU 上更高效。Dust 則在啟用上搜尋,並為每個 token 使用獎勵,以從每一次前向傳播中提取更多貢獻。
For fine tuning, MeZO (Malladi et al., 2023Sadhika Malladi、Tianyu Gao、Eshaan Nichani、Alex Damian、Jason D. Lee、Danqi Chen 與 Sanjeev Arora。只用前向傳播進行語言模型微調。在 Advances in Neural Information Processing Systems,第36卷,2023年。)顯示語言模型僅用前向傳播即可調整,記憶體使用接近推論。Evolution Strategies at Scale (Qiu et al., 2026Xin Qiu、Yulu Gan、Conor F. Hayes、Qiyao Liang、Yinggan Xu、Roberto Dailey、Elliot Meyerson、Babak Hodjat 與 Risto Miikkulainen。大規模演化策略:LLM 微調超越強化學習。在 Proceedings of the International Conference on Machine Learning,2026年。URL https://arxiv.org/abs/2509.24372。arXiv:2509.24372。)展示了在擁有數十億引數的語言模型中,使用 ES 對所有引數進行微調。Neural Thickets (Gan and Isola, 2026Yulu Gan 與 Phillip Isola。Neural thickets:多樣化任務專家聚集於預訓練權重周圍。在 arXiv preprint arXiv:2603.12228,2026年。URL https://arxiv.org/abs/2603.12228。)通過隨機擾動預訓練權重、挑選最佳候選者並整合其預測,找到有用的任務專家。這些結果顯示搜尋在預訓練模型周圍可以達成的程度。我們的實驗針對從零開始預訓練學習表示本身。
Our activation perturbations build on node perturbation (Werfel et al., 2003Justin Werfel, Xiaohui Xie, and H. Sebastian Seung. Learning curves for stochastic gradient descent in linear feedforward networks. In Advances in Neural Information Processing Systems, volume 16, 2003.). GEMINI (Le Cun et al., 1988Yann Le Cun, Conrad C. Galland, and Geoffrey E. Hinton. GEMINI: Gradient estimation through matrix inversion after noise injection. In Advances in Neural Information Processing Systems, volume 1, pages 141–148, 1988. URL https://papers.neurips.cc/paper_files/paper/1988/file/a0a080f42e6f13b3a2df133f073095dd-Paper.pdf.) injected noise into the first hidden layer and recovered layerwise gradient estimates through iterative matrix inversion. Zoop (Hu et al., 2025Xixi Hu, Bo Liu, Qiang Liu, Xiaocong Du, Bhargav Bhushanam, Louis Feng, Chengyue Gong, and Kaizhao Liang. Zoop it! Efficient zero-order optimization with output perturbation. In ICML Workshop on Tiny Titans: The next wave of On-Device Learning for Foundation Models, 2025. URL https://openreview.net/forum?id=Tc8vFyRhPO.) uses output perturbations for language model fine tuning, converting estimated output gradients into parameter updates with local derivatives. Scaling Forward Gradient with Local Losses (Ren et al., 2023Mengye Ren, Simon Kornblith, Renjie Liao, and Geoffrey Hinton. Scaling forward gradient with local losses. In International Conference on Learning Representations, 2023.) combines activation perturbations and local losses with forward mode automatic differentiation to reduce estimator variance. Forward gradients with multiple tangents (Flügel et al., 2025Katharina Flügel, Daniel Coquelin, Marie Weiel, Charlotte Debus, Achim Streit, and Markus Götz. Beyond backpropagation: Optimization with multi-tangent forward gradients. In International Joint Conference on Neural Networks, 2025. URL https://arxiv.org/pdf/2410.17764v2.) also use forward mode differentiation, combining multiple directional derivatives through orthogonal projection to improve gradient estimates.
A separate line of work replaces the global backward pass with local learning dynamics: Sakana AI’s PC-ALM (Seely and Gould, 2026Jeffrey Seely 與 Julian Gould。Augmented Lagrangian predictive coding。arXiv preprint arXiv:2605.31022,2026。URL https://arxiv.org/abs/2605.31022.) propagates credit through local predictive coding dynamics and Lagrange multipliers. It uses local derivatives and is evaluated on image classification tasks. Dust estimates credit from forward perturbations during transformer pretraining, combining rewards for each token with local targets for attention outputs.
參考文獻
Nathan Allaire、Mahsa Ghazvini Nejad、Sébastien Le Digabel 與 Vahid Partovi Nia。Zeroth order optimization for pretraining language models。In Proceedings of ICPRAM,第 113–121 頁,2025 年。doi: 10.5220/0013261100003905。URL https://doi.org/10.5220/0013261100003905。
Nathan Allaire, Sébastien Le Digabel, Dominique Orban, and Vahid Partovi Nia. Zeroth-order Kronecker optimization for pretraining language models. SN電腦科學, 7, 2026. URL https://www.gerad.ca/en/papers/G-2025-44. Article 162.
Katharina Flügel, Daniel Coquelin, Marie Weiel, Charlotte Debus, Achim Streit, and Markus Götz. Beyond backpropagation: Optimization with multi-tangent forward gradients. In 國際神經網路聯合會議, 2025. URL https://arxiv.org/pdf/2410.17764v2.
Yulu Gan and Phillip Isola. Neural thickets: Diverse task experts are dense around pretrained weights. arXiv 預印本 arXiv:2603.12228, 2026. URL https://arxiv.org/abs/2603.12228.
Wes Gurnee, Nicholas Sofroniew, Adam Pearce, Mateusz Piotrowski, Isaac Kauvar, Runjin Chen, Anna Soligo, Paul Bogdan, Euan Ong, Rowan Wang, Ben Thompson, David Abrahams, Subhash Kantamneni, Emmanuel Ameisen, Joshua Batson, and Jack Lindsey. Verbalizable representations form a global workspace in language models. arXiv 預印本 arXiv:2607.15495, 2026.
Xixi Hu, Bo Liu, Qiang Liu, Xiaocong Du, Bhargav Bhushanam, Louis Feng, Chengyue Gong, and Kaizhao Liang. Zoop it! Efficient zero-order optimization with output perturbation. In ICML 小巨人工作坊:基礎模型在裝置端學習的下一波, 2025. URL https://openreview.net/forum?id=Tc8vFyRhPO.
Yann Le Cun, Conrad C. Galland, and Geoffrey E. Hinton. GEMINI: Gradient estimation through matrix inversion after noise injection. In 神經資訊處理系統進展, volume 1, pages 141–148, 1988. URL https://papers.neurips.cc/paper_files/paper/1988/file/a0a080f42e6f13b3a2df133f073095dd-Paper.pdf.
Timothy P. Lillicrap, Adam Santoro, Luke Marris, Colin J. Akerman, and Geoffrey Hinton. Backpropagation and the brain. 自然評論神經科學, 21 (6): 335–346, 2020. doi: 10.1038/s41583-020-0277-3.
Jack Lindsey, Wes Gurnee, Emmanuel Ameisen, Brian Chen, Adam Pearce, Nicholas L. Turner, Craig Citro, et al. On the biology of a large language model. Transformer 電路線索, 2025.
Shengchao Liu, Dimitris Papailiopoulos, and Dimitris Achlioptas. Bad global minima exist and SGD can reach them. In 神經資訊處理系統進展, volume 33, 2020.
Sadhika Malladi, Tianyu Gao, Eshaan Nichani, Alex Damian, Jason D. Lee, Danqi Chen, and Sanjeev Arora. Fine-tuning language models with just forward passes. In 神經資訊處理系統進展, volume 36, 2023.
Yurii Nesterov and Vladimir Spokoiny. Random gradient-free minimization of convex functions. 計算數學基礎, 17 (2): 527–566, 2017. doi: 10.1007/s10208-015-9296-2.
Xin Qiu, Yulu Gan, Conor F. Hayes, Qiyao Liang, Yinggan Xu, Roberto Dailey, Elliot Meyerson, Babak Hodjat, and Risto Miikkulainen. Evolution strategies at scale: LLM fine-tuning beyond reinforcement learning. In 機器學習國際會議論文集, 2026. URL https://arxiv.org/abs/2509.24372. arXiv:2509.24372.
Mengye Ren, Simon Kornblith, Renjie Liao, and Geoffrey Hinton. Scaling forward gradient with local losses. In 學習表示國際會議, 2023.
David E. Rumelhart、Geoffrey E. Hinton 與 Ronald J. Williams。透過反向傳播誤差學習表示。Nature,323:533–536,1986。
Tim Salimans、Jonathan Ho、Xi Chen、Szymon Sidor 與 Ilya Sutskever。演化策略作為可擴充套件的強化學習替代方案。arXiv preprint arXiv:1703.03864,2017。
Bidipta Sarkar、Mattie Fellows、Juan Agustin Duque、Alistair Letcher、Antonio León Villares、Anya Sims、Clarisse Wibault、Dmitry Samsonov、Dylan Cope、Jarek Liesen、Kang Li、Lukas Seier、Theo Wolf、Uljad Berdica、Valentin Mohl、Alexander David Goldie、Aaron Courville、Karin Sevegnani、Shimon Whiteson 與 Jakob Nicolaus Foerster。演化策略在超大規模環境。arXiv preprint arXiv:2511.16652,2025。
Jeffrey Seely 與 Julian Gould。增強拉格朗日預測編碼。arXiv preprint arXiv:2605.31022,2026。URL https://arxiv.org/abs/2605.31022。
David Silver、Julian Schrittwieser、Karen Simonyan、Ioannis Antonoglou、Aja Huang、Arthur Guez、Thomas Hubert、Lucas Baker、Matthew Lai、Adrian Bolton、Yutian Chen、Timothy Lillicrap、Fan Hui、Laurent Sifre、George van den Driessche、Thore Graepel 與 Demis Hassabis。不依賴人類知識掌握圍棋。Nature,550(7676):354–359,2017。doi:10.1038/nature24270。
Richard S. Sutton。苦澀教訓。http://www.incompleteideas.net/IncIdeas/BitterLesson.html,2019。部落格貼文。
Akshay Vegesna 與 Samip Dahal。神經網路訓練中搜尋與學習的分離。arXiv preprint arXiv:2509.10973,2025。
Justin Werfel、Xiaohui Xie 與 H. Sebastian Seung。線性前饋網路中隨機梯度下降的學習曲線。收錄於 Advances in Neural Information Processing Systems,第 16 版,2003。
Bernard Widrow 與 Michael A. Lehr。30 年的自適應神經網路:感知機、Madaline 與反向傳播。Proceedings of the IEEE,78(9):1415–1442,1990。doi:10.1109/5.58323。
A 架構 B 餘留檢查點的餘弦梯子 C 以代幣預算和人群為基礎的測試損失 D EGGROLL 的外推人群 E 人群配置 F 調校 G 餘弦擬合引數
來源:hackernews100 · qlabs.sh