Dust: Pretraining Transformers Without Backpropagation
AIThis post was created with the assistance of artificial intelligence (AI).

TL;DR

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get school and study supplies delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Q Labs Research describes Dust, a zeroth-order method that estimates learning updates by perturbing transformer activations rather than calculating gradients through backpropagation. The researchers report competitive results in small-scale pretraining tests and say larger populations improve alignment with backprop, but the evidence is not a demonstration of large language model training at production scale.

Q Labs Research has published a report on Dust, a method for training transformer language models without the usual backpropagation step. In experiments described by the researchers, Dust perturbs model activations and uses the resulting loss changes to estimate updates; they say it can compete with backprop in some small-scale pretraining settings, though the report does not establish that it can replace backprop for large-scale language models.

Dust is a zeroth-order optimization method: rather than calculate derivatives of the loss through the network, it changes activations and observes how those changes affect the loss. The report says perturbations are applied independently at each token. That lets the researchers treat tokens as members of a virtual population and evaluate their contributions in parallel in a forward pass, rather than creating and running a separate model for each member.

The authors report that Dust’s estimates become more aligned with backpropagation as the population grows. They also say that, in multiple tested settings, Dust matched or exceeded backprop. Those findings are claims from the Q Labs report, not evidence that the method has been independently validated across model families or training conditions.

The report’s efficiency comparison is also an extrapolation. Q Labs estimates that from one million tokens upward, Dust is roughly 1,000 to 10,000 times more efficient than a transformer implementation of EGGROLL, a weight-space evolution-strategy method. The report does not present this figure as a universal measured advantage across all workloads; it depends on the authors’ extrapolations and the comparison setup.

At a glance
reportWhen: Report dated October 2026
The developmentQ Labs Research has published a report describing Dust, an activation-perturbation method tested for pretraining small transformer language models without a backward pass.

A Different Route to Training

Backpropagation is central to modern neural-network training because it computes gradients that indicate how model parameters should change to reduce loss. Dust explores whether a search-based method can provide useful learning updates without that backward calculation. If activation perturbations can be scaled efficiently, they could broaden the set of training approaches researchers can test and potentially change how compute is used.

For now, the significance is primarily research-oriented. The reported ability to handle a virtual population within a forward pass addresses a cost associated with evolution strategies, which typically require multiple candidate models to be evaluated. But Dust’s claimed competitiveness is not a result showing that it is faster, cheaper, or more effective for training frontier-scale systems. The report itself describes larger populations as requiring substantially more compute.

Amazon

transformer model training hardware

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

From Weight Search to Activations

Many zeroth-order and evolution-strategy approaches search by perturbing model weights, then comparing the resulting losses. In large networks, evaluating many weight-perturbed candidates can be costly. Dust instead perturbs activations inside the network, aiming to represent many perturbations through token-level computations in one forward pass.

Q Labs frames the work against the prevailing reliance on backpropagation, while acknowledging its efficiency in low-compute settings. Its central proposal is not that derivatives have been shown to be unnecessary in all training, but that search-based credit assignment may become more attractive as available compute grows. The report’s evidence concerns the experiments and comparisons it describes, not a completed shift in standard training practice.

“We present the first zeroth-order method that is competitive with backprop at pretraining transformer language models.”

— Q Labs Research, in the report’s TL;DR

Amazon

neural network activation perturbation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Limits of the Reported Evidence

The report’s conclusions remain bounded by the experiments it describes. It is not yet clear from the supplied material whether Dust has been independently reproduced, how performance compares across a broad range of datasets and model architectures, or whether the method remains practical at the scale of current large language model training.

The claimed efficiency advantage over EGGROLL is based on extrapolations, and the report says that larger populations use substantially more compute. The material also does not establish whether Dust’s reported results translate into better language-model quality, lower total training costs, or advantages after accounting for tuning and implementation. Those questions require direct, clearly specified comparisons.

Amazon

machine learning research books

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Replication and Larger-Scale Tests

The next useful evidence would be independent replication and controlled comparisons that report model size, data, compute, population size, and final training outcomes. Such tests could clarify whether Dust’s gradient estimates remain useful as workloads grow and whether the extra population compute is offset by its activation-based design.

Q Labs lists code alongside the report, which may allow other researchers to inspect or reproduce the method. The report does not provide a confirmed timetable for further experiments or an independent evaluation, so the scale and timing of follow-up work remain unknown.

Amazon

AI model training optimization

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What is Dust?

Dust is a zeroth-order training method from Q Labs Research. It perturbs transformer activations and uses changes in loss to estimate updates without the standard backpropagation pass.

Does Dust eliminate backpropagation for large language models?

No such conclusion is established by the report. Q Labs describes experiments it says are competitive with backprop in some settings, but the material does not show that Dust can replace backprop for production-scale or frontier-scale training.

How does Dust differ from evolution strategies?

Many evolution strategies perturb model weights and evaluate separate candidates. Dust perturbs activations at each token and treats tokens as a virtual population that can be evaluated in parallel during a forward pass.

What does the efficiency comparison with EGGROLL mean?

Q Labs estimates Dust is about 1,000 to 10,000 times more efficient than a transformer implementation of EGGROLL from one million tokens upward. The report characterizes this as an extrapolation, not a universal measured result.

What evidence is still needed?

Independent replication, broader tests across data and architectures, and direct measurements at larger scales would help show whether Dust’s reported results hold beyond the experiments described by its authors.

Source: hn

HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Spatial Focus Room: Make Distraction Impossible

A new deep-work app for Apple Vision Pro, Spatial Focus Room, removes distractions by immersing users in focused environments, redefining productivity tools.

Four Questions On Electricity Access For US Data Centers

Rymvard published illustrative scenarios showing how grid delays, curtailment, heat and tariffs can limit usable data center capacity.

M 4.9 – Southern Mid-Atlantic Ridge

A magnitude 4.9 earthquake occurred on the southern Mid-Atlantic Ridge, confirmed by USGS. No immediate reports of damage or injuries.

God Help Us, Let’s Try To Learn About Mechanistic Interpretability Techniques

Search interest in mechanistic interpretability techniques is surging amid growing concerns over AI transparency, though specific developments remain unconfirmed.