Dust: Pretraining Transformers Without Backpropagation
AIThis post was created with the assistance of artificial intelligence (AI).

TL;DR

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get school and study supplies delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Q Labs Research describes Dust, a zeroth-order method that trains transformer language models by perturbing activations rather than calculating gradients with backpropagation. The researchers report competitive results in experiments and say Dust’s estimates become more aligned with backprop as population size grows, but the method’s compute costs and performance at larger training scales remain important open questions.

Q Labs Research has introduced Dust, a zeroth-order optimization method that it says can pretrain transformer language models without backpropagation. In a report dated October 2026, the researchers describe experiments in which Dust’s performance was competitive with backprop in selected settings, putting forward a different approach to assigning credit for a model’s errors.

Dust perturbs a model’s activations—the intermediate values processed inside the network—rather than changing its weights to create a population of candidate models, as conventional evolution-strategy methods do. The method estimates a direction for learning by measuring how perturbations affect the loss, then averaging perturbations weighted by their rewards. Q Labs says perturbations are applied independently at each token, allowing tokens in a forward pass to act as members of a virtual population.

The report says Dust’s estimates become more closely aligned with backpropagation’s gradients as population size increases. The authors report that this alignment remains strong across the scales they tested, up to 1 billion tokens, and that Dust matches or exceeds backpropagation in multiple experimental settings. Those are findings reported by the researchers; they do not establish that Dust is generally faster or more effective across language-model training.

Q Labs also compares Dust with EGGROLL, a weight-space evolution-strategy method. The report estimates that, from one million tokens upward, Dust is roughly 1,000 to 10,000 times more efficient than a transformer implementation of EGGROLL. That figure is an extrapolation, not a broadly validated measurement across deployed training runs. The authors also say a 243-million-parameter model outperformed a model 120 times smaller at most tested population sizes.

At a glance
reportWhen: Report dated October 2026
The developmentQ Labs Research has published a report describing Dust, a method for pretraining transformers without a backward pass.

A Different Route to Transformer Training

Backpropagation is central to training current deep-learning systems: it computes gradients that tell model parameters how to change in response to errors. A method that can train transformers without that backward calculation could broaden the set of optimization approaches researchers can test, and could matter if available computation grows faster than the efficiency of established methods.

Dust’s proposal is not simply to replace one training recipe with a proven cheaper one. Its reported advantage over EGGROLL concerns the cost of evaluating a population, while its comparison with backprop includes settings in which the method uses larger populations and therefore more computation. The practical question is whether the method’s learning quality justifies that expense at realistic model sizes and training budgets.

The work also challenges a common expectation that zeroth-order methods become less useful as networks grow. Q Labs reports that its larger model was more population-efficient than its smaller comparison model in most tested settings. If that result holds in further experiments, it could motivate more research into search-based learning. The report does not show that Dust can replace backprop in production-scale model development.

Amazon

transformer training hardware

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

How Dust Builds a Virtual Population

Evolution strategies typically search by perturbing model weights, evaluating candidate versions and using their outcomes to guide updates. That can require separate candidate evaluations, making larger populations costly. Dust instead applies perturbations in activation space, the intermediate representations generated as data passes through the network.

Q Labs calls the resulting setup a virtual population: each token receives its own perturbation, and one forward pass processes them together. The report frames this as a way to keep the parallel evaluation benefits of a population without materializing a separate set of model weights for every candidate. Earlier work on node perturbation also changed activations to explore learning without standard gradient calculations; Dust applies that idea to transformer pretraining.

The authors’ broader motivation is that backpropagation’s efficiency has shaped neural-network architectures, optimizers and hardware. They argue that search-based learning may become more appealing if computation becomes abundant. This is a research motivation, not evidence that a shift away from backprop is already underway.

“We present the first zeroth-order method that is competitive with backprop at pretraining transformer language models.”

— Q Labs Research, in its report’s TL;DR

Amazon

AI model training accelerators

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Limits of the Reported Results

The supplied report summary does not provide enough detail to independently evaluate Dust’s experimental setup, compute accounting, benchmark choices or statistical uncertainty. It is not clear how the reported comparisons change across different architectures, datasets, optimizers or training budgets, or whether the reported performance advantage persists in a full-scale language-model training run.

The efficiency comparison with EGGROLL is explicitly based on extrapolations. The report’s results up to one billion tokens are encouraging to its authors, but they do not establish performance at the substantially larger token counts used for major language models. The provided material also does not establish whether the work has undergone peer review or whether independent research groups have replicated the results.

Amazon

GPU for deep learning

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Replication and Larger Training Runs

The next useful tests are independent replications and comparisons that hold model, data and compute budgets constant. Those would help determine whether Dust’s reported alignment with backprop translates into comparable training quality, and what additional compute is required to reach it.

Q Labs’ report links to code, which could let other researchers inspect and test the implementation. No release schedule, independent replication or larger-scale result is specified in the supplied material. Until those details emerge, Dust is best understood as a research result proposing an alternative training method, not a demonstrated replacement for backpropagation.

Amazon

high performance computing for AI

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What is Dust?

Dust is a zeroth-order optimization method proposed by Q Labs Research for training transformers. It uses changes to activations to estimate learning directions rather than relying on backpropagation’s backward pass.

Does Dust eliminate backpropagation?

The report says Dust can train models without a backward pass in the experiments described. It does not show that researchers or developers can broadly replace backpropagation in practical language-model training.

How does Dust create its virtual population?

Dust applies independent perturbations to activations at each token. The researchers say this allows one forward pass to process the tokens as members of a virtual population, rather than evaluating many separately perturbed sets of weights.

How strong is the evidence so far?

Q Labs reports competitive results in selected experiments, gradient alignment with backpropagation at larger population sizes, and tests up to one billion tokens. The supplied material does not establish independent replication, peer-review status or performance in a full-scale production training run.

Is Dust more efficient than backpropagation?

The report’s estimate that Dust is 1,000 to 10,000 times more efficient applies to a comparison with a transformer implementation of EGGROLL, from one million tokens upward, and is based on extrapolation. It is not a general finding that Dust is more efficient than backpropagation.

Source: hn

EVERGREEN BESTSE

Evergreen bestsellers Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Whole House Surge Protection Explained: What It Does (and Doesn’t) Protect

Discover how whole house surge protection shields your home from voltage spikes—and what it doesn’t cover—so you can better safeguard your appliances and safety.

Dual Fuel vs Inverter Generator: What Fits Different Backup Plans

When choosing between dual fuel and inverter generators, understanding their differences helps you select the best option for your backup needs.

What to Do the Minute You Think You’ve Been Hacked

Optimize your response immediately after suspecting a hack to prevent further damage—discover essential steps to protect yourself now.

Electric Bike Classes Explained: Class 1 Vs 2 Vs 3

A comprehensive comparison of Class 1, 2, and 3 electric bikes reveals key differences that can help you choose the perfect ride for your needs.