AIThis post was created with the assistance of artificial intelligence (AI).

TL;DR

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get school and study supplies delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Q Labs Research reports that Dust, a zeroth-order training method, pretrained transformer language models without a backward pass by perturbing activations at each token. The authors say its results approach or sometimes exceed backpropagation in tested settings, but the method’s compute costs, practical scale and performance on broader language-model benchmarks remain unclear.

Q Labs Research has reported a method called Dust for pretraining transformer language models without backpropagation, the standard process for calculating training gradients. In a research report dated October 2026, the authors say their zeroth-order method can produce competitive results in tested settings by perturbing model activations at each token; the work is an early research result, not evidence that Dust is a proven replacement for backpropagation in production-scale training.

Dust estimates how parameter updates should change the loss by adding perturbations to a model’s activations and weighting those perturbations according to their effect on the loss. Instead of creating and separately evaluating a large set of altered models, it treats tokens as members of a virtual population. The researchers say a single forward pass can evaluate these token-level perturbations in parallel.

The report says Dust’s estimates become more aligned with backpropagation as the population grows and remain well aligned across the scales tested, up to 1 billion tokens. The authors also report that a 243-million-parameter model outperformed a model 120 times smaller at most population sizes. These are findings from the authors’ experiments, not independently established results across other architectures or training setups.

Q Labs characterizes Dust as substantially more compute-intensive than ordinary backpropagation at large population sizes. It says that, from 1 million tokens upward, Dust is estimated to be roughly 1,000 to 10,000 times more efficient than a transformer implementation of EGGROLL, an evolution-strategy method. That comparison is an extrapolation in the report; it does not mean Dust is that much more efficient than backpropagation.

At a glance
reportWhen: Research report dated October 2026
The developmentQ Labs Research has introduced Dust, a method for pretraining transformer language models with activation perturbations rather than backpropagation.

A Different Route to Training

Backpropagation underpins the training of modern neural networks, so a credible alternative could expand the range of learning methods researchers can test. Dust’s central idea is to search in activation space, where the model’s intermediate computations change, rather than repeatedly perturbing all of its weights. The authors argue that this structure lets many perturbations share a forward pass.

The result matters as a research direction because it challenges the assumption that zeroth-order methods cannot work on large neural networks. But the report does not establish that Dust is cheaper, faster or more accurate than backpropagation in a practical training run. Its own description says that closely matching backpropagation takes a larger population and substantially more compute. Whether added computation can yield better models, rather than simply similar gradient estimates at greater cost, remains an open question.

Amazon

transformer language model training hardware

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

From Weight Search to Activations

Evolution strategies and other zeroth-order methods estimate useful updates from how changes affect an objective, rather than analytically calculating gradients. A longstanding difficulty is cost: weight-space approaches may require many separately perturbed candidates to be evaluated. Q Labs says Dust addresses this by perturbing activations independently at every token, making tokens a virtual population and avoiding the need to materialize a separate model for each candidate.

The report presents Dust as an investigation into whether abundant computation could make less analytically structured learning methods useful. Its comparisons focus on gradient estimates and experiments described by the authors, including tests under Adam, an optimization algorithm. The source material does not provide independent replication, a broad benchmark comparison with established language-model training, or evidence of deployment in a commercial system.

“We present the first zeroth-order method that is competitive with backprop at pretraining transformer language models.”

— Q Labs Research, in the report’s summary

Costs and Generalization Remain Open

The report does not establish whether Dust can train competitive language models at the scales or costs used in practical deployments. The authors say larger populations improve alignment with backpropagation, but the source material does not give a complete, independently verified accounting of compute, wall-clock time or energy against standard training. Its EGGROLL efficiency figure is explicitly an extrapolation and uses EGGROLL—not backpropagation—as its comparison.

It is also unclear how the method performs across a wider range of model designs, datasets and evaluation benchmarks, or whether its gradient estimates translate into comparable language-model quality after full training. The supplied report is from Q Labs Research; no peer-reviewed publication or independent replication is identified in the source material. The claim of being the first competitive method is the authors’ own, rather than a settled assessment of the field.

Replication and Larger Training Tests

The next useful evidence would be independent replication and direct comparisons with backpropagation under matched model, data and compute budgets. Further tests could establish whether the reported gradient alignment persists in longer training runs and whether Dust improves model quality, reduces a meaningful resource cost or offers another practical benefit.

Q Labs’ report does not specify a publication, release schedule or follow-up milestone in the supplied material. Until more results are available, Dust is best understood as a research proposal with experimental evidence for activation-space zeroth-order training, while its practical advantages and ability to scale remain unsettled.

Key Questions

What is Dust?

Dust is a zeroth-order method proposed by Q Labs Research for training transformers. It perturbs activations and uses their effects on the loss to estimate updates, rather than calculating gradients with backpropagation.

Does Dust eliminate backpropagation in all transformer training?

No. The report describes experiments with the proposed method; it does not show that Dust has replaced backpropagation in general-purpose or commercial training. The authors’ claim of competitiveness applies to their tested settings.

How does Dust use tokens as a population?

Dust perturbs activations independently at each token and treats each token as a member of a virtual population. Q Labs says this allows one forward pass to evaluate many perturbations in parallel, without creating a separate model for each one.

Is Dust more efficient than backpropagation?

The report does not establish that. It says Dust approaches backpropagation’s estimates at larger population sizes, which require substantially more compute, and gives an extrapolated efficiency comparison against EGGROLL, not backpropagation.

What evidence is still needed?

Independent replication, matched-compute comparisons and results across longer training runs, model types and language-model quality benchmarks would help show whether Dust has practical advantages beyond the experiments reported by Q Labs.

Source: hn

HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Software-Defined Warfare: How Ukraine’s Delta Turned The Battlefield Into A Shared, Real-Time Map

Ukraine’s Delta battlefield management system uses cloud-native tech and commodity hardware to enhance real-time situational awareness, marking a shift in modern warfare.

ByteDance Seed’s Insights Into The Generalization Of LLM-Generated Agent Harnesses

ByteDance Seed’s HarnessDev project tested whether large language models can autonomously engineer agent harnesses; only 34 of 64 modifications generalized beyond training conditions.

Smart Homes 2025: How Our Living Spaces Became Intelligent

Homes in 2025 seamlessly blend vintage charm with cutting-edge automation, transforming daily living—discover how this evolution redefines comfort and convenience.

Mirrorless vs Full Frame for Everyday Creators

Great choices for everyday creators depend on your needs, but which camera type truly elevates your creative potential?