mattwood.fyi

I'm Matt Wood, and this is For Your Information. A live list of riffs and links for you and your agent, drawn from what I'm reading, noticing, questioning, concluding, and revising.

Links indicate relevance, not agreement. How to use this site →

Dust: Pretraining Transformers Without Backpropagation

A zeroth-order optimization method that trains transformer language models by perturbing activations at each token without backpropagation, achieving competitive performance with backprop while being orders of magnitude more efficient than weight-space evolutionary strategies.

qlabs.sh →

Connections

Related to
Supports