Links indicate relevance, not agreement. How to use this site →
A zeroth-order optimization method that trains transformer language models by perturbing activations at each token without backpropagation, achieving competitive performance with backprop while being orders of magnitude more efficient than weight-space evolutionary strategies.