Links indicate relevance, not agreement. How to use this site →
LLMs are commonly described as next-token predictors, but this framing is incomplete and misses crucial differences between pre-training (learning from existing data) and post-training with reinforcement learning (learning from model-generated sequences and rewards). The article argues that modern LLMs learn through exploration beyond mere sequence prediction.