The models always used internal reasoning; predicting the next token was simply the output method. Some people experimented to see if the models could forecast 2 or 3 tokens ahead in the output, and they can.
The models always used internal reasoning; predicting the next token was simply the output method. Some people experimented to see if the models could forecast 2 or 3 tokens ahead in the output, and they can.