Words, Tokens, and Predicting What Comes Next
Under the hood, a language model does not read words the way you do. It chops text into small pieces called tokens, which can be a whole word, part of a word, or even just punctuation. Then it does one job over and over: given everything so far, what token is most likely to come next?
That is the whole loop. It picks a likely next token, adds it to the text, and asks the same question again. A full paragraph is just thousands of those small guesses chained together. The reason it feels like conversation is that the patterns it learned include how conversations go.
There is a bit of controlled randomness in the picking, which is why you can ask the same question twice and get two different answers. Neither one is the 'real' answer. They are both plausible continuations.
Takeaway: if an answer is not what you wanted, rephrasing or regenerating is not a trick, it is simply asking for a different set of likely continuations.