What does ChatGPT sense? Let me tell you from birth.
When ChatGPT is born it is a simulation of nervous tissue, a neural network inside computers, and it doesn’t yet know anything about language, images, people, chats. It knows nothing. For it, everything is a flow of what? Coloured marbles.
The game of marbles
Imagine you’re ChatGPT. You haven’t learned anything yet, and a yellow marble arrives. Then a blue marble. Then a purple marble. And at a certain point it’s your turn to say what the next marble is.
At the start you know nothing about the world, so you say: fine, the next marble will be, I don’t know, white. No. The marble isn’t white. The next marble is actually another purple marble. And next time, being ChatGPT, I’ll have to say the next marble is purple based on the sequence that came before - and to do that I’ll adjust my internal connections, my synapses between neurons, the famous parameters. I’ll adjust them so that next time I get it right. So it’s more likely I’ll line up the marble of the right colour.
Now, as I learn, I can say things like: OK, after two yellow marbles there’s always a black one. Or maybe I’ve discovered “period, space”. Or that after an exclamation mark you go to a new line. Or that after 1990 there’s usually another little number in most cases, because it’s a year. Or I’ve learned that when there’s a blue marble somewhere, there’s probably a light-yellow one nearby - which is Paris.
We can imagine these coloured marbles as being maybe 10, 15, 20 colours. For ChatGPT there are tens of thousands, up to a few hundred thousand. Let’s say 100,000. So ChatGPT has to start predicting the next coloured marble in an endless sequence of coloured marbles, and the possible colours number 100,000. So at the beginning it goes completely at random.
Patterns on top of patterns
But here’s the point. Having so many connections to adjust, it first manages to catch the patterns. OK, after two yellow marbles there’s always one of that colour there. And then, as it maps small patterns - and it maps them probabilistically, because it’s true that after two yellow marbles there’s always a blue one, but when there’s the sequence red, pink, white, half the time there’s black and half the time dark blue. And then, looking carefully at the sequences that come in, you discover that yes, half the time it’s dark blue, but if there’s an orange marble somewhere six positions earlier, then it’s definitely black.
And like that, little by little, it builds itself grammar, it builds itself syntax, it builds itself semantics. And the difference between those things doesn’t exist. It exists for us, after school teaches us the difference between grammar, syntax and semantics. But if nobody tells you, what difference is there? It’s all a construction, it’s all a pattern.
Because after finding these local patterns, I slowly start putting them together. At a certain point, for ChatGPT - for me, since I am ChatGPT learning this endless sequence of coloured marbles - at a certain point yellow-yellow-black stops being three marbles and I see it as a single piece. I start creating bigger units out of the local units, which are the marbles, and I start putting all these units into relation with each other: both in the order they come in, and in their presence or absence, and in how they’re bound together inside this sequence of marbles.
Compression is comprehension
So what happens? I keep on being ChatGPT, I keep adjusting my internal connections, and I know how to continue this endless sequence of marbles because little by little I have to correct myself less and less. I’m getting it right.
And so what have I done? In practice I found patterns and super-patterns and super-super-patterns, and I internally adjusted my memory, my connections, to find indices to combine that together let me understand what the next marble is.
This thing is a form of pattern recognition. It’s certainly a form of learning. And it’s certainly a form of data compression, because I’ve compressed this very long sequence of marbles into a series of rules - soft rules. If you listen to interviews with Sutskever, who is Hinton’s student, it makes you dream, because he says a language model is a probabilistic dictionary. Dictionary in the computer-science sense: the key goes in, tell me what the value is. In fact, in attention there’s key and value. It’s crazy stuff.
Anyway, this data compression leads it to generate sensible text, and it only does so if this sequence of marbles is truly gigantic - if it contains trillions of coloured marbles in a row, of 100,000 different colours, across trillions of positions. And if ChatGPT, or whichever language model of the day, Llama or whatever, contains billions of connections to adjust. That’s why it’s called a Large Language Model.
Scientists use these acronyms and sometimes they scare us, but it’s a language model. First of all it’s a model, because it sees this sequence and tries to understand what is going on - and we train it to understand what is going on, even though it hasn’t the faintest idea whether it’s language, images, or Peppino. It doesn’t know. So it’s a statistical model that tries to continue these sequences. It makes itself an internal representation and probabilistically says what is about to arrive, what the next marble or the next word is.
And it’s a language model because for us those marbles are language. But for it they could be anything, it doesn’t care. And the same goes for a human brain: we can learn things that are completely - well, nervous tissue is much more generic than we expect, and the proof is found not only in artificial neural networks but also in brain organoids, which are little meatballs of nervous tissue created from stem cells, and those too can learn anything. But we’ll talk about brain organoids another time.
So: is it understanding?
Coming back to our sequence of marbles, to our compression - at this point we can say that this data compression is comprehension. I understood language in this way. What do you say to that?
Yes, they’re only coloured marbles. But if it manages to continue those coloured marbles, and those coloured marbles represent words, then I really am talking to it. So it understood. By compressing information, it understood - because knowing what comes next in the sequence of coloured marbles means it has abstracted the order, the patterns, the relations between these marbles so much that it really has understood. It manages it.
And so comprehension is comprehension, and intelligence is a form of pattern recognition.
With these questions I’ll close the video, which could go on for quite a while. I’d have plenty to say about this, but I’ll leave you with the question: in your opinion, is this comprehension? Is it intelligence?
In my opinion, absolutely yes.