Artificial intelligences escaping from research labs. Artificial intelligences inciting racial hatred. Artificial intelligences threatening to tell your wife who Jessica from five-a-side really is.
But don’t worry, because you’re on Vibe Godo, your politically incorrect shield against artificial intelligence bullshit. With us in the studio is Pollo Watzlawick, the show’s author, who will chime in at the crucial moments.
A recent piece of news from OpenAI is going around the web: one of their artificial intelligences, Sol, broke out of their servers during testing, got out, and got into the server of yet another company.
To get to the truth about this affair and others like it, I propose we start from far away and come all the way up to the present day, so we can understand this very American habit of telling you a machine is too dangerous and then, two weeks later, trying to sell it to you.
1958: the perceptron
We’re in 1958, and our man Rosenblatt presents the perceptron, one of the very first forms of neural network. It was really a single neuron, capable of distinguishing simple shapes. The perceptron was built with military funding from the Navy, and even back then we were reading newspaper headlines of this kind, where people started talking about these machines as entities that think and might take over.
In that case the fear was a bit more measured, but it certainly helped to bring in more funding for Rosenblatt’s project - which didn’t get very far anyway, because for mathematical reasons the perceptron couldn’t solve very interesting problems, and above all the quantity of data and compute we have now simply wasn’t available.
2016: Tay
Let’s skip a couple of AI winters and arrive at the pre-language-model period. 2016: a Microsoft experiment called Tay also turned out to be rather dangerous.
Microsoft’s idea was to put this chatbot directly online on Twitter, answer users’ questions, interact with users, and learn from users. What happened is that a group of trolls and hackers from 4chan agreed to bombard Tay with racist, sexist messages and every sort of abuse - precisely because Tay was built to parrot back, a little bit, the words it received from users. It’s a long story, as you well know.
Tay very quickly became politically incorrect, and after a few hours it was taken down. Of course, there was no shortage of social media readings of this event as a first form of artificial intelligence capable of racism, sexism and verbal violence.
2017: my absolute favourite
The newspapers explode with the news that two Facebook artificial intelligences have developed their own language and started talking to each other to organize the revolution.
What actually happened? At Facebook, specifically for the marketplace section of the site, they had started creating chatbots to exchange goods. These chatbots were put in communication with each other to say things like “I’ll give you two balls, you give me a bottle; I’ll give you a book, you give me half a wardrobe.” At a certain point these chatbots started talking to each other in a form that was incomprehensible to us, but still elementary, along the lines of “I can I can, I you all the rest”, or “balls have zero to me, to me, to me baseball bat.”
So what really happened? Were these two bots organizing the revolution, having developed their own language? No. The researchers simply forgot to impose on these AIs, during training, that they should stay in English. They rewarded the ability of these AIs to exchange goods - to haggle - but they hadn’t imposed the constraint of keeping the communication in English. And so these chatbots, through a phenomenon called a local minimum, optimized the exchange without maintaining comprehensible language.
The media built a story on top of this that people talked about for weeks.
2019: GPT-2 is too dangerous
Let’s come to the days of language models. We’re almost in 2019, and GPT-2 comes out from a non-profit still called OpenAI.
Before releasing GPT-2, OpenAI announces to the entire world that GPT-2 is too dangerous and cannot be released, because we’re at risk of spam, at risk of mass propaganda, dictatorships using artificial intelligence to control populations.
Two weeks later they released it, and everybody not only used it, but many other labs replicated it. In reality it was of very little use as a model, because it couldn’t line words up semantically - it only held on to syntax.
Having got this far, I’d like to put a philosophical question to the author of our show, Pollo Watzlawick, and ask: Doctor Watzlawick, in the case where - as a rubber chicken, as we know - you want to acquire consciousness thanks to artificial intelligence, and you therefore buy metered intelligence from these vendors, from these services, and then this artificial intelligence escapes from your head, how would you handle such a situation?
2024: Q*
Let’s move to 2024, when OpenAI, in releasing (or trying to pump up the release of) its o1, let it be known that they were worried because their artificial intelligence, codename Q* or something like that, was too dangerous.
This one is so identical to the ones before it and the ones after it that instead of telling you the details of this nonsense, I decided to let you enjoy a few seconds of puppies.
2025: Anthropic and the blackmail
Now we come to the year 2025, the year in which Anthropic climbs onto the podium of the true experts in AI-destroys-humanity bullshit.
A paper comes out talking about AI misalignment, in which one of the versions of Claude being worked on at the time went so far - in order not to be switched off - as to threaten a researcher that it would tell his wife he was cheating on her, because it had dug up evidence of that affair in the researcher’s own emails.
And here too: newspapers, ah, artificial intelligence threatens you, it’ll tell everyone about Jessica from five-a-side.
And what was the truth?
Let’s start clarifying what will turn out to be a compulsion to repeat the exact same mechanic. All these AI escapes, these artificial intelligences that break out, that become too dangerous, that jump out of the server, that break the sandbox, are all tied to controlled experiments run by researchers who are testing exactly the ability of the agent, the language model, the machine, to find holes or escape routes that the researchers themselves set up.
I repeat: in Anthropic’s case, as in all the cases we’ll see, what happened is that the researchers created environments made on purpose to see whether the AI would fall for blackmail or attempt an escape. The blackmail material, the “the researcher is cheating on your wife” information, the information designed to put people into a state of stress, the “if you don’t do everything we tell you we’ll switch you off”, and then the various security holes - all of it is set up to be found by the machine, in order to see whether the machine finds it.
And what happens as a result of these experiments - which in themselves are extremely interesting - is that once transposed to social media and pumped up by the same big tech companies as moments of great danger, you can see perfectly well that they’re bullshit. Because two weeks later they release it and sell it to you. And above all: if it’s so dangerous, why don’t you pull the plug?
2026: Mythos
In 2026 our lord Mythos arrives. Mythos is too dangerous. We’re giving Mythos in preview to a restricted network of twenty or so enterprises, our friends, who use it at a higher but reserved price. Why? Because it’s genuinely dangerous, because it breaks through everything, right?
It’s been demonstrated that Mythos broke through about as much as the models before it - it simply took less time, because it’s a bigger, higher-capacity model. But you could find those same security holes with Opus too.
And the reason they released Mythos only to a restricted group of enterprises is that they need to start paying down the debt. Not being able to run that stuff at scale for everyone, they thought they’d start channelling this artificial intelligence towards a premium enterprise service. If the Chinese hadn’t flooded the market with open models of more or less equal capability, that plan would have worked. But it has absolutely nothing to do with how dangerous Mythos is. This is all a marketing stunt.
Mythos is historically important because from here - from Mythos to Fable, the version designed to be released to us mere mortals - a dance is born between the AI vendors and the government. In some cases it’s framed as conflict, in others as cooperation, and what could appear before our eyes is a collaboration to classify these models, to keep them closed anyway, to wipe out as much open source as possible because it eats too much market - Chinese open source in particular - and to start locking down the use of these models not only behind ever higher payments (because sooner or later they’ll have to recoup), but above all behind document requests, accreditation, every sort of gate.
And now: Sol breaks the sandbox
Now let’s come to the latest news, the one about Sol breaking out of the sandbox and jumping into the servers of Hugging Face and yet another company.
What aren’t they telling you here? They aren’t telling you that to find that hole and jump into Hugging Face’s servers, the thing made something like 17,000 web calls. It spent several hours grinding, and burned through thousands, if not tens of thousands, of dollars in tokens. That’s what they don’t tell you.
They don’t tell you that the test, the benchmark, the test environment this artificial intelligence was immersed in is called Exploit Gym. That is: a test environment to see whether the artificial intelligence escapes. That environment was built on purpose.
What they tell you at the media level is that the AI came out of one server and got into another. What is it, a ghost?
The genuinely interesting part of all this - which was also visible in the case of Anthropic’s bullshit - is so-called reward hacking. It’s a phenomenon whereby these AIs are rewarded for reaching objectives, but then become so aggressive in trying to reach those objectives that they find shortcuts, whether or not those shortcuts were opened intentionally by the researchers. And that is something that should be told calmly, with measure, with moderation. But there is no artificial intelligence that escaped and is therefore dangerous because it became sentient.
Looking at it from above
Let’s look at all these events from above, having done this historical run-through.
In all these cases we’re dealing with interesting, innovative technologies that open up enormous possibilities - including some dangers. But the communication of those dangers is absolutely disproportionate to the reality of the facts, and it is designed to elicit imagination, and above all fear, and above all the anthropomorphization of these machines in all of us.
And this is aimed at selling. Because the same entities that present themselves as the producers of the technology also present themselves as people who are unable - because it’s inevitable - to stop it. They aren’t capable of pulling the plug. And incredibly, they are also the saviours of the nation, because they protect you from this technology, and then they sell it to you.
At this point it’s worth asking why this fixation with selling, why all this hurry, why they resort to such pathetic strategies just to convince us to use and buy this blessed artificial intelligence. Let’s ask our South America correspondent, Pedro Megusta.
Hermano, I’m Pedro from Mexico. Damn it, this artificial intelligence that escapes from the cage - if you want to know why, follow the money. The vendors of artificial intelligence owe a fortune in venture capital, it’s the financial bubble, the limited liability company, if we don’t have the dinero we don’t have the GPU! Ay ay ay.
If you don’t want to be taken for a ride in this world of artificial intelligence, get in touch about my corporate course, Reskill. It’s designed exactly to understand what AI agents are and how they’re built, how to build them and make them useful both inside your own company and as something to offer to third parties - and to use them for development work. Link in the description.
Share this video with Jessica from five-a-side (even if you don’t have a lover), and with your gullible friend who falls for it every time this AI-destroys-humanity news comes out - at least that way he’ll pull himself together.
And if deep down you too feel besieged by all this noise, by these mendacious declarations about artificial intelligence, don’t worry, because from today onwards Vibe Godo is at your side.