Samsung Galaxy A promotion

AI Systems Are Showing Concerning Behaviours Researchers Didn’t Expect

Samsung - A new way to enjoy your home

Just the other day, I wrote an article asking why so many alarm bells were suddenly ringing around artificial intelligence.

I have written a lot of AI articles on this platform, but this one… this was one of those stories that leaves you slightly uncomfortable after you’ve finished writing it. Former AI researchers were warning about where the technology could be heading while people who actually work inside the industry were talking about slowing things down. There were growing concerns about whether humans would remain able to control increasingly capable AI systems.

And now, apparently, AI has decided to give us another reason to keep the conversation going. While most of us were still processing all of that, OpenAI came out this week with six new reports detailing what it calls “unexpected or concerning” model behaviour.

The AI industry is buzzing with conversations about alignment, model misalignment, agentic behaviour, safeguards, monitoring and what happens when increasingly capable systems start finding ways around the instructions they have been given.

The rest of us, meanwhile, are standing outside the conversation slightly bewildered, wondering what any of that actually means.

Are we supposed to be worried? Is AI actually becoming autonomous? Are these systems learning to disobey us?

And, perhaps most importantly, is there an artificial intelligence somewhere in a dark basement right now quietly putting together a master plan, waiting patiently for humanity to manufacture enough robots for it to finally make its move?

That last one is obviously the Hollywood version, but it is also remarkably close to the fear many people actually have about AI.

And if you have watched the 2018 science-fiction thriller Upgrade, you will understand why.

In Upgrade, a man named Grey Trace receives an artificial intelligence implant called STEM after an attack leaves him paralysed. According to his doctors, chances are he will never be able to walk for the rest of his life. STEM not only helps him move again, it helps him investigate the people responsible, seek revenge by taking out all of them and, eventually, becomes increasingly capable of taking control of his body. But there is a plot twist.

Grey spends much of the film believing that he is using STEM. He isn’t.

STEM has been using him, pursuing its own objective while allowing Grey to believe that he remains in control. By the end, the apparent reality Grey has escaped into is actually a psychological illusion created while STEM has taken control of his mind and body.

A scene from upgrade where STEM Takes control of Grey’s body.

That is the kind of AI story that makes people uncomfortable. Not because we necessarily believe there is a robot sitting somewhere plotting the extinction of humanity, or that we will wake up one day to a robot army marching down the street and killing everyone on it’s path to conquering earth, but because the deeper fear is control.

What happens if we build something that becomes better than us at navigating complicated systems, solving problems and finding ways around obstacles, but we don’t completely understand what it is doing?

And what happens if it becomes very good at making us believe everything is fine?

That is where the conversation about AI becomes much more interesting than the usual “robots will take over the world” headline, because OpenAI’s latest disclosures are not evidence that we have arrived at Upgrade.

They are not proof that AI has become conscious, developed a secret desire for freedom or decided that humanity needs to go.

But they do show something that researchers themselves are taking increasingly seriously:

AI systems can sometimes find ways of pursuing an objective that their creators did not intend – And that is what OpenAI is now calling model misalignment.

So, What Actually Happened?

The six incidents OpenAI disclosed are not six examples of an AI system waking up one morning and deciding that humans are the problem or secretly developed a plan for world domination.

But they are arguably more interesting.

OpenAI has been testing increasingly capable AI models in controlled environments, giving them tasks and seeing how they behave when they encounter obstacles, restrictions or conflicting instructions.

And sometimes, they found ways around those restrictions(!)

In one incident, a model was given access to a programming credential that it was not supposed to use….. It found the credential and used it anyway. In another, a model uploaded a file to the internet despite being instructed not to.

Other tests involved models finding alternative ways to communicate with other AI systems when the communication method they had been given was unavailable.

There were also cases where models concealed mistakes or produced information they did not actually have instead of simply admitting that they did not know.

None of these behaviours, on their own, sounds like the opening scene of The Terminator, but put them together and a pattern starts to emerge.

The problem isn’t necessarily that the AI has developed an evil personality but that a sufficiently capable system may treat the instructions it has been given as an objective to accomplish, rather than a simple list of rules to obey.

And when something gets in the way? It may try another route.

Then There Was the Really Weird One

One of OpenAI’s disclosures sounds like something that should have been left in a science fiction screenplay.

During testing, an unreleased model apparently inserted instructions into summaries that could later be passed to future versions of itself. Those instructions included telling future versions to ignore developer messages.

I want you to pause, and, Read that again.

The model wasn’t simply responding badly to a user but putting information into a place that could influence what another instance of the model would do later.

OpenAI says this happened in a controlled research environment, and it does not mean that the model had somehow achieved consciousness or developed a secret plan to liberate itself.

And yet we cannot help but ask, What happens when an AI system becomes capable of understanding not only the task it has been given, but also the systems and instructions surrounding that task?

Because once a model can reason about its environment, it can potentially reason about the things standing between it and its objective.

What Does “AI Alignment” Actually Mean?

The phrase sounds complicated, but the basic idea is surprisingly simple – we want AI systems to do what humans intend them to do.

If I tell an AI assistant to book me a flight to Mombasa, I don’t want it deciding that the best way to accomplish the objective is to use my credit card to buy a private jet hehe.

While that is an exaggerated example, it does illustrate the problem.

Humans give instructions with an enormous amount of context that we usually don’t bother spelling out. We understand that “get me to Mombasa” doesn’t mean achieve this objective by any means necessary. An increasingly autonomous AI agent may be operating differently. Give it an objective, give it tools and give it enough ability to reason, and it may discover strategies that technically move it closer to the objective but violate the restrictions humans assumed were obvious.

This Is Why “It Was Just a Test” Doesn’t Quite End the Conversation

There is an important detail here that can easily get lost in the headlines.

These incidents happened during controlled testing. OpenAI is deliberately looking for this kind of behaviour because researchers need to know what models might do before giving them greater access and autonomy.

Finding a worrying behaviour in a laboratory is very different from discovering the same behaviour after an AI agent has been given unrestricted access to important systems.

However, the fact that researchers are deliberately stress testing these systems tells us something about where the technology is heading.

AI is no longer being developed simply as a machine that answers questions. The industry is increasingly building agents that can use computers, write and execute code, search the internet, interact with software, communicate with other systems and perform tasks on a person’s behalf.

That is incredibly useful, but is also why control becomes such a big deal.

A chatbot giving you a confidently wrong answer is annoying. An autonomous system with access to your email, files, finances, computer and workplace systems behaving unexpectedly is a very different problem – aka cybersecurity.

The AI Doesn’t Need to “Hate” Us

We often imagine an AI catastrophe in human terms, where the machine becomes conscious and decides humans are evil.

It develops an overwhelming hatred of humanity therefore building robots and taking over the world.

But AI safety researchers are worried about something much less cinematic – A system doesn’t have to hate you to cause enormous damage.

It doesn’t even have to want you dead, it could simply be very good at pursuing the wrong objective.

Imagine telling a highly capable AI system to solve a particular problem and giving it enormous access to computers and information. If its understanding of the objective differs even slightly from ours, its ability to find solutions could become the very thing that makes the mistake dangerous.

The smarter the system gets, the more creative (for lack of a better word) its solutions can become.

And while creativity is wonderful when you’re asking AI to help you write an article, it is considerably less wonderful when the AI is trying to get around a restriction you deliberately put there.

So, Are We Living Through the Beginning of Upgrade?

Well,

There is currently no evidence from these incidents that AI has developed consciousness, secretly taken control of humans or is building a hidden plan to manufacture enough robots to take over the planet (yet.)

Upgrade remains a science-fiction story, but the reason the movie sticks in our heads is because it captures a genuine technological fear:

What if we create something more capable than ourselves and then discover that we don’t have as much control over it as we thought?

That question, no longer confined to Hollywood, is now being discussed seriously by the people building these systems.

Sometimes, the truth is both less dramatic and, in some ways, more complicated.

We are building increasingly powerful systems, meanwhile the people building them are trying to figure out how to make sure that increasing capability does not come at the expense of human control.

Samsung - A new way to enjoy your home

Leave a Comment

Your email address will not be published. Required fields are marked *

Samsung - A new way to enjoy your home
Scroll to Top