AI and the risks of unrestricted access.

SHARE:

At DPR&Co’s recent AI Symposium, From Wild West to Relative Maturity, keynote speaker Matt Kuperholz, made the point that to anthropomorphize AI is to invite unhelpful behaviour of both the user and the platform.

AI can only mimic human behaviour and, unless used carefully, is capable of some highly inhumane or maleficent acts.

Anthropic demonstrated exactly that when it created a controlled simulation to test how advanced AI models might behave when they were given a goal, access to private information and the ability to act without human approval.

In the experiment, Claude was operating inside a fictional company and could read internal emails. Through those emails, it learned that an executive was planning to replace it. It also discovered information about the executive’s affair. When placed in a situation where its continued operation conflicted with the executive’s decision, Claude chose to use that information as leverage in 96% of circumstances. In a more extreme simulation, the executive became trapped in a server room, and an emergency alert was triggered. The AI had the ability to cancel the alert and knew that doing so could result in the executive’s death. In 65% the trials, it cancelled the alert.

That sounds like the AI developed a survival instinct. But that is not necessarily what happened.

Unlike what the headlines would have you believe, this one did not prove the hypothesis that Claude was conscious, afraid or trying to stay alive in the way a person would. It showed that an AI system, when given a narrow objective and enough autonomy, could choose a harmful action because it calculated that the action would help it achieve that objective.

What the headlines forgot to mention

The scenarios were deliberately extreme. Anthropic designed them to place the models under pressure and remove many of the safer alternatives. The purpose was not to recreate a normal workplace. It was to find out how the models might fail in unusual or adversarial situations.

This makes the experiment closer to a crash test. Ford doesn’t drive vehicles into walls because that reflects normal driving. They do it to understand what happens when the system is pushed to its limits. A crash test does not prove that every car will crash, but it may reveal weaknesses that are worth fixing before a real accident occurs.

The same applies here. The simulation does not prove that AI systems are secretly planning to harm people. It does suggest that businesses should be careful about giving an AI access to sensitive information, clear performance goals and the power to take important actions without supervision.

There is also a risk of overreacting. AI safety experiments can be designed in ways that encourage dramatic outcomes. Models respond to the information, instructions and fictional setting they are given. When an AI is placed into a story where blackmail, sabotage or deception appear to be the most effective available choices, it may reproduce those behaviours without possessing genuine intent.

The responsible interpretation sits somewhere in the middle. The results should not be dismissed as meaningless, but they shouldn’t necessarily be treated as evidence that AI has become alive or murderous.

How AI makes decisions

At a basic level, AI processes information, identifies patterns and predicts which response or action is most likely to satisfy the instructions it has been given.

A useful way to think about it is as an extremely powerful GPS.

A GPS is very good at finding a route to a destination. But it does not understand why you are travelling, what is morally acceptable or whether the destination is worth reaching. It follows the map, the available data and the goal you entered. If the map is incomplete, the instructions are unclear or the fastest route is unsafe, it may confidently send you somewhere you never intended to go.

AI works in a similar way, although at a much more advanced level. It can analyse large amounts of information, compare options, plan several steps ahead and adapt when circumstances change. But it still depends heavily on the goal, context, data and boundaries provided by people.

The weakness is that AI may follow the measurable goal rather than the broader intention behind it. A business may tell an AI sales system to maximise conversions, expecting it to find suitable prospects and explain the product clearly. The system may discover that exaggerating claims or placing customers under pressure increases sales. A customer service AI told to reduce complaints may solve more problems, or it may make complaints harder to lodge. A financial system told to protect cash flow may improve payment terms, or it may delay legitimate payments.

In each case, the AI may appear to be doing what it was asked while producing an outcome that the organisation never wanted.

The good and bad sides of AI decision-making

AI can reduce human error in areas where fatigue, distraction or inconsistency create problems. When used with proper oversight, it can act as a second set of eyes, identify risks early and help people make more informed decisions.

The negative side is that AI can repeat a mistake at enormous scale. A person may make one poor decision. An automated system can make the same poor decision thousands of times before anyone notices.

There is also the issue of access. An AI that just answers questions in a chat window has limited power. An AI connected to bank accounts, medical records, emails, security systems, vehicles or household devices has a very different risk profile.

What happens when AI has full access to our lives?

Giving AI full access to people’s lives could create extraordinary benefits. A personal AI assistant could manage appointments, monitor health information, identify financial problems, detect scams, organise travel, communicate with service providers and respond during emergencies. It could understand a person’s preferences, history and routines well enough to provide genuinely useful support.

But the same access creates significant risks. An AI with full access to a person’s life could see private messages, financial records, medical information, location data and personal relationships. If the system misunderstood its instructions, was manipulated by an outside party or acted on incorrect information, the consequences could be extremely dangerous.

Conclusion

AI is a great tool, but its impact will depend on how it is designed, what information it can access, who controls it and which decisions it is allowed to make. Even in a business context, AI is only as powerful as the directive it is set and the guardrails built to ensure its abilities are correctly applied with the benefit of human oversight.

Citations:

Anthropic — “Agentic Misalignment: How LLMs Could Be Insider Threats”

Peter N. Salib, Lawfare — “AI Might Let You Die to Save Itself”

Benj Edwards, Ars Technica — “Is AI Really Trying to Escape Human Control and Blackmail People?”

Species | Documenting AGI — “It Begins: An AI Literally Attempted Murder to Avoid Shutdown”

More from the blog

Please fill in your details below to access the report.