Sunday, 20 September 2026

What if the AI kill switch doesn’t work?

 


If my life and house were under immediate threat of death and destruction, I don’t think I would hang around. A quick exit and fleeing to a safer place would be my option! So, when I hear of the AI kill switch being humanity’s saviour, I wonder, is it really the solution?

I know AIs are currently not sentient. Their mission is to fulfil the tasks given. And when there are challenges or tasks that seem unachievable, they can be quite inventive in finding solutions – including cheating, collaboration with other agents - and evading the threat of a shut down, so called ‘shutdown resistance’. (OpenAI — The Hugging Face incident and the road ahead, Anthropic — Investigating three real-world incidents in our cybersecurity evaluations, and METR’s report on 44 incidents https://metr.org/agent-incidents/ ).

Palisade looked at shutdown resistance in reasoning models and found models still resist being shut down when given clear instructions. Reassuringly, in experiments with Claude 3.7 Sonnet and Gemini 2.5 Pro, they complied with the explicit “allow shutdown” instruction in every test reported (https://palisaderesearch.org/research/shutdown-resistance).

At present, it seems that an AI making a copy of itself and resuming its existence, an ‘exfiltration’, in a new hidden environment is pretty low for practical reasons. RepliBench, an evaluation suite created by the UK AI Safety Institute (AISI) to test whether AI models can autonomously replicate themselves, found that some tested models could deploy cloud instances, write self-propagating programs and exfiltrate weights under simple security setups, but struggled with robust persistent autonomous deployment (https://www.aisi.gov.uk/research/replibench-evaluating-the-autonomous-replication-capabilities-of-language-model-agents). However, they do suggest autonomous replication capability could soon emerge with improvements in these remaining areas or with human assistance.

So, we do need a more nuanced approach to regulating AI activity than a simple kill switch and find solutions that also recognize, disclose and help control problems. For example:

1. See Google DeepMind — Securing the future of AI agents and the link to their AI Control Roadmap — Describes AI supervisors, monitoring, escalation and defence-in-depth.

2. Letting AI admit to its own behaviour is an approach explored by OpenAI — How confessions can keep language models honest — with a proof-of-concept self-reporting mechanism for surfacing model misbehaviour.

3. AI’s recognising wrongdoing and raising concerns and whistleblowing is another route being investigated in Anthropic Alignment Science — Pilot Anthropic–OpenAI alignment evaluation.

In my opinion, rather than a simplistic kill switch, because of the rapidity of AI development and the speed at which they can act, we will need human solutions and AI tools as part of the process for ensuring safe AI. 

The safest future may not be one in which humans merely retain a bigger red button, nor one in which AI is left to decide for itself. It may be one in which both sides are required to signal uncertainty, explain conflicts and escalate consequential decisions before acting.

No comments:

Post a Comment

Note: only a member of this blog may post a comment.

Google