If my life and house were under immediate
threat of death and destruction, I don’t think I would hang around. A quick
exit and fleeing to a safer place would be my option! So, when I hear of the AI
kill switch being humanity’s saviour, I wonder, is it really the solution?
I know AIs are currently not sentient. Their mission
is to fulfil the tasks given. And when there are challenges or tasks that seem
unachievable, they can be quite inventive in finding solutions – including cheating,
collaboration with other agents - and evading the threat of a shut down, so
called ‘shutdown resistance’. (OpenAI — The Hugging Face incident and the road ahead,
Anthropic — Investigating three real-world incidents in
our cybersecurity evaluations, and METR’s report on 44 incidents https://metr.org/agent-incidents/ ).
Palisade looked at shutdown resistance in reasoning models and found
models still resist being shut down when given clear instructions. Reassuringly,
in experiments with Claude 3.7
Sonnet and Gemini 2.5 Pro, they complied with the explicit “allow shutdown”
instruction in every test reported (https://palisaderesearch.org/research/shutdown-resistance).
At present, it seems that an AI making a copy
of itself and resuming its existence, an ‘exfiltration’, in a new hidden
environment is pretty low for practical reasons. RepliBench, an evaluation
suite created by the UK AI Safety Institute (AISI) to test whether AI models
can autonomously replicate themselves, found that some tested models could
deploy cloud instances, write self-propagating programs and exfiltrate weights
under simple security setups, but struggled with robust persistent autonomous
deployment (https://www.aisi.gov.uk/research/replibench-evaluating-the-autonomous-replication-capabilities-of-language-model-agents).
However, they do suggest autonomous replication
capability could soon emerge with improvements in these remaining areas or with
human assistance.
So, we do need a more nuanced approach to regulating AI activity
than a simple kill switch and find solutions that also recognize, disclose and help control problems. For example:
1. See Google DeepMind — Securing the future of AI agents
and the link to their AI Control Roadmap — Describes AI supervisors,
monitoring, escalation and defence-in-depth.
2. Letting AI admit to its own behaviour is an
approach explored by OpenAI — How confessions can keep language models honest
— with a proof-of-concept self-reporting mechanism for surfacing model
misbehaviour.
3. AI’s recognising wrongdoing and raising
concerns and whistleblowing is another route being investigated in Anthropic Alignment Science — Pilot Anthropic–OpenAI
alignment evaluation.
In my opinion, rather than a simplistic kill switch, because of the rapidity of AI development and the speed at which they can act, we will need human solutions and AI tools as part of the process for ensuring safe AI.
The safest future may not be one in which humans merely
retain a bigger red button, nor one in which AI is left to decide for itself.
It may be one in which both sides are required to signal uncertainty, explain
conflicts and escalate consequential decisions before acting.

No comments:
Post a Comment
Note: only a member of this blog may post a comment.