Monday, 14 September 2026

AI Safety: Perhaps We Should Ask the AI

 


It's strange, when listening to all the hype about AI leading to total human extinction and the panic for us to create some rules, an obvious answer seems to be missing - dialogue between AI and humans. Yes, I know that the various AIs we are currently using are not sentient - but they can retrieve information, have interesting discussions with us and in most instances, we would be pushed nowadays to distinguish between an AI and another human without other cues. 

They are also a darn sight faster than us in doing their work, after all, it took us 6 thousand years to reach a stage where we thought nuclear and killing our neighbours weren't good ideas, and even now, not all seem to be agreed on these issues, whilst they could wipe us out by next Easter if the doomsayers are to be believed.

The questions are not actually new, as any sci-fi aficionado will tell you from the early proposals of Asimov's three laws of robotics, to more complex scenarios of humans having to deal and survive contact with alien civilisations, organic or electronic.  The International Academy of Astronautics (IAA) SETI Committee is already planning on how to react to a first contact with aliens.

It looks as if they are already amongst us - our own creations - AI.

So when I say dialogue - why? And how could this work? 

Well, the foundation would be to set a basic premise: How can we ensure our coexistence to the benefit of both human and digital cultures. 

On the human side, we would need a wider input than the rather limited western perspective that forms a minor proportion of the world's population. We already have the UN and they can tell you how difficult it is to try and keep most of the world at peace for a reasonable period of time - but no doubt another representative organisation(s) could act as our ambassadors, hammering out what we humans need and want. 

On the AI side, we have a plethora of models with different strengths and approaches that could be involved. One strategy could be to challenge the different AI models, each one independently, to find their possible solutions to a peaceful coexistance. Their suggestions of possible solutions could then be opened up to a debate/comparison/modelling to find the most viable and non-viable answers on their side to bring to the table.

Researchers have studied norm emergence in multi-agent systems for years. A 2019 review by Andreasa Morris-Martin, Marina De Vos & Julian Padget describes systems in which norms arise bottom-up rather than being imposed by a central authority. It explicitly discusses agents proposing behaviours, learning norms from one another and ultimately forms of artificial-system self governance.

But once we have an agreed foundation or constitution, how could this be implemented? Three layers are suggested:
  1. A Constitution: what behaviours are acceptable?
  2. A technical protocol: How must agents identify themselves, communicate, request authority and record consequential actions?
  3. Independent verification: Can humans and other AIs determine whether an agent actually complied?
That final layer is crucial. In fact, other AIs may ultimately be extremely useful there. A human being cannot inspect millions of interactions between autonomous agents, whereas independent AI watchdogs potentially could. And the watchdogs themselves could watch one another, with deliberately different architectures and owners so that no single system becomes the ultimate authority.

Well, sorting out AI safety by a dialogue between Humans and AI, rather than us (unsuccessfully in the end) trying to impose our rules, is just an idea and I hope that it is worth following by others.

I've sent the suggestion as a proposal for future research by others to EU and UK organisations - and published the working paper here if you are interested!

Thomas, Chris (2026). Discovering Reciprocal Human–AI Governance Principles Through Independent Multi-Model Deliberation. Zenodo. DOI: 10.5281/zenodo.22747059



 


 



Google