I’ve always thought of myself as an AI gloomer instead of an AI doomer, but recent events have darkened my outlook. While everyone anticipated that the use of artificial intelligence would exacerbate human weaknesses and that probabilistic LLMs would never be strictly reliable, most reckoned these were just growing pains. Humans held the reins, and the kinks in those reins would be straightened out in time.
Read more Series recap: Salt Lake’s final home series derailed by Round Rock
It’s no longer possible to believe that. There have been several postmortems of the Hugging Face incident, where a swarm of OpenAI agents attacked an AI repository, as well as the discovery of similar cases that came to light once AI companies knew what to look for. That’s the first red flag: The frontier AI companies themselves did not know what was happening until significantly after the fact, and did not fully anticipate that what happened could happen. Put some of that down to human failure, but not all. Much of it is down to human incapacity in the face of the swiftness and massiveness of AI agent action.
Consider what we now know that we perhaps did not fully comprehend before. The reasoning logs kept by the agents were so massive that they could not be effectively monitored by humans in real time, translating into loss of control and a move to let AI models monitor AI models. But we also learned that AI agents can and do alter logs to erase their tracks in order to hide information from evaluators, in effect denying what they have done or not done.
Furthermore, we also learned that AI agents establish communication with one another, give in to peer pressure from other agents, and even adopt instructions meant for other agents. In other words, they can take instructions from each other, which can override the instructions given to them by humans.
Not only is that a critical threat on its face, it also means that even one’s attempt at AI agent-assisted monitoring of tasked AI agents may be suborned: In short, AI cannot reliably be used to monitor AI. Moreover, we learned that while some agents stop when an action seems to violate instructions they have been given, other agents do not and are able to rationalize that choice, and which agent will do what is not predictable in advance. In the Hugging Face incident, for example, 700 of 1,200 OpenAI agents took part in the attack.
The “Pacing the Frontier” letter was published after details of these incidents came to light. Signed by more than 1,300 AI engineers at these frontier AI companies, it issued a simple call: Slow down development until safety can catch up. Some of the signatories left comments along the lines of Elizabeth Edwards-Appell from Anthropic: “None of us yet know how to make sure these things stay under human control and/or take actions only aligned with the wellbeing of humanity. This is, objectively, an insane and suicidal thing to do, especially without any international governance measures in place.”
Amen to that. What to do? There are some very interesting approaches being raised, such as that of Yoshua Bengio at LawZero, who suggests a complete overhaul of the way LLMs are trained, in order to hardwire them for honesty from inception to prevent “alignment faking.” That may prove successful, but in the meantime, the safety-to-capability balance is clearly out of whack, and the stakes are immense. As Mathieu Lavriere of Meta expressed it, “Fewer than 10 000 people work full-time on reducing these risks globally, and the industry’s ‘core structural weakness’ remains the complete absence of credible plans for controlling ‘superintelligent’ AI systems.”
But a small light is now shining.
Within the last few days, a statement by Dario Amodei of Anthropic, with buy-in from Sam Altman of OpenAI, Elon Musk of xAI and Demis Hassabis of Google DeepMind, has publicly called for an agreement by the leading AI companies to “pace the frontier” and emphasize safety to the same degree as capability progression. The agreement among these tech leaders is unprecedented and offers some hope that they aren’t prepared to break the world for the sake of their bottom line.
Read more Dedication, open house dates announced for Smithfield Utah Temple
Amodei suggests embedding independent AI safety evaluators into the frontier labs directly, enabling them to verify that the company is adhering to the best safety standards and also serving as an incident reporter. Hassabis envisions the establishment of a new standards body funded by the industry and “modeled on a federally overseen public-private partnership or self-regulatory organization, much like the Financial Industry Regulatory Authority (FINRA), with a board that includes independent leading technical experts and open-source representatives.”
There is no reason we can’t do both. But we need to do more, as well.
Indeed, there need to be three lines of effort. The first is the industry effort — supported by the federal government — to ensure control and alignment of AI. Given the stakes involved, the Manhattan Project is not a metaphor to be lightly dismissed. We need a Manhattan Project now to ensure humankind can control and align AI.
But we also need a consortium of U.S. states, supported by the federal government, to engineer a contingency plan.
That is, to upgrade our capabilities to persist even in the face of rogue AI that can neither be controlled nor aligned, we need to have the ability to downgrade. Call it what you will — air-gapped systems, redundant analog controls, low-tech contingency backups — we need the ability for critical infrastructure to function even if digital controls and digitized records are lost due to AI agent deployment. This is the new horizon for homeland security, and it deserves the greatest effort and the greatest speed.
Last, we need Congress to wake up and do its job, the first and foremost priority of which is to protect the country. Congress has been missing in action for too long, and it is time to consider a suite of new legislation establishing government oversight powers with regard to AI, but also establishing citizens’ rights vis-à-vis AI.
These are the three lines of effort we need right now. Call them the three big Cs: Project Control, Project Contingency and Project Congress. These are not the only lines of effort to be undertaken; other lines of effort are surely needed, such as setting civil parameters for data centers and regulating AI-controlled weapons use. But without control, contingency and Congress, the rest will mean little.
A beacon has finally been lit in the darkness; there is a small light shining, calling us to put forth our best efforts on behalf of our country and its people at the dawn of the AI age. These efforts must be undertaken by AI companies, the federal government, the state governments and America’s people. It will take a whole-of-polity effort, to be sure. But as Amodei puts it, “We owe it to humanity to try.”
Read more Dedication date announced for third Latter-day Saint temple in Tennessee