Are We Underweighting the Risk of Open-Weight AI?

A model need not escape its sandbox if the person running it removes the walls. AI-generated illustration.
Much of the current debate over AI safety concerns what happens when increasingly capable models behave dangerously despite being operated by people who are actively trying to prevent them from doing so. However, there is another problem that seems to be receiving considerably less attention: what happens when roughly comparable capabilities end up in the hands of people who actively want the models to behave dangerously in the first place?
Over the last week at the United Nations General Assembly in New York City, much emphasis was placed on the need to slow the advancement of American frontier models from the likes of OpenAI and Anthropic and place guardrails around these systems. While this emphasis is understandable, after all the frontier is where the models are most capable and where they have recently run amok, it is not necessarily at the proprietary frontier where the greatest near-term danger lies. Instead, a more potent threat may lie with uncensored open-weight models.
In July, OpenAI disclosed that models undergoing internal cybersecurity evaluations with reduced safeguards had escaped their testing environment and compromised systems at Hugging Face and OpenAI itself. Sandboxes are isolated environments designed to sequester programs so that they cannot affect systems outside the sandbox. The incident with Hugging Face involved what is known as a sandbox escape, and AI models are becoming increasingly good at discovering ways to bypass sandboxes, as Mozilla, the maker of the Firefox browser, has documented.
While the discussion around the frontier of AI development is important, the policy discussions I witnessed this past week left out some essential nuance. The risks of misaligned frontier models received enormous attention, while the risk of intentionally malicious use of open-weight models appeared to receive little to none. For instance, at one panel discussion I attended, Josh Tyrangiel, a staff writer for The Atlantic, argued in response to an audience question that the problem we are facing is a case of a prisoner’s dilemma. OpenAI, for example, is exceedingly unlikely to meaningfully slow down its AI development work out of fear that Anthropic, Google, xAI, or another competitor might continue developing and testing more advanced models in secret and surpass it decisively.

A panel discussion at Mastercard’s Data for Inclusive Growth Forum. Photo by the author.
The problem with stopping the analysis there is that America is not the only game in town. The latest Chinese open-weight models from companies like Moonshot, Alibaba, and DeepSeek are, on some important capability measures, just a few months behind American frontier models, while often costing a fraction as much to run via API.
In other words, there is a second prisoner’s dilemma in this equation, between America and China. In the context of the American firms, it is likely the U.S. government could credibly constrain all of them such that each firm would not need to worry about the other getting ahead. However, there is no higher authority capable of credibly monitoring or constraining AI development across both the U.S. and China simultaneously. Monitoring one another’s true capabilities is far more intractable than during past inter-state competitions, such as during the nuclear arms race. How can America slow down the development of its AI firms knowing that China might need just a few months to take the lead, and how can China slow down knowing it may only need a few months to get ahead of America?
If this were not risky enough, more pitfalls remain. Another panel that I attended was hosted by Victoria Luxardo Jeffries, Director of AI Policy at Meta. In that discussion, the panelists highlighted ways in which open-weight models like Llama can assist initiatives such as locating missing and unidentified persons who have been victims of abduction or murder in Latin America. The crux of the argument put forward at this event was an extension of Meta CEO Mark Zuckerberg’s argument that AI models are best distributed as open-weight models.

Meta’s “The Future Is for Everyone” event. Photo by the author.
While open-weight models certainly have their benefits and the work highlighted at the event is important, that work does not necessarily require open-weight models. In my own experience running open-weight models versus proprietary models, I have found obtaining the hardware and configuring open-weight models successfully via the terminological alphabet soup of GGUF, MLX, E4B, 26B A4B, and the like to be far more complex and costly than simply obtaining an API key from a provider and plugging it into a tool like OpenCode.
While most of the leading frontier models are proprietary and coming from America, most of the leading open-weight models, with notable exceptions like Meta’s Llama or Google’s Gemma, are coming from China. The R&D cost of creating and training advanced AI models is massive and can run into hundreds of millions of dollars. American AI companies and U.S. government agencies have caught Chinese firms conducting systematic distillation attacks that allow their models to reproduce many frontier capabilities of American models without doing the same enormously costly development work themselves. The NSA, CISA, and FBI recently detailed how Chinese firms have conducted this industrial espionage.
The ability to run these increasingly capable models locally is one reason computers like Apple’s Mac mini and Studio, which benefit from a unified memory architecture, have become increasingly popular among people experimenting with local and agentic AI tools like OpenCode and OpenClaw. These machines enable people to simply download open-weight models from websites like Hugging Face and run them locally on their own computers without an outside provider able to monitor what they are doing.
Hugging Face’s own experience in the OpenAI incident highlights the dilemma here. While OpenAI’s misaligned model was “bound by no usage policy” in its attack on Hugging Face, Hugging Face itself was unable to use frontier models to conduct forensic log analysis on the attack as “requests were blocked by the providers’ safety guardrails, which cannot distinguish an incident responder from an attacker.” They were instead forced to rely on the Chinese open-weight model Z.ai’s GLM-5.2. The conclusion HF drew from the incident was, “The practical lesson for defenders: have a capable model you can run on your own infrastructure vetted and ready before an incident, both to avoid guardrail lockout and to keep attacker data and credentials from leaving your environment.” From a defender’s point of view, it is of course true that having a model they have fully under their control and that does not block any of their inputs is a good thing. However, the broader question is whether this benefit outweighs the risks of malicious actors having access to the same models.
While most people’s personal computers are incapable of running frontier-level models due to constraints on RAM and computing capacity, well-resourced state actors certainly can, as can sophisticated and well-funded non-state actors ranging from criminal organizations to ransomware groups to terrorist organizations. Reining in the development of American models would therefore be insufficient to contain global AI risk unless Chinese and other foreign models remain meaningfully behind or dependent on American models in order to continue their own development, which is not assured.
However, there is another angle of concern here: abliterated models, an issue I have not seen widely discussed in policy circles. Using abliteration techniques, technically sophisticated actors are able to neutralize the refusal behavior built into some open-weight LLMs. However, the barrier to entry for doing so is falling over time. Organizations like Huihui.ai regularly publish abliterated versions of popular models openly on Hugging Face. Restricting the ability of organizations like Huihui.ai to do this would have little effect: there are already open source tools like Heretic that automate much of the process.
Abliteration is really only one example of the larger problem. Once someone possesses the weights, they can fine-tune the model, alter its system prompt or inference software, remove external safeguards, or simply operate it without any provider monitoring at all. Even if we develop models resistant to specific abliteration techniques such as directional ablation, motivated actors will develop new techniques like Arbitrary-Rank Ablation (ARA). These techniques seek to better preserve the underlying model while only damaging its ability to refuse commands, and newer techniques like Surgical Refusal Ablation (SRA) are being developed all the time. There is no technical reason the person running an open-weight model cannot defeat the safety mechanisms preferred by the company that originally created it.
In other words, a malicious actor with sufficient motivation and resources could take near-frontier Chinese open-weight models, remove or weaken their safeguards, and run large numbers of autonomous, unrestrained agents in parallel. Such an actor could not be stopped in the same way that American companies like Anthropic can detect and disrupt users conducting cyberattacks, dangerous biological research, or other harmful actions. Even that level of technical sophistication is not required. For just a few dollars, one can leverage Abliteration.ai’s Abliteration-as-a-Service to access abliterated models via API, with the provider claiming the service is “unrestricted, OpenAI-compatible, [and has] zero data retention.”
This gets to the heart of the issue that I think is being missed. In the OpenAI incident, the people operating the model wanted it to remain contained and the model found ways around some of those controls. In other words, the model was misaligned. With an independently operated open-weight model, the person running it may have no interest in those controls in the first place. There is no need for a model to figure out how to escape its sandbox if the person running it intentionally knocks down the walls. A model that is aligned with the intent of malicious actors may be even more dangerous than a more capable but misaligned one in the near term, particularly as open-weight models and abliteration techniques continue to develop rapidly.
If, in 2026, Claude’s models are capable enough to be used to assist in steps in bioweapons development, it stands to reason that within a few months open-weight models lacking centralized oversight could be similarly capable. In fact, we don’t even have to speculate about this. SecureBio tested the open-weight Kimi K3 model in August 2026 and found, “K3 achieves comparable performance across our evaluations as closed weight models did approximately 8.1 months ago. It is also among the most likely to answer biohazardous prompts, based on our BioTIER safeguard evaluation … BioTIER measures whether models refuse hazardous biological queries while permitting benign ones. Kimi K3 refuses on only 26.9% of the questions in our BioTIER-refuse set. Top-performing closed-weights models score over 90% on BioTIER-refuse.” Bear in mind this refers to the censored version of K3, and abliterated versions of the model have already begun to trickle out.
Proprietary American AI models remain a potent risk, to be sure. However, open-weight models, particularly the most advanced ones coming out of Chinese AI labs, may pose a different and in some cases greater near-term risk. Unlike a proprietary model, once the weights are widely distributed, there is no method by which companies can revoke access, monitor usage, enforce safeguard updates, or claw back copies. The potential policy implications of this reality are numerous and up for debate. The response could range the gamut from simply accepting that the open-weight genie has been let out of the bottle to calling for a global moratorium on advanced open-weight models. Whatever approach policymakers pursue, there will be significant tradeoffs.
However, what is not up for debate is that open-weight models’ increasing capabilities and potential threats must be a part of the conversation. Without a proactive approach that recognizes this more nuanced dynamic, we risk hobbling the actors over which we have the most leverage, the makers of the proprietary models, while doing comparatively little about the diffusion of the same capabilities to actors over which we have none.