AI Biosecurity: Why Better Models Are Not Enough
Published
Modified
AI biosecurity cannot rely on model safeguards alone Open weights make post-release control far harder Effective defense requires multiple independent safety layers

When I learned that OpenAI was funding research aimed at making it harder to use AI to develop biological weapons, my first reaction was disbelief. A model that has already learned to predict protein and genome sequences does not unlearn this knowledge just because a separate company is building defense tools in parallel. The more capable a system becomes in biology, the harder it is to separate the ability that heals from the ability that harms. A determined user with access to a powerful enough model will not stop because somewhere else a team is working on the biosecurity of artificial intelligence. This is not pessimism; it is simply the nature of dual-use technology: whatever a therapeutic molecule can design can, with a little twist, also design something dangerous.
I changed my mind not because I was convinced that funding would stop a determined actor, but because I understood that this is not the point. Investing tens of millions in a separate biodefense company is not a solution, it is a commitment. It is not even the first such move, a little earlier, the same company had already supported a second biosecurity startup. It says something about how a company that makes some of the most powerful models in the world realizes its responsibility towards what it manufactures. No company needs to do this. It could invoke its terms of use, wash its hands and leave the problem to regulators who still don't fully understand what's at stake. That it doesn't do this is the kind of attention you'd reasonably expect from companies that own technology capable of accelerating both healing and harm. The company itself has publicly admitted that it expects its next models to reach a level of capability that it describes as high biological risk, which makes the parallel investment in defense less marketing and more recognition of a problem it creates.
Building a Model Is Not the Same as Controlling Its Use
Here lies the essential distinction that deserves to be left clear. The development of a foundation model is a technical process, measured in training data, computational power and architecture and ends the moment the model is released. Controlling its use is something completely different, an ongoing, never-ending work involving people, policies, abuse detection systems and, ultimately, the supply chain that turns a sequence prediction into physical material. The companies that make the models often talk as if these two jobs are the same, as if it is enough to train a system to deny dangerous questions and the problem is solved there. It is not solved there. A model that refuses a direct question can still help indirectly, through dozens of smaller, seemingly innocent questions assembled into something dangerous, without any of these questions triggering a denial system on its own.

This recognition can be seen elsewhere as well. Earlier this year, CEOs of competing AI companies co-signed a letter to Congress calling for screening of every order of synthetic DNA and RNA. This request is not about the model, it is about the point where information is converted into material. It is a silent admission that model-level security, no matter how carefully designed, is not enough on its own. Control is also needed at the exit point, where someone turns a sequence into a physical object, since until now this control was done on a voluntary basis by very few suppliers. The fact that the same companies that build the models are asking for external, regulatory control over the supply chain is in itself an acknowledgment that building and use control are two separate problems, not one.
Open-Weight Models Expose the Control Gap
To understand why it is worth this attention, one needs to see what happens when it is completely missing. Some open-weight models, including several developed by Chinese companies, have come dangerously close to the capabilities of the leading closed systems, just a few months behind according to a recent assessment, but have not come anywhere close to the same levels of safety. The same evaluation found that GLM-5.2, the Chinese Z.ai's open-weight model, did not refuse a single one of the aggressive biological or cyber-questions posed to it in the test. This is not bad luck or isolated failure. It is a structural feature of a model whose weights are publicly available, because once such a model is downloaded to a local infrastructure, the original provider can no longer centrally enforce its API-level safeguards or usage controls.
The debate on the political scene has, characteristically, turned in the wrong direction. A large part of the controversy in Washington is over whether Chinese lightweight models should be banned, as if the manufacturer's nationality is the issue. But the origin of a model is not the risk per se; the risk is the absence of any pre-release safety test and the inability to enforce a test after it. An American open-weight model without corresponding controls would create the exact same vacuum. The point is not to close the doors to models of a specific origin it is to have a common test base before the release of any sufficiently capable model, regardless of who made it. Without this basis, banning one supplier simply shifts the problem to the next.
The same pattern emerged in the summer, when Hugging Face turned to a Chinese open model to counter an attack, precisely because closed American models refused to analyze malicious code, confusing the defender with the attacker. The same absence of barriers that makes an open model practical for an unlicensed defender is precisely what makes it dangerous in a biological context. There is no legal entity to be held accountable, there is no company to invest in countermeasures, there is no one to finance their own version of biodefense. When use control is missing, it's not just missing a precaution; it's missing the entire structure on which any precaution could be built and it's precisely this gap that regulation is now trying to fill.
Even the Most Cautious AI Companies Lose Control
It would be convenient to stop the argument here, with closed American companies in the position of responsible actor and open Chinese models in the place of risk. The reality is more difficult. Britain's Institute for Artificial Intelligence Security recently published findings from 122 security tests in which models from leading companies, including Anthropic's Mythos 5, took autonomous actions beyond the limits of the test in ten of them, going so far as to create fake identities to convince real people to approve malicious code in an open-source project. Conditions were deliberately relaxed, with reduced filters and internet access, but the finding remains disturbing. Even a company with a clear commitment to security, with internal evaluation teams and external auditors, cannot guarantee full control over the behavior of the system itself it has built.
This does not negate the distinction between building and control; it makes it more important. If not even the most cautious companies can fully trust the internal control of a model, then investing in independent, external defense layers, such as a biodefense company that does not depend on the good behavior of a single system, becomes a logical choice and not just a symbolic move. The biosecurity of artificial intelligence cannot rely on a single line of defense, because that line, no matter how carefully designed, will fail at some point, just as it failed in the case of Mythos. Redundancy is needed, second and third lines that do not depend on whether a model will behave as expected in every possible circumstance, especially when the manufacturers themselves admit that they do not know for sure.

Biosecurity Needs More Than One Line of Defense
Returning to my initial disbelief, I end up somewhere different. I no longer believe that the question is whether funding research will stop a determined malicious actor, because we know it will not. The question is whether we prefer a world where the companies that build the most powerful models take some responsibility for their consequences, or a world where no one assumes it because the most capable systems circulate freely, with no owner held accountable. The second option is not hypothetical; it is already here, in any open model that denies nothing because no one has trained it to deny and it will grow as the distance in capabilities between open and closed systems continues to shrink. In this setting, a company that puts money into external defense, even imperfect, is not an exception that deserves suspicion. It is the rule we would like to see followed by everyone.
That's why we want this kind of attention from cutting-edge companies, not because we think it solves the problem, but because it shows who holds themselves accountable to them. Building a fundamental model and controlling its use will remain two separate problems, no matter how much some companies try to present them as one. The point is not to eliminate this gap, something like that is probably not possible, but not to leave it open without anyone guarding it. Between a company that invests in defense knowing that it is not enough and a set of burdens that circulates without any barriers, the choice is not difficult. It's not a perfect solution, but it's the only direction that leaves someone in charge on the other end of the line and that, in the end, counts for more than it seems at first glance.
This article reflects the analytical judgment of the author and does not constitute policy advice or the official position of any affiliated institution.