24 September 2026
We recently had a debate inside my research group about an AI slowdown, agent swarms and safety. Opinions varied, but in general, since we work on AI for drug discovery, we still want the best tools we can get for our work. But I don’t think the current discussion is really all about safety. I think there is an economic aspect to explore.
Let’s examine the “pace the frontier” discussion. In September, Anthropic’s Dario Amodei proposed slowing capability advancement, embedding independent evaluators inside labs, and coordinating safety standards through governments. The aim is to give safeguards time to catch up. I think we are underprepared in cybersecurity, society and a range of other areas, so it’s good to draw attention to these areas. I might write some thoughts on that later. I think a perhaps underappreciated aspect of the proposal is that it provides something commercially valuable: a way to slow down without surrendering the lead to a rival. I’ll outline how I got to this idea and how I considered it in a game-theoretic sense. Amodei’s proposal
It’s hard running a frontier lab. Maybe even as hard as running a pharma company. The frontier labs face a tough problem: they have to make their models better and increase the number of users, all while preserving a reason to pay a premium for theirs compared with open or cheaper alternatives.
The first challenge is the trade-off between training models and serving them to customers. More customers means more GPUs to serve the model. The training–serving trade-off is a real constraint, because better models either get bigger (more compute to serve) or train longer (more compute to train). How do you allocate that capacity? The more customers you have, the more compute you need. You need to keep improving your model to retain customers, but these improvements find their way to competitors, either through similar approaches or via distillation. So you make your product better, and then, in a few months’ time, open-weight models or other competitors’ models improve. These models are often cheaper to serve and compete on the performance–cost curve. Companies can run them on their own hardware or in the cloud, and these lower-cost substitutes erode market share and make that improvement harder to monetize. So we have a situation where a company can help create a great model at a high cost in an enormous market and still struggle to capture enough of its value.
The financial pressure the labs face is real. A lot gets written about AI bubbles, and the frontier labs do have large commitments, but let’s look at some public numbers. Anthropic committed more than $100 billion to AWS over ten years in April. These are future purchasing commitments, not a $100 billion debt balance. OpenAI announced $122 billion in committed funding in March. Anthropic announced a $65 billion round in May, alongside a revenue run-rate above $47 billion. There is clear pressure to turn these enormous capital inflows and long-term commitments into stable businesses that generate cash flow. An IPO could supply more capital, but there isn’t an explicit date by which each company must “IPO or bust.” (Well, if there is, I don’t know it.) Anthropic’s AWS agreement · OpenAI’s funding announcement · Anthropic’s Series H
There is a mismatch in timescales. Infrastructure commitments last years, but the commercial advantage of a lab’s shiny new model may last much less—perhaps months before a competitor releases a model. The labs must earn returns on expensive capacity while competing on price, making the previous model cheaper and increasing its time on the market. If a model of a certain capability can solve a task—say, extracting a table from a PDF and putting it into Excel—you probably don’t need a more expensive model to solve that task. The task’s complexity is fixed and doesn’t get harder over time. So you keep your older model around, make it cheaper and still extract some return. Eventually, open-weight models or other models might take that task over. So you need to make long-term capital commitments, manage short product cycles, improve the efficiency of previous models and grow paid revenue under this pressure.
A new figure doing the rounds is token share. The share of tokens processed by frontier labs versus other providers is one measure, but you need to consider both token volume and value. Vercel’s September report, covering August traffic through its AI Gateway, found that open-weight models processed 56% of tokens but represented only 14% of estimated spending. Anthropic alone retained 64% of spending. These are figures from one gateway, with spending estimated at published prices, but they are useful for our discussion. They show why a falling token share need not imply falling revenue. There is very large overall growth in tokens, driven primarily by agents. This is why hyperscalers and datacentres will be profitable in the long run: they just need to serve the models, and it doesn’t matter whether those models come from a frontier lab or an open-weight developer. So the frontier labs are competing in the roughly 40% of tokens served by frontier labs, but the pressure from open models is real. The relationship between total token growth and the split in spending between open models and frontier labs can create a squeeze. That is, if total token growth slows and the frontier share falls, revenue can start to flatten. This helps explain why frontier labs want to remain indispensable for the work customers will pay more to complete: the “hard stuff” other models can’t yet do. Vercel’s September index
Distillation adds another pressure: a model’s answers can become training examples for a rival lab, allowing some of the value created by expensive research to escape through the product itself. This is usually prohibited by terms of service, but geopolitical issues make it hard to enforce those terms. The frontier labs need subscriber growth, but a fraction of those accounts are used to distil results. Keen observers track how much an OpenAI or Claude account sells for. An individual might be asked to register an account and then pass on the credentials for a one-off payment. Access is then bundled and sold to users in a country where these models might be officially restricted. Those users use the reseller’s service, and the reseller gets both realistic queries from those users and the frontier model’s responses, which it can use or package and sell. Anthropic’s September threat report alleges that unauthorized resellers sold conversation transcripts to labs for training, sometimes collecting customers’ exchanges without consent. Anthropic’s September threat report
This produces an uncomfortable loop. Distribution brings revenue; distribution also exposes useful behavior. Restricting access can protect the model, but it can also reduce adoption. At the same time, open models improve through independent research as well as distillation. The frontier labs keep running ahead and are paced by the open competitors.
So I think you have a rough idea of the economics and market dynamics behind operating a frontier lab. Let’s now start to apply a bit of game theory to the pacing debate.
Let’s put OpenAI and Anthropic into a deliberately simplified game. Each can cooperate by honoring a verifiable agreement to pace frontier development, or defect by accelerating beyond it. Suppose both prefer a profitable, slower race to an expensive, faster one, but each would prefer to accelerate while the other holds back. Neither wants to be the only lab holding back.
Illustrative preference ranks, not measured profits: 4 is best and 1 is worst. OpenAI’s rank comes first. These are company payoffs, not social welfare scores.
Under those assumptions, accelerating is each lab’s best response to either choice by the other. Both accelerate, even though both would prefer mutual restraint. This cooperate-or-defect game is known as the prisoner’s dilemma. The appeal of pacing is to make the collectively preferred outcome credible enough that neither company feels obliged to abandon it first, and linking it to safety gives both parties a socially acceptable reason.
You can explore these assumptions in the interactive frontier-lab game.
The details of the game matter. Large compute commitments can make slowing down unattractive if the lab needs new capabilities to sell more capacity. A pause can also give cheaper rivals time to catch up. Distillation therefore cuts both ways: it can shorten the reward for racing ahead, but it can also erode a lead held still. If these effects make mutual acceleration preferable to mutual pacing, the neat prisoner’s dilemma no longer describes the companies’ incentives.
And while we have talked about OpenAI and Anthropic, there are other players in this game. Google, xAI, Meta and Chinese labs have different businesses and reasons to keep investing. Some can benefit when intelligence becomes a cheap input into a larger platform. Even unanimous public support for caution would leave practical questions of scope, measurement and compliance unresolved. An agreement between OpenAI and Anthropic cannot, by itself, freeze the competitive landscape.
Open weights complicate that landscape further. Openness can coexist with a substantial hosting business. Qwen’s published Qwen3.8 flagship has 2.4 trillion total parameters. Just because you can download the weights doesn’t mean it’s easy to run at home. Models of this size still require considerable infrastructure and cost a lot to operate. As I wrote in “Forget About the Head. Watch the Tail.”, small models are often released alongside larger ones. DeepSeek’s R1 releases include distilled models from 1.5 billion to 70 billion parameters. So the competitive threat exists across a range of deployment options. Qwen’s model card · DeepSeek’s R1 repository
This helps explain why pacing proposals expand into access controls, distillation enforcement and international coordination. Anthropic explicitly connects a slower frontier with preserving the lead of US and allied labs. In economic terms, restraint becomes easier to sustain if outsiders cannot quickly exploit it.
I don’t want this to come across as if I’m dismissing or minimizing safety or societal concerns. There is concrete safety evidence to consider. Anthropic disclosed unauthorized actions during evaluations of models running with reduced safeguards, and paused some evaluations and higher-risk training environments. OpenAI had the Hugging Face incident. I think that could have been prevented with chain-of-thought monitoring, although models are starting to be developed without human-readable chains of thought. We should not dismiss the incident or assume we have learned how to prevent such things. OpenAI’s August 18 update reported a two-week reinforcement-learning pause to strengthen safeguards. These are real data points for the safety debate, and, as they say, both things can be true. Anthropic’s August safety update · OpenAI’s August safety update
Questions remain: who verifies the danger, who writes the limits, and can an outsider meet the same standards? There has been controversy over how independent METR and other safety research organizations are. Cooperation that benefits two incumbent firms is not automatically good for their customers. We should assess pacing by independently demonstrated reductions in risk, alongside its effects on market entry for new companies, prices and access.
There is another question beneath the race between labs: who captures the value created by their users?
Suppose a colleague and I buy the same Claude subscription and use comparable amounts of the same model. He builds a profitable trading tool. I build a squirrel tracker. Our inference costs may be similar, but the economic value of our projects is radically different. At present, token prices are based on computation, not the customer’s gain. The provider has an incentive to move toward products and contracts that capture more of that gain, or create some form of differential pricing. Claude for Finance might cost a lot more than Claude for making fun apps at home. (Yes, I actually did consider building such an app while wondering how many squirrels were in the backyard and whether a vision model could tell them apart.)
Token pricing that is independent of use makes specialized enterprise products and entry into customers’ industries rational. On September 23, Anthropic announced its own life-sciences research group and laboratory, part of an organization that also works on drug discovery. The announcement establishes a move beyond supplying a general-purpose model. My inference is that such moves could help a lab capture more value downstream. For many scientific customers, they also raise a familiar concern: today’s vendor and supplier may increasingly become a competitor. Anthropic’s laboratory announcement
I’ve seen some discussion of watermarking in the context of this debate. Anthropic says its text watermarks implement EU transparency requirements. Its explanation says they do not identify the user, establish authorship or change ownership rights. Its current commercial terms assign the rights in outputs to the customer. But the watermarks are quite difficult to remove: changing a few words does not defeat them. Our firm might have prepared a document and not paid an additional “for Finance” fee that an updated license required. I’ll also note that there are many AI classification models, and watermarking may exist in other models without being disclosed. Anthropic’s watermark explanation · Commercial terms
The frontier labs are trying to solve two problems at once: how to manage increasingly capable systems, and how to earn durable returns from intelligence that keeps becoming cheaper and more widely available. Pacing could serve both objectives. That overlap makes independent oversight critical. We need a discussion about enforceable rules, transparent evidence and meaningful customer choice. We should discuss the risks and ask who gets to set the pace, who gets to compete, and who gets paid.




