Notes On Lab Profitability
Chains and Networks
Note: This was originally submitted as part of Dwarkesh’s essay writing competition; as such, it’s a bit more formal than I typically write. I hope you enjoy!
The agentic paradigm will drive labs toward long-term structural profitability. Several trends are likely to increase frontier model providers’ lead on value per inference dollar, while reducing the value of distillation and increasing switching costs, creating durable profitability in the process.
First, the general structure. Simply put, models succeed by automating business tasks. Such tasks used to be relatively simple, but are growing increasingly complex. For example, over the last three years in the legal field, models have gone from barely being able to summarize a case to drafting reliably accurate simple motions. A more complex task typically requires a longer chain of steps to complete. One could reasonably characterize this longer chain as a series of more complex, but similarly numbered steps. However, most complex steps can be disaggregated into simpler subcomponents. For example, making Eggs Benedict requires either a single complex step—poaching an egg—or a series of simpler steps: grabbing a pot, filling it with water, turning the burner on, cracking the egg, etc. For simplicity, I call the least complex way to describe a step an atomic step. At a broad level of generality, success at step chains is defined by the number of 9s of accuracy that a model can generate per atomic step. Per-step accuracy compounds multiplicatively across the chain: a 99%-reliable model on a 50-step task succeeds about 60% of the time; a 99.9%-reliable model, about 95%. The return profile for any given 9 is therefore shaped by both diminishing gains in absolute reliability and the economic returns to any given unit of reliability—as per-step accuracy approaches 100%, the latter can rise even as the former falls, because each additional 9 unlocks longer feasible chains.
The shift from chatbots to agents is largely a shift toward these longer chains. To explain requires a brief digression into harnesses. A harness is the software that turns a base model into a functional agent. Claude Code is one, as is Codex. Inside the harness, the model runs in a loop, taking in an input, taking an action, feeding the output back into the model, and taking another action until a stop condition is reached. This loop requires context and memory management, verification and error recovery, and, increasingly, orchestration of subagents. This new paradigm advantages the leading models as efficient context window management is the name of the game. Notably, despite increases in price per input and output token, 5.5 is reportedly cheaper on a per-task basis. This per-task performance is useful for cheap tasks but critical for complex tasks. Because of the multiplicative compounding of failure rates, as long as the length of tasks demanded is not satisfied by the number of steps compressible into the effective context window, larger, more token-efficient models will dominate cheaper, distilled ones. Distilled models will simply be incapable of task completion and require human labor substitutes; a dramatically more expensive proposition on all relevant near- and medium-term margins. Additionally, as models are increasingly trained in a harness-specific manner, the relative effectiveness of distillation in producing the aforementioned 9s should further decrease.
As they are deployed, these long-chain agents will compound their advantages through unique access to real-world deployment training data. Thus far, most training has been conducted through benchmarks, web text, corpuses of created problems, etc. With Codex and Claude Code, the main downside of these data sources (their inherently limited correlation with real-world tasks), can be mitigated by training on enterprise data. Every day, millions of people, through their actions, are telling coding agents which responses matched with their requests, which need to be modified, which failed and needed a reset, etc. This is a goldmine for whoever decides to seize it. Currently, both enterprise Codex and Claude Code default away from training on these tasks, but expect to see carrots (priority access to frontier models, discounts on compute) used aggressively as the value of this data grows with scale. And, because distillation requires training on a set of problems similar to those the model is being used to solve; the ability to distill will be limited by the inability to effectively recreate a similar business environment that led to the valid answer, further decreasing distillation’s effectiveness in matching the 9s of the leading models.
Where there is genuine competition, then, is not from distillation, but from Google and Microsoft, which have much more trusted relationships with a substantially greater number of enterprise users; potentially allowing them to get a jump on the labs in enterprise integration despite worse model-harness sets to start.
There are two reasons I suspect the labs will overcome these competitors. First, given the rapidity with which agent quality is improving, the option facing most companies will not be: “Can I get 80% as effective a product for 20% of the price?” Rather, it will be “Can I wait 9 months to get 80% as effective a product for 20% of the price, at which point my competitor will get access to entirely new capabilities with their frontier model?” Making the second pitch will, reasonably, be seen as unwise. As a result, the labs are near-certain to win the first enterprise sales cycle.
The question then becomes whether model providers can keep switching costs high enough to maintain healthy margins. Beyond traditional enterprise switching costs, I believe agentic implementation will create extraordinarily high switching costs. This portion is more speculative, but I contend that—as they are deployed—agents will create internal networks between each other and with other employees. As described by Metcalfe’s law, these webs of context will drive a quadratic increase in value. Much like social media providers, labs will have little-to-no incentive to improve interoperability within such networks with other model providers. That means the relevant margin of replacement for an enterprise considering switching model providers will not be the relative cost of a new model to replicate tasks, but rather the cost to replace the entire network. As the value of that network rises, so will the size of the moat and relevant pricing power.
In sum, such networks are how the labs could win.



