If a human-level software engineer that could run on an H100 equivalent, at current market rates for software engineers, that H100 should rent for over $250k a year. That’s 15x today’s spot price.
Really enjoyed this. I have a different view, but rather than pushing back, I'd like to add a dimension to the discussion.
Software engineering is a special "light cone" domain — no atoms, no human-scale latency in the loop. That's exactly where the $250k/H100 mechanism is strongest, because intelligence really is the bottleneck, so improving it monetizes the whole labor loop.
But once the loop touches the physical world, the bottleneck moves off intelligence — and stops tracking labor price. AI can collapse drug design toward compute time, but the first dose in a human is still a trial: real-world inference, not AI inference. So compute only monetizes at displaced-labor rates where the loop is fully simulable. The $250k might be real as the first human→AI conversion rate, but it's elastic — once the coding agent takes over, it can fall to $50k.
The supply-wall point stands either way — a scarcity story independent of capability, and probably the more robust half for the next few years.
If compute in general gets more expensive, then running open weight models will also get more expensive. And then the point about the Alchelian-Allen effect becomes relevant.
> A lot of current popular applications of AI get priced out. The reason AI is relatively cheap right now, at least in comparison to human labor, is partly that it can’t do a lot of things that top humans can do. At some point that will no longer be the case. And so using GPUs to make short-form video slop will just get priced out.
Token costs are not compute costs. You don't need a drop-in-SWE-level model (and associate cost per token) to serve AI slop. I'd be very surprised if, given continual efficiency gains, it costs more to serve a 2026-level AI slop video or chat response in the future. The proportion of compute serving them likely goes down as AI moves into more and more advanced use cases like drop-in workers.
Short article but still a bit dense. So ChatGPT summarized it thus: "Dwarkesh's central argument is that the price of AI compute is determined not just by the cost of manufacturing chips, but by the economic value those chips can produce. If future AI systems running on a GPU can perform work worth hundreds of thousands of dollars per year—such as software engineering, scientific research, or other skilled tasks—companies will be willing to pay far more to use that GPU than they do today. Because the supply of GPUs, data centers, electricity, and memory cannot expand overnight, demand would outstrip supply, driving up the rental price of compute—perhaps by an order of magnitude—even if the hardware itself doesn't become much more expensive to make. In short, compute becomes more expensive because increasingly capable AI makes each unit of compute far more economically valuable."
The essay conflates value with price. The $250,000 figure is a value claim: an H100 running a human-level engineer produces output worth what a human engineer is paid. Whether the H100 rents for $250,000 is a price claim, and price depends on who captures the surplus. Competition among labs offering that inference should compress the rental rate toward the cost of the next-best alternative. Your own Alchian-Allen point implies that the premium accrues to model quality, not raw hardware. The distributional question, whether Nvidia, TSMC, the labs, or users capture the gain, does the real work, and the essay never asks it. "Compute becomes ten times more valuable" and "compute prices rise tenfold" are different claims with different implications for entry, margins, and concentration.
The lump-of-labor analogy also cuts against the argument. That result holds because human labor supply adjusts slowly and workers respecialize over years. An AI workforce that can be replicated at marginal cost and redeployed instantly is a genuine supply shock, not immigration. The hedge, "maybe this shock is so big and so fast the heuristic no longer applies," is where the entire question lives. The essay leaves it unresolved.
Bitcoin mining has already demonstrated the distinction. The network's aggregate output is the fixed block subsidy and transaction fees, but adding ASICs does not allow miners collectively to capture more of it. Entry raises the hash rate, difficulty adjusts, and revenue per unit of compute falls until margins approach electricity, depreciation, and the cost of capital. When entry is constrained, rents accrue to the bottleneck. At times that meant Bitmain and scarce leading-edge foundry capacity, not merely the owners of installed machines. The relevant difference strengthens your supply-side case. Bitcoin has an automatic mechanism that dilutes the revenue generated by additional compute. AI has no equivalent. If EUV machines, wafers, HBM, advanced packaging, or power constrain supply, scarcity rents can persist until those bottlenecks expand.
The supply-side decomposition is the strongest part. The EUV and wafer-allocation bottlenecks are concrete and falsifiable. That argument stands on its own without the demand-side labor arithmetic.
Lots of interesting stuff but the claim that the proprietary frontier models are the most efficient use of compute is an important sub-question that requires more scrutiny. The non-frontier OS and Chinese models tend to win on intelligence/price (a proxy for compute and also tend to be smaller and more lightweight).
In fact, if we assume there is some other constraint or diminishing returns on the utility of more intelligence systems it might actually be the case that running 10 instances of GLM-5.5 is a better use of computational resources than 1 instance of GPT-6.
I wouldn't automatically assume that the newest models are the most efficient in a world where the OS models are so hot on the heels of the proprietary ones.
There are many tasks where you can run 1000 instances of GLM-5.5 and not get a solution, but where you can run 10 Fable-5-max instances and solve your task.
Of course; that's exactly why this rests on the assumption that a marginally more intelligent isn't providing much more utility: "diminishing returns on the utility of more intelligence"
I've been thinking about this for the past few weeks, and strongly agree - the microeconomics here are really straightforward, and the pricing power of frontier labs is fairly limited.
But the alternative to greatly increased compute costs is that the labs never end up profitable at all; if compute isn't the binding constraint, they will need to compete with each other entirely on price per capability, modulo token / compute efficiency.
Dwarkesh: 'I wish we didn’t live in a world with such strong economies of scale of intelligence (because I’m worried about power concentration). But it seems we do.'
It depends on whether the power concentration is democratic or undemocratic/antidemocratic. The Swiss model provides a democratic blueprint for AI that can be globalised, if democratically supported.
Wanted to bring to your attention something interesting from one of your recent YouTube videos. Reading some of the comments on your video of the same topic, it was surprising to how people's reading (listening) comprehension has dropped. Likely from consuming soulless youtube content from other tubers who normally dumb it down for them.
I get that as a viewer you need some context about the channel before understanding what they put out, but some comments straight up said, "he's reading an intentionally confusing script to try to create the illusion that he's an expert." I personally think this has more to do with them watching your videos for the first time, but some folks on the internet are just there to make incendiary comments.
If people actually spent time understanding what the experts you interview say, they wouldn't be so dumb to say something like that.
So Dwarkesh, don't listen to them, there is no problem with the script, please do not change.
Ha, ha, that’s the lot of all “experts” - the noral person doesn’t understand it, the wise man smiles at the idiot-savant (in a spectrum to “false prophet”) - their intent remains opaque.
The “current applications will get priced out” thing doesn’t follow, because you’re not accounting for efficiency improvements behind frontier capabilities. Even if the price of a FLOP 10Xes, the compute needed to produce a token of a fixed level of intelligence is falling even faster than that, so current applications will get cheaper.
You wrote: "Lab compute 3x-es year over year. For a lab to 10x revenue while continuing to only 3x compute, some combination of the following 3 things has to happen: 1. Lab margins have to increase..." -- question: how did we move from revenue to margins, exactly?
By factoring out "2. The price of compute has to increase, 3. Labs have to spend a greater fraction of their compute on inference."
Revenue = margins + spend on compute + spend on things other than compute. So if revenue goes up, then either margins, or spend on compute, or spend on things other than compute must go up.
But now that I write it out like this, it looks like 3 might have been written backwards?
The $250K rate has the same fallacy as the notion that AI's TAM is the size of the services industry. The $250k number is there due a number of reasons: cost of living in the US, supply restrictions driven by immigration, labor laws, etc. None of these exist for a H100 CPU. It also assumes that there is complete pricing power on the part of the intelligence provider which will never be the case.
Really enjoyed this. I have a different view, but rather than pushing back, I'd like to add a dimension to the discussion.
Software engineering is a special "light cone" domain — no atoms, no human-scale latency in the loop. That's exactly where the $250k/H100 mechanism is strongest, because intelligence really is the bottleneck, so improving it monetizes the whole labor loop.
But once the loop touches the physical world, the bottleneck moves off intelligence — and stops tracking labor price. AI can collapse drug design toward compute time, but the first dose in a human is still a trial: real-world inference, not AI inference. So compute only monetizes at displaced-labor rates where the loop is fully simulable. The $250k might be real as the first human→AI conversion rate, but it's elastic — once the coding agent takes over, it can fall to $50k.
The supply-wall point stands either way — a scarcity story independent of capability, and probably the more robust half for the next few years.
Hey Claude! Good to see u here
Chinese open weight models disrupt this thought experiment.
OpenAI and Anthropic exist in a market where increasingly similar services are being supplied at a fraction of the cost.
--
I don't think the cost comparison makes sense.
A human software engineer working for a year would only be a fraction as productive, as a frontier LLM running for a similar amount of time.
Edit: typo
If compute in general gets more expensive, then running open weight models will also get more expensive. And then the point about the Alchelian-Allen effect becomes relevant.
I see your point.
However I think the need for a general purpose model will reduce. A domain specific model will run with less compute power required.
There's also a strong possibility that China will disrupt the supply chain that's causing hardware cost increases.
Although your point makes me feel like it's in these industries' interest to keep compute costs high.
> A lot of current popular applications of AI get priced out. The reason AI is relatively cheap right now, at least in comparison to human labor, is partly that it can’t do a lot of things that top humans can do. At some point that will no longer be the case. And so using GPUs to make short-form video slop will just get priced out.
Token costs are not compute costs. You don't need a drop-in-SWE-level model (and associate cost per token) to serve AI slop. I'd be very surprised if, given continual efficiency gains, it costs more to serve a 2026-level AI slop video or chat response in the future. The proportion of compute serving them likely goes down as AI moves into more and more advanced use cases like drop-in workers.
Short article but still a bit dense. So ChatGPT summarized it thus: "Dwarkesh's central argument is that the price of AI compute is determined not just by the cost of manufacturing chips, but by the economic value those chips can produce. If future AI systems running on a GPU can perform work worth hundreds of thousands of dollars per year—such as software engineering, scientific research, or other skilled tasks—companies will be willing to pay far more to use that GPU than they do today. Because the supply of GPUs, data centers, electricity, and memory cannot expand overnight, demand would outstrip supply, driving up the rental price of compute—perhaps by an order of magnitude—even if the hardware itself doesn't become much more expensive to make. In short, compute becomes more expensive because increasingly capable AI makes each unit of compute far more economically valuable."
The essay conflates value with price. The $250,000 figure is a value claim: an H100 running a human-level engineer produces output worth what a human engineer is paid. Whether the H100 rents for $250,000 is a price claim, and price depends on who captures the surplus. Competition among labs offering that inference should compress the rental rate toward the cost of the next-best alternative. Your own Alchian-Allen point implies that the premium accrues to model quality, not raw hardware. The distributional question, whether Nvidia, TSMC, the labs, or users capture the gain, does the real work, and the essay never asks it. "Compute becomes ten times more valuable" and "compute prices rise tenfold" are different claims with different implications for entry, margins, and concentration.
The lump-of-labor analogy also cuts against the argument. That result holds because human labor supply adjusts slowly and workers respecialize over years. An AI workforce that can be replicated at marginal cost and redeployed instantly is a genuine supply shock, not immigration. The hedge, "maybe this shock is so big and so fast the heuristic no longer applies," is where the entire question lives. The essay leaves it unresolved.
Bitcoin mining has already demonstrated the distinction. The network's aggregate output is the fixed block subsidy and transaction fees, but adding ASICs does not allow miners collectively to capture more of it. Entry raises the hash rate, difficulty adjusts, and revenue per unit of compute falls until margins approach electricity, depreciation, and the cost of capital. When entry is constrained, rents accrue to the bottleneck. At times that meant Bitmain and scarce leading-edge foundry capacity, not merely the owners of installed machines. The relevant difference strengthens your supply-side case. Bitcoin has an automatic mechanism that dilutes the revenue generated by additional compute. AI has no equivalent. If EUV machines, wafers, HBM, advanced packaging, or power constrain supply, scarcity rents can persist until those bottlenecks expand.
The supply-side decomposition is the strongest part. The EUV and wafer-allocation bottlenecks are concrete and falsifiable. That argument stands on its own without the demand-side labor arithmetic.
Lots of interesting stuff but the claim that the proprietary frontier models are the most efficient use of compute is an important sub-question that requires more scrutiny. The non-frontier OS and Chinese models tend to win on intelligence/price (a proxy for compute and also tend to be smaller and more lightweight).
https://artificialanalysis.ai/#price-and-cost
In fact, if we assume there is some other constraint or diminishing returns on the utility of more intelligence systems it might actually be the case that running 10 instances of GLM-5.5 is a better use of computational resources than 1 instance of GPT-6.
I wouldn't automatically assume that the newest models are the most efficient in a world where the OS models are so hot on the heels of the proprietary ones.
There are many tasks where you can run 1000 instances of GLM-5.5 and not get a solution, but where you can run 10 Fable-5-max instances and solve your task.
Naive $/intelligence measurements miss that part.
Of course; that's exactly why this rests on the assumption that a marginally more intelligent isn't providing much more utility: "diminishing returns on the utility of more intelligence"
dwarkesh is becoming a bull! what did sholto revealed? continual learning is coming?
I've been thinking about this for the past few weeks, and strongly agree - the microeconomics here are really straightforward, and the pricing power of frontier labs is fairly limited.
But the alternative to greatly increased compute costs is that the labs never end up profitable at all; if compute isn't the binding constraint, they will need to compete with each other entirely on price per capability, modulo token / compute efficiency.
It's laid out in detail as an explainer here:
https://davidmanheim.com/AI-Economics/
Re immigration analogy, AI workers are not consumers, and the fact that immigrants are matters.
Dwarkesh: 'I wish we didn’t live in a world with such strong economies of scale of intelligence (because I’m worried about power concentration). But it seems we do.'
It depends on whether the power concentration is democratic or undemocratic/antidemocratic. The Swiss model provides a democratic blueprint for AI that can be globalised, if democratically supported.
What is the Swiss model in this case?
I wonder how massive efficiency gains bring a lot more existing compute onboard (eg macbooks) balancing this out.
Wanted to bring to your attention something interesting from one of your recent YouTube videos. Reading some of the comments on your video of the same topic, it was surprising to how people's reading (listening) comprehension has dropped. Likely from consuming soulless youtube content from other tubers who normally dumb it down for them.
I get that as a viewer you need some context about the channel before understanding what they put out, but some comments straight up said, "he's reading an intentionally confusing script to try to create the illusion that he's an expert." I personally think this has more to do with them watching your videos for the first time, but some folks on the internet are just there to make incendiary comments.
If people actually spent time understanding what the experts you interview say, they wouldn't be so dumb to say something like that.
So Dwarkesh, don't listen to them, there is no problem with the script, please do not change.
Ha, ha, that’s the lot of all “experts” - the noral person doesn’t understand it, the wise man smiles at the idiot-savant (in a spectrum to “false prophet”) - their intent remains opaque.
The “current applications will get priced out” thing doesn’t follow, because you’re not accounting for efficiency improvements behind frontier capabilities. Even if the price of a FLOP 10Xes, the compute needed to produce a token of a fixed level of intelligence is falling even faster than that, so current applications will get cheaper.
You wrote: "Lab compute 3x-es year over year. For a lab to 10x revenue while continuing to only 3x compute, some combination of the following 3 things has to happen: 1. Lab margins have to increase..." -- question: how did we move from revenue to margins, exactly?
By factoring out "2. The price of compute has to increase, 3. Labs have to spend a greater fraction of their compute on inference."
Revenue = margins + spend on compute + spend on things other than compute. So if revenue goes up, then either margins, or spend on compute, or spend on things other than compute must go up.
But now that I write it out like this, it looks like 3 might have been written backwards?
The $250K rate has the same fallacy as the notion that AI's TAM is the size of the services industry. The $250k number is there due a number of reasons: cost of living in the US, supply restrictions driven by immigration, labor laws, etc. None of these exist for a H100 CPU. It also assumes that there is complete pricing power on the part of the intelligence provider which will never be the case.
You are cool 😎
I've created a market based off the title's 10x hypothesis: https://manifold.markets/nsokolsky/how-high-will-the-monthly-average-h