One of the amazing things is that when every has one GPU, they will actually have 1k-10k agents at their disposal.
LLMs and KV caches have amazing performance characteristics with concurrent throughput. It scales very non linearly. So the token throughput within a batch scales WAY faster than the tokens per second of each user.
This is the reason the LLM providers have such crazy margins on their costs.
They had incredible revenue growth the last few years and just broke 100M in revenue. I don't know what their internal spend was , but that was almost half of their recent round in ARR. Mostly likely they were profitable or on a clear trajectory to revenue growth. Huggingface hosts a lot of data and models, but mostly static cold storage is pretty cheap tbh.
This is a wildly incorrect and myopic view on the world.
Finetuning model is cheap and incredibly useful for deployment. You don't need to pre-train a frontier llm from scratch to make useful models.
There is tons of domains where you and fine-tune llms and deploy them for value in companies and for your own entrepreneurship ambitions. I have made this a big part of my career for the last few years and now I'm working on finetuning models for starting my own companies.
Both of you are right. There is demand for tailored (fine-tuned) models; almost every enterprise would theoretically benefit from them.
But there are also a lot of prerequisites, namely does the enterprise have its sh*t together on a technical level. Does it have the processes and data pipelines available to train and benefit from these models? Probably not!
Applied ML is at the crown of a tech pyramid whereas most enterprises are still struggling at ground level. Being able to build from be ground is likely a safer skillset than only knowing how to work at the (non-existent) apex.
I find the fine tune approach more interesting than straight to RAG and MCP.
End of the day they're all customized data stores and protocols to interact with them. May as well stick to a uniform toolkit with fine-tunes.
Not that other tools aren't useful. But reaching straight for a bunch of infrastructure reliant services is like jumping in with k8s when you're still at a stage where basic mocks in code are sufficient.
I won't roll my own encryption or UI lib but want to stay focused on the incompleteness of the project I have to ship not all the buttons and knobs of some dependency or framework. Same old manage context switch problem.
I think this is largely true. Theoretically it follows that mature industries would become this way more and more.
In my software engineering jobs there was always strong top down product direction. In my AI research jobs people are often staring at me blankly waiting for me to tell them what we can do.
A young industry the only people who really understand the potential of what the technology can do are the experts and practioners.
Overtime all the best ideas get picked off, and the general population who are non practitioners gain enough knowledge that they can understand what the technology can do better than the experts how can impliment it.
I think a lot of it is just time. The quality of a model is E * C
Where:
E = Efficiency, and efficiency gains come from quality of data, quality of algorithms.
C = Compute (Size of model, flops of train run)
So a better company can train a bigger and better model with less required compute which let's anthropic get there first. If another company does the same thing with a worse: model architecture, kernel, optimizer, etc... They will get there as well if they just run there train run with more flops for longer
Mythos was actually ready about 6 months ago. So if you have 6 months later or hardware setup and time to train you can get a lot done.
FWIW it is entirely possible to square the notion that small models will still be hosted on cloud hardware with the idea that the data centre buildout will end in tears.
Many analysts (and Microsoft) think even now that if everything committed gets built there will be considerable oversupply and there is not the revenue to pay for it.
If small models do continue to improve in unusual ways (I think there are limits) then the marginal need for cloud AI compute could fall precipitously beyond current estimates. The marginal need for consumer AI could almost totally collapse if someone makes good progress on very small reasoning and tool-calling models (which is a modestly big if)
The possibility of the data centre boom resembling the Chinese real estate bubble is not inconsiderable.
So far everyone seems to be consistently GPU-poor, despite the huge buildout, and usage keeps going up drastically. I don't know what would make usage drop.
Every time they've made smarter models we've wanted the smarter ones, and local models runnable on typical hardware are still very far behind in speed and intelligence (as neat as they are)
> So far everyone seems to be consistently GPU-poor, despite the huge buildout
That is seemingly not the case. The buildout is actually slow; almost nothing of these giant projects has been completed. Nobody will say how much of anything they have actually finished. And Nvidia have made huge, huge buy-and-hold deals for GPUs that do not have data centres to go into.
Everyone is GPU poor because stuff hasn't been finished but large numbers of GPUs are spoken for, but they are GPU poor on therefore much less demand than is being built for.
Look at how tiny SpaceX's deal is with Anthropic, for example. This meaningfully turned around Anthropic's prospects — allowing them to radically lift rate limits beyond what many users needed -- but it was for just 300 megawatts. Tiny compared to the 31 gigawatts allegedly under construction by the end of last year.
So the picture is partly illusory. GPU prices and RAM prices have been pushed up by the AI firms booking them for data centres they haven't even started building yet, as well as the ones that they've only completed a tenth or an eighth of.
There will be significant oversupply. And if open weights models keep getting good and staying fuel-efficient, that picture gets worse.
If I understand correctly, you're saying people are compute-poor but not necessarily GPU-poor because there's a lot of GPUs out there but nowhere to plug them into? If so, I'm not sure that distinction matters to the GP's point that there is too much demand to call this an oversupply.
I also doubt we can estimate the level of demand based on a single deal between Anthropic and SpaceX (despite which, note, Claude still stuggles at times.) Consider other signals, like Google, who we thought had an insurmountable infra advantage, also renting compute capacity from SpaceX and limiting Meta's usage (along with other clients apparently) to conserve capacity: https://www.cnbc.com/2026/06/28/google-limits-metas-use-of-i...
I am not sure Microsoft thinks there will be an oversupply either; last earnings they announced bumping up their CapEx spend, along with all the other hyperscalers.
Here's a way to estimate how much room there is for demand to grow. Various sources (linked in this comment, along with more analysis: https://news.ycombinator.com/item?id=49089296) indicate that even though a large number of people (50 - 60%) are now using AI at work, they use it for only 6% of their work hours.
That means, even if AI can only address 30% of all work, there is still 5x potential demand growth left! Note, the sources above indicate that AI is even being used in non-knowledge work industries, so the scope is already larger than we thought. This is in addition to the remaining 40 - 50% of people are still not using AI at work. Plus we know that agentic workloads consume way more tokens, so that's yet another multiplier.
But will that demand keep growing? Well, some of those same sources above mention that most executives are planning on ramping up their AI spend in coming years.
Putting all this together explains the hyperscalers' quarterly bemoaning of how strapped for compute they are and why they are spending so much to add more capacity. Given this, an oversupply seems pretty unlikely.
> If I understand correctly, you're saying people are compute-poor but not necessarily GPU-poor because there's a lot of GPUs out there but nowhere to plug them into? If so, I'm not sure that distinction matters to the GP's point that there is too much demand to call this an oversupply.
It's a distinction without a difference if your issue is getting hold of a GPU.
But there's a significance to it if you are trying to use demand for a GPU as a proxy for demand for AI. That is where the industry is making serious mistakes.
> I am not sure Microsoft thinks there will be an oversupply either
The rest of your comment I am not going to address because it's kind of unfalsifiable. Hypotheses about what AI might be able to do in principle aren't all that useful when talking about even medium-to-long-term demand for what LLMs and GANs can provably do now.
Oh I think I recall this. However Satya said that more a year and a half ago, eons in AI time, and maybe he even believed it at the time. But since then Microsoft's annual CapEx spend has gone from $65B to $116B. Not the actions of somebody concerned about an overbuild.
> But there's a significance to it if you are trying to use demand for a GPU as a proxy for demand for AI. That is where the industry is making serious mistakes.
Ah I see what you're saying. The AI industry is certainly not looking at just GPU demand as a proxy for AI demand, primarily because they themselves are creating the GPU demand -- but that is in response to other signals they are getting. Actually that's what I hoped to illustrate with the rest of my analysis; rather than a hypothesis of what AI might be able to do, it was more a hypothesis of the signals the industry is looking at (including various economic data sources) that drive its current spending frenzy.
When the glut of GPU arrives I'm sure humanity will find a good use for all that excess compute, like finally getting back to signing monkey pictures and excreting endless hash based pyramid schemes.
It may not make financial sense for someone retired, not into tech, and/or data privacy to host their own LLMs. However if usage of AI in day to day lives continues to increase, I think it will eventually make sense for the majority.
Many tasks suited for AI assistants are background asynchronous tasks. They can run in the downtime where immediate demand is low, keeping overall utilization high enough.
Your argument is similar to those who argue that owning a GPU for gaming doesn't make sense when you can stream from something like GeForce Now. However like with gaming locally (improved latency) there are also benefits to local AI (data privacy).
The power of small models isn't only that you can run them on local hardware. You can also fully own your data and workflow, and choose/fine-tune models for your specific use-case.
for clarity, I'm not agreeing with GP that small models will mean doom for data center projects
> Small llms are still way more efficiently server on big GPUs.
Yes, but the privacy aspect means that for many, many applications slower local will still be preferable to faster remote so long as the actual model performance is the same.
Yes latency, and the usual preference of ownership over rentership. Similarly their are benefits to running local AI too, like data privacy and control.
yea, a lot of my prior work was in the "code is the easy part" I was a frontend engineer for years. And something like 90% of my job the code was not the hard part.
I loved writing GPU shaders or optimizing visualization performance, but most of the time it was wiring up netcode to UI elements that exist.
Ironically as I've moved into focusing on more GPU and kernel programming AI is now lapping me there anyway, however the impact of knowing what sort of algorithsm are state of the art in papers, what is causing memory bandwidth issues etc... does a lot to drive the machine.
reply