Why do everyone assume they are subsidized? When we seemingly have no idea what it costs? Maybe average subscription is breaking even and token spend is pretty much pure profit?
Case in point, claude code seems hell bent on increasing usage at all cost. Which makes sense in the growing phase (get people hooked) but it does not make sense given the hardware shortage. So, which is it?
Anthropic admitted last year to losing money on inference, it had negative margins. The margins have improved and are now positive but there’s still a significant cost. If plans aren’t being subsidized it would mean that the margin on inference is ~99%+ which would mean OpenAI and Anthropic should be wildly profitable but both are still losing money. So, it’s mathematically impossible that they’re not subsidizing plans.
The most widely accepted estimates (though I disagree with them) are that Anthropic’s margin on API inference is ~70% from which people extrapolate what their token usage would cost via the API and compare that to what their plan costs.
re: increasing usage with resets, it’s because they’ve overblown usage and need to show that usage is growing ahead of the IPO. They’re increasing usage on fixed price plans without increasing the cost, the only plausible explanation is they have unused capacity. If they were capacity constrained then the last thing they would do is give away more usage for free.
The last I saw with Claude was that the plans are subsidized somewhere around 10x. The usage of Claude and similar products (for people not paying the actual API token costs) would be lot less if they were paying 10x more per account/seat.
A lot of people are using Claude and ChatGPT for all kinds of minor things at work, and they probably wouldn't be if they were paying the true cost of the product. And this is all the while their work product is suffering because AI is not a great fit for a lot of use cases.
Depends on whether the costs they have are inflated ten-fold due to hardware shortages.
Ignoring current state of shortages, the interesting part (for me at least) is how much a SOTA token costs on current hardware if you exclude the costs for paying for current TSMC shortage and without factoring in the cost of new datacenters being rushed and renting hardware from your competitors (who are also supply limited).
The "actual" cost is of course also relevant and interesting but it is another thing entirely.
Have they actually announced that they're charging less than the marginal cost of inference, or that the revenue they're taking in doesn't make up for the cost of training each newer and better model?
The former does not pass the smell test when there are random providers selling tokens for competitive open weight models.
Case in point, claude code seems hell bent on increasing usage at all cost. Which makes sense in the growing phase (get people hooked) but it does not make sense given the hardware shortage. So, which is it?