AI’s compute crunch: Why Kimi K3 and Fable 5 are both rationing access

Two highly touted AI models from this month have faced similar problems in quick succession, and there is nothing coincidental about that; it’s simply a lack of chips under a different name. While Kimi K3, the 2.8 trillion-parameter open-weight AI model from Moonshot AI, initially managed to out-benchmark a great number of other models upon release, after just less than 48 hours the creators decided to halt new subscription intake. The Chinese AI model had to suspend new subscriptions due to the fact that the overwhelming user demand brought it close to its capacity limits. Moonshot’s official account on X admitted that the model has gained more popularity than the company has expected, and their GPU clusters were having trouble keeping up. They are currently scaling their computing capabilities, and bringing new subscribers in smaller controlled batches instead of opening everything up at once. Current paying users are unaffected, but everyone who was planning to get in after the initial launch weekend is shut out by an impenetrable barrier of GPU shortages.

Also read: HuggingFace hacked: How RCE Dataset Loader exploited AI playground

But Anthropic’s new flagship, Claude Fable 5, which heads up its new Mythos tier, has had a far more chaotic set of weeks. It first debuted on June 9, then went offline on June 12 after export controls were put into place by the US government, returned on July 1 when the controls were removed, and has found itself engaged in a drawn-out process of negotiation with its own users about what the scope of access on the paid plans should be. It pushed the deadline for free-inclusion of usage from July 7 to July 12 to July 19, all with hours left to go before the deadline in question. For this week on, access to Fable 5 on both Max and Team plans will be limited to 50 percent of the maximum usage limit, with the company citing unanticipated demand. This is quite an admission for a firm that just secured a $45 billion compute deal with SpaceX, as well as potentially securing another $10 billion in computing power from Meta.

Also read: Rise of Friendslop games: Best couch co-op games you need to try

What unites these two stories is not the competition between Chinese and American AI, nor open vs. closed weighting. What is happening here is that the inference bottleneck has come into play, and not the training bottleneck. Each of the labs created their own impressive model; each got far more traction than was predicted by internal capacity planning; and each answered to that the same way – restricting new access, protecting the existing paying customers, and promising to do so only temporarily. For Kimi, the solution is to literally depend on open weights (which come due on July 27th) to allow developers to bypass Moonshot’s infrastructure completely.

For developers, what matters is not who’s winning. What matters is that at this point, no lab can give you guaranteed access to frontier-model inference at any fixed cost. If you’re planning to use Fable 5 or K3 in your product, the sensible strategy is to have a backup model for all the routine inference calls, and use the frontier model where it is really needed. Compute capacity, not inference ability, is the constraint on this generation of products.

Also read: Why your next phone will cost more and do less: State of the smartphone industry 2026

Vyom Ramani

A journalist with a soft spot for tech, games, and things that go beep. While waiting for a delayed metro or rebooting his brain, you’ll find him solving Rubik’s Cubes, bingeing F1, or hunting for the next great snack.

Connect On :