OpenAI engineers say they more than halved inference cost through software
OpenAI engineers say they found a way to more than halve inference cost using software alone, per The Information. In one case, powering ChatGPT for logged-out users needed only a couple hundred Nvidia $NVDA GPUs. It is an internal efficiency target, not a launched price cut.
OpenAI engineers told colleagues in June that they had found a way to more than halve the cost of running its models, and the gain comes entirely from software rather than new hardware, according to The Information, with coverage from Seeking Alpha and Germany's heise. The improvement is about squeezing more efficiency out of existing GPU servers. In one striking example, when the techniques were applied to power ChatGPT for visitors without an account, the number of Nvidia $NVDA GPUs needed at one point dropped to just a couple hundred. The specific methods have not been confirmed, though analysts point to familiar levers like quantization, key-value caching, batching and routing simpler tasks to lighter models. A caveat matters here: this is an internal efficiency claim, not a launched product or a price cut customers can see yet. Still, the implication is worth sitting with. Much of the AI industry's spending story has been about buying ever more chips. If a leading lab can cut serving costs in half through software alone, it suggests the economics of AI may hinge as much on optimization as on raw compute. That would change how investors judge the enormous infrastructure budgets these companies are defending, and how quickly cheaper inference reaches the products people actually use.