Earlier this month, @OpenAI engineers developed an optimization that more than halves inference costs for the models it…
By Mark Kretschmann · AI Infra
Earlier this month, @OpenAI engineers developed an optimization that more than halves inference costs for the models it was applied to. When deployed on logged-out ChatGPT traffic, it reduced the number of GPUs needed to power that traffic to just a couple hundred. Reporting from The Information.