OpenAI’s applied AI engineers told developers to judge models on cost per task rather than the list price per token, in a DevDay 2026 session that OpenAI posted to YouTube on 7 October 2026. “The metric that developers and businesses should be looking at is cost per task not cost per token,” said Mandeep, who leads OpenAI’s applied AI engineering team in San Francisco.
The panel backed the advice with customer results reported on stage. The largest: Blitzy, an AI software development company, reported 87% cheaper cost compared with GPT-5.4 Mini after moving to Luna’s tool-calling loop, with cache reuse rising from 24% to 90%.
Business Pill · PAY FOR THE JOB, NOT THE WORD
A one-minute explainer of the idea behind this story: cost per task. It teaches the concept, not this story’s figures.
The key insight: As we read it, OpenAI is asking customers to stop comparing price lists. Its engineers argue a more expensive model per token can be the cheaper one per task, which moves the comparison onto ground each customer has to measure for itself.
The Customer Numbers
According to the panel, Perplexity reported that GPT-6 Astra produces 9% more accurate results at half the cost on its research benchmark, compared with the previous model. Notion reported better accuracy and half the cost per task after migrating from GPT-5.5 to GPT-5.6 Sol.
Clio, which makes cloud software for law firms, reported 38% fewer prompt tokens in multi-step document analysis by using programmatic tool calling, with no loss of quality, the panel said.
For Blitzy, the panel added that across thousands of production calls the company reported “8.5x fewer output prompt tokens” and handled 2.2 times more context.


Why the Token Price Misleads
Mandeep gave two reasons a model that is cheap per token may not be cheapest per task: a cheaper model may consume or produce more tokens to finish, or it may not finish at all and a human has to step in.
His example was an airline support chatbot that could not rebook his flight, so a human had to complete the task. The cost of that task, he said, was the tokens plus the human’s time.
The panel described Perplexity’s case the same way: GPT-6 Astra may cost more per token, but uses fewer tokens and gets the task done.
The Four Levers
Beyond choosing a model, the panel listed four cost levers: prompt caching, programmatic tool calling, reasoning effort, and batch or flex processing.
On caching, the moderator said OpenAI’s price list shows up to 90% lower costs for cached inputs, and that reasoning effort can now be changed in the API without breaking the cache, which the moderator said was released with Astra. OpenAI has also shipped a prompt-caching dashboard and a diagnostics API that shows where cache misses happen, according to the panel.
On batch processing, the panel said OpenAI’s batch APIs offer a 24-hour processing window at 50% lower token pricing than standard synchronous requests.
On model choice, Mandeep said customers report that Luna works for 80 to 90% of their workflows, and recommended testing a smaller model early for simple tasks such as classification.
The Structural Read
The unit of comparison changes. The panel said a cheaper model may use more tokens, or fail and need a human, so the cost that matters includes the work around the model.
Most of the levers sit outside the choice of model. Of the four the panel listed, programmatic tool calling and reasoning effort cut the work sent to the model, while prompt caching and batch processing cut the price paid for it.
The figures are customer-reported and different in kind. Blitzy’s 87% is a cost comparison with GPT-5.4 Mini, Clio’s 38% is a token count, and the panel gave no absolute costs.
Mandeep, OpenAI applied AI engineering, DevDay 2026
“So the metric that developers and businesses should be looking at is cost per task not cost per token.”
Three Implications
MEASURE THE TASK The panel told developers to define the task and its accuracy threshold first, then find the cheapest configuration that meets it.
CACHE WHAT REPEATS The panel said cached inputs cost up to 90% less on OpenAI’s price list.
WAIT WHEN YOU CAN OpenAI’s batch APIs offer 50% lower token pricing within a 24-hour window, the panel said.
The Business Engineer Lens
This story maps onto the Business Engineer framework Tokenomics: The Economics of AI.
The framework’s starting point: “It is simultaneously the unit of cognition the model produces, the unit of compute the data center serves, the unit of price the lab charges, and the unit of value the enterprise extracts.”
As we read it, OpenAI’s panel separates the last two: the token stays the unit of price the lab charges, while the engineers tell customers to count value per task, where fewer tokens and fewer human hand-offs can outweigh a higher price per token.
What Is Not Established
We read OpenAI’s uploaded transcript of the session. The customer figures are as reported by OpenAI staff on stage; we did not read the customers’ own accounts, and the panel did not give the benchmarks’ details or absolute costs. We did not contact OpenAI or the customers.
The session was recorded at DevDay on 29 September 2026 and posted on 7 October.
The Bottom Line
OpenAI’s engineers told developers to compare models on cost per task, not per token, and cited customers on stage: Blitzy reported 87% cheaper cost than GPT-5.4 Mini on Luna, Perplexity half the cost with GPT-6 Astra, Notion half the cost per task on GPT-5.6 Sol, and Clio 38% fewer prompt tokens.
94,000+ executives read Business Engineer for the AI strategy frameworks cited by ChatGPT, Claude, and Perplexity.
A note on sourcing. We read the full transcript of OpenAI’s DevDay 2026 session as posted to its YouTube channel on 7 October 2026. Customer results are as reported on stage by OpenAI staff; we did not read the customers’ own accounts. We did not contact OpenAI or the customers. Nothing here is a forecast, and nothing here is financial or investment advice.
Sources: OpenAI on YouTube: Stop Overpaying for Intelligence, DevDay 2026 (posted 7 Oct 2026)









