Two releases in a week aimed at agent token costs, which tells you where the pain is
Gemini 3.6 Flash and NVIDIA's Nemotron 3.5 Lightning both landed with always-on agents and their running costs as the headline.
Source: Oracle ↗
- AI economics
- AI agents
- Inference cost
- Model selection
Within a week, Gemini 3.6 Flash arrived with enterprise agent token costs as its stated focus, and NVIDIA's Nemotron 3.5 Lightning became available on Oracle's cloud pitched at always-on agents. When two releases target the same problem in the same week, the problem is the story.
The constraint moved
For two years the question was whether a model could do the task. That question is mostly settled for the work most companies want automated. The live question is what it costs to run continuously, because an agent that watches a queue all day makes calls a chatbot never did.
What I'd do about it
Instrument cost per outcome before you scale anything, not after. The number that matters is not cost per token or per call, it is cost per invoice matched, per ticket resolved, per order routed. Teams that skip this discover the economics on an invoice, usually one quarter after the pilot everyone was pleased with.
It is also worth designing for model substitution from the start. Two releases in a week is the tempo now, and a system welded to one provider cannot take advantage of it.
