Attending Workday Rising? Connect with us at the event.

Tokens, credits, and capacity limits make today’s AI pricing feel like a fact of life. It isn’t. Inference costs keep falling, and a new crop of cheap, credible models is already pulling the rug out from under the current pricing playbook. The real challenge for leaders is getting value now without getting locked into a meter — or a set of assumptions about which model will still matter next year.

AI Economics

The leadership decision: Invest in what pays off at today’s prices, but keep the freedom to switch providers, models, and contract terms as the market keeps shifting.

To a business leader sizing up an AI investment, a rate card can look a lot like a forecast. Tokens, credits, task-based units — these give Finance something concrete to budget against and Procurement something they can compare line by line. That precision feels reassuring. It might also be the shakiest part of the whole decision.

Three things are moving at once while you’re being asked to commit: what it actually costs to run these models, how providers choose to package that cost, and how much compute is even available to meet demand. So the real question isn’t which plan is cheapest this quarter. It’s whether you move now, wait for prices to keep falling, or risk building your organization around today’s pricing unit before anyone — including the vendors — has figured out what customers are really paying for.

The early mobile phone market is a decent point of comparison. It doesn’t mean AI is headed toward free or unlimited everything. But it does show how a market changes once a meter that used to


The first meter rarely becomes the final market

Anyone who had a cellphone in the 1990s or early 2000s remembers rationing minutes. Plans counted incoming and outgoing calls. Texts cost extra. People waited until nights and weekends to call anyone, kept conversations short, and sometimes just didn’t pick up because they knew the meter was running on both ends. The bill was a precise record of network usage — it just didn’t say much about whether the conversation was worth having.

Carriers didn’t jump straight from metered minutes to unlimited plans. They worked up to it: bigger buckets, free off-peak calling, bundled long distance, family plans that pooled minutes across a household. The FCC’s 2003 review of U.S. mobile competition found that competitive pressure was pushing carriers toward progressively larger bundles, to the point that an extra minute started to feel essentially free to the customer. Those bundles weren’t just discounts — they changed how people behaved by removing the penalty for using the phone.

A few things had to happen first. Digital networks got a lot more efficient with spectrum. Coverage expanded, and so did switching capacity. Building the network was still expensive, but the marginal cost of one more ordinary call kept falling. That shifted the competitive calculus: carriers had more to gain from getting people to adopt, use, and stick with the service than from making them anxious about every minute they spent on it.

Going unlimited didn’t make network capacity infinite — it just moved the meter out of the customer’s line of sight. Scarcity never actually disappeared. Carriers still had to manage congestion, prioritize traffic, cap hotspot usage, and deal with the small slice of customers whose usage looked nothing like everyone else’s. What changed was the interface between the network’s real economics and what the customer experienced day to day. Routine use became predictable enough to bundle, while anything unusual or premium stayed metered.

That distinction is worth sitting with, because it maps pretty well onto where AI is headed. The likely destination isn’t a world with no metering at all — it’s one where routine, everyday intelligence gets bundled in, while the expensive stuff, deep reasoning, autonomous execution, premium performance, bursts of priority capacity, stays visible on the bill.


AI is similar — and fundamentally harder to flatten

Tokens are a rough proxy for the work a model actually did. They’re closer to a measure of infrastructure consumption than to something like a seat license, but they’re still a poor stand-in for business value. The same token count can produce a throwaway draft, a genuinely useful answer, or a decision with real consequences attached. Swap in a different model and the same task can end up costing more or less, with wildly different quality and review burden.

Enterprise credit models try to abstract this up a level. Workday Flex Credits, for instance, are sold as an annual pool that gets drawn down when specific agent skills or platform capabilities get used in production. That’s easier for a customer to plan around than raw token math, because the unit is tied to something recognizable. But it’s still just a commercial estimate of usage — it doesn’t tell you whether the work actually moved the needle inside a particular organization.

AI also behaves differently than a phone call ever did. A call is naturally bounded by how much time two people have. An agent isn’t bounded that way — it can search, reason, call other tools, try different approaches, and repeat the whole process without much stopping it. One employee kicking off a task can fan out into hundreds of machine actions before it’s done. So the cost of getting something done swings with how ambiguous the task is, how high the quality bar is, which model you picked, and how many attempts it takes to get a reliable result.

Tokens tell you what got consumed. Credits tell you the commercial unit. Neither one tells you whether the work was actually worth doing.


The meter will move up the stack

As inference keeps getting cheaper, the least differentiated slices of AI are going to be a hard sell as standalone premium features. Drafting, summarizing, searching, routine assistance — these are already showing up inside tools organizations already own. Competition is going to keep pushing more of this into subscriptions, the same way big minute bundles eventually made an ordinary phone call feel free.

Where the meter is likely to stick around is wherever the vendor faces real variable cost, or wherever the customer is getting something genuinely more valuable: deeper reasoning, autonomous completion, high-volume execution, specialized models, low latency, reserved capacity, stronger guarantees. Expect providers to keep experimenting with tokens, credits, task-based pricing, tiers, and outcome-based pricing, because none of these approaches perfectly balances variable compute costs against a customer’s need for budget certainty.

Realistically, the market that emerges will be a hybrid. Routine intelligence gets bundled in or covered by generous allowances. The expensive stuff gets metered. Priority capacity costs extra. Some enterprise tasks will get priced by the completed action, because that’s just easier to understand than counting model calls. True outcome-based pricing will stay rare, since a vendor rarely controls the process design, adoption, judgment, and follow-through that actually determine whether an AI action creates value on the customer’s end.

Low-cost Chinese models will accelerate the price reset

This pressure isn’t only coming from efficiency gains inside the big U.S. labs anymore. A growing set of Chinese-developed models — DeepSeek, Alibaba’s Qwen, Moonshot’s Kimi, among others — are closing the capability gap while landing at much lower prices, often as downloadable open weights. Stanford’s 2026 AI Index found the top U.S. and Chinese models were only 2.7 percent apart in performance as of March 2026, after the lead had already traded back and forth between the two countries since early 2025. The real implication goes beyond any single benchmark score: once credible substitutes exist, it gets harder to hold premium pricing on intelligence that customers can get somewhere else.

But for enterprise buyers, a cheaper model isn’t automatically a lower-risk one. A hosted service raises real questions about where the data lives, which laws apply, how long it’s retained, whether it’s used for training, what security assurances exist, and what contractual recourse you have. DeepSeek’s current privacy policy states that personal data used with its services is collected, processed, and stored directly in the People’s Republic of China, and says the service isn’t designed for sensitive personal data. That might be fine for some use cases and a non-starter for others — the right answer depends on the information involved, the jurisdictions in play, and what’s at stake if something goes wrong, not on which country the vendor happens to be based in.

Open-weight deployment changes this calculus somewhat, but it doesn’t make it disappear. Running a model inside an environment you control can address some of the data-residency and access concerns, but you still need to dig into licensing, model and training provenance, security testing, how updates are handled, content behavior, support, auditability, and whether you actually have the in-house expertise to run it. A model that’s cheaper per token can still produce a more expensive enterprise service once you factor in the extra hosting, controls, evaluation, and human review it demands.

The cheapest capable model isn’t necessarily the cheapest enterprise decision. That’s exactly why model selection belongs inside your operating model, not a procurement spreadsheet. You still have to decide what work is appropriate to hand off, how much authority the model gets, what information it can touch, how you’ll evaluate and run it, and what the full economics look like once it’s actually in production. Chinese and other open-weight models add real competition and optionality to the mix — they’re not an excuse to skip that discipline. If anything, they make it more valuable.


Scarcity and overbuild can be true at the same time

Today’s supply constraints are real. On Microsoft’s fiscal 2026 fourth-quarter earnings call, the company said Azure demand was still outpacing available capacity, and reported $41 billion in capital expenditure for the quarter — roughly two-thirds of it on shorter-lived CPUs and GPUs. If your application depends on a specific model, region, latency requirement, or service level, take today’s supply limits seriously.

But near-term scarcity doesn’t settle the long-run capacity question. Stanford’s 2025 AI Index found that querying a system performing at roughly GPT-3.5’s level cost $20 per million tokens in November 2022 and had fallen to $0.07 by October 2024, more than a 280-fold drop. Hardware price-performance and energy efficiency improved right alongside it. Capacity can grow two ways at once: more infrastructure gets built, and each unit of that infrastructure does more.

Which is where the fiber-optic comparison actually helps. The late-1990s telecom boom laid down enormous amounts of network capacity. When demand didn’t show up as fast as expected, and new transmission technology let more traffic ride the same fiber, a lot of that capacity sat dark and infrastructure values collapsed. But the cheap bandwidth left behind is exactly what made a whole generation of new digital businesses possible. The capital cycle punished a lot of the builders while handing the benefits to customers and whoever built on top of them next.

The Federal Reserve noted in 2026 that the recent surge in AI-related IP and equipment investment looked comparable to the 1990s tech boom, and that a future capital overhang was a real possibility. That doesn’t mean the investment is irrational or the technology isn’t real — transformative technologies can attract too much capital in some layers while still generating enormous value across the broader economy.

AI infrastructure isn’t a perfect match for fiber, though. GPUs and CPUs age fast, which raises the risk tied to any given hardware generation, but it also lets providers throttle back future orders if demand shifts. Data-center sites, power, and cooling last much longer. The result could be messy and uneven: shortages in one valuable model or region sitting right alongside excess capacity somewhere else, both shifting faster than a typical procurement cycle can keep up. Overbuilding can hurt infrastructure returns while making intelligence cheaper and more useful for everyone downstream — at the same time.


Do not make your AI strategy a forecast of token prices

The wrong move here is turning your AI strategy into a bet on when compute becomes abundant. If capacity stays tight, disciplined selection and cost control will matter most. If it becomes abundant, your ability to actually apply intelligence to valuable work will matter even more. Either way, the goal is to come out stronger — not to guess correctly.

That should reshape where you invest. Trusted data, clear ownership of processes, reusable integrations, evaluation methods, governance, the ability to adapt, and real evidence of how work is improving — these are durable. Large prepaid credit pools, bespoke dependence on a single model, dedicated infrastructure, and plans that only pencil out if prices collapse the way you’re hoping — those are not.

And don’t assume cheaper intelligence automatically shrinks your AI bill. Lower prices tend to expand demand: more employees start using it, agents attempt more work, and things that used to be uneconomical suddenly make sense. Providers may also hang onto some of those efficiency gains themselves, or redirect them into higher-quality service instead of passing them straight through. So model cost per useful outcome and cost at scale — don’t just celebrate a lower price per token.


The practical questions leaders should answer

The point isn’t to eliminate uncertainty. It’s to make commitments that generate value and learning without betting the company on being right about the entire market. Five questions are a decent filter.

  • Does this use case pay off at today’s price? If your business case only works after a tenfold cost drop, don’t scale it as if that were already true. A small, contained pilot can still be worth running if the learning transfers, the risk is bounded, and you can walk it back.
  • What happens to your bill if adoption actually succeeds? Pilot-stage averages are a bad guide to production reality. Model out transaction volume, concurrency, agent fan-out, retries, and how much human review you’ll need at scale. Success shouldn’t blow up your costs the moment people start actually using the thing.
  • Which layer of capability are you really paying for? Generic assistance is going to keep getting cheaper and more bundled. Contextual intelligence tied to something that’s actually differentiating for you might justify a deeper investment. Spend based on the value of the work, not how novel the model feels.
  • How much room do you have left after signing the contract? Look closely at credit expiration, minimum commitments, exclusivity clauses, data portability, observability, integration design, and how hard it would be to switch vendors. A flexible-looking rate card can hide a very rigid technical architecture underneath.
  • What’s still valuable if the pricing model changes tomorrow? A good investment leaves you with better data, cleaner workflows, stronger evaluation, real governance, and people who know how to use this stuff. If those assets hold up regardless of vendor or pricing model, falling prices only strengthen the case — they don’t undercut it.

Move on learning. Stay flexible on capacity.

Waiting to see what happens can sound like fiscal discipline, since prices are probably going to keep moving. But it has a real cost of its own: you can’t buy back operating experience once the market gets cheap. It takes time to learn what work actually matters, where AI falls short, what context it needs, how people actually use it day to day, and which controls make it dependable.

Move now if the problem is real, the economics work today, the risk is bounded, the learning carries over, and you can reverse course if you need to. That covers cleaning up your data and workflows, using what’s already sitting inside platforms you own, running real experiments, and putting governance and evaluation in place before scale makes retrofitting them a nightmare.

Move more slowly when the capability is generic and likely to get bundled anyway, when reliability isn’t there yet, when the business case only works if prices fall dramatically, or when the commercial terms lock you into a technology generation that won’t last. Be especially careful with workforce reductions or big fixed capacity commitments — hold off until you have actual production evidence to back them up.

So it’s not really early adoption versus prudent waiting. It’s a portfolio: commit to the customer and operating problems that are actually worth solving, learn by deploying things that work at today’s prices, and stay flexible on the infrastructure, model, and pricing unit you use to deliver them.

Move early on what you’re learning as an organization. Stay flexible on the meter. The leaders who come out ahead in the next phase of AI won’t be the ones who called the cost curve correctly — they’ll be the ones who picked the right work, understood the full economics, and built an organization that can turn cheaper intelligence into better outcomes no matter how the market shifts underneath them.

The operating model matters more than the model

The market might bundle routine intelligence, meter the premium stuff, or swing back and forth between the two every time a new provider or architecture resets the cost curve. None of that changes the core leadership job: pick the work worth changing, decide how much authority the machine gets, ground it in trusted context, run it over time, understand what it actually costs end-to-end, and build the capability to keep making these calls well. Price changes your options. It doesn’t make the decisions for you.

References 

  • Federal Communications Commission. Eighth Report, 2003. 
  • Workday. Enterprise AI consumption and annual credit-pool model. 
  • Microsoft. Capacity demand and capital-expenditure discussion. 
  • Stanford Institute for Human-Centered AI.  Inference cost, hardware price-performance and efficiency trends. 
  • Stanford Institute for Human-Centered AI.  U.S.–China performance convergence and open-versus-closed model performance. 
  • Stanford Institute for Human-Centered AI. The distinction between open weights and fully open-source models, 2026. 
  • DeepSeek. Current API pricing and time-of-day discounts. 
  • DeepSeek. Data collection, storage and sensitive-data terms. 
  • Board of Governors of the Federal Reserve System. FEDS Notes, July 6, 2026.