Skip to main content
AI & Automation

Azure OpenAI Cost Control: Provisioned Throughput vs Pay-As-You-Go

The Azure OpenAI bill that surprises a finance team is the one that scales with success — every extra user, every longer prompt, every retry adds tokens. The two ways to pay behave very differently once the volume is real, and choosing wrong is expensive in both directions.

Amit Kumar Singh - Technology Consulting Partner at MyData Insights

Technology Consulting Partner · MyData Insights

14+ years in industrial data · Former Accenture & EY · India, GCC, SEA

28 September 2026 · 9 min read

The bottom line

Azure OpenAI can be paid two ways. Pay-as-you-go bills per token — input and output priced separately, output usually dearer — which is ideal for pilots and spiky, low-volume workloads because you pay only for what you use. Provisioned Throughput Units (PTUs) reserve dedicated capacity for a fixed periodic cost, giving predictable spend and latency, and become cheaper than pay-as-you-go once volume is high and steady. The crossover is a volume-and-steadiness question, not a preference. Either way, the levers that cut cost most are the same: shrink the context you send, cache and reuse where you can, route easy work to a smaller cheaper model and reserve the frontier model for what needs it, and cap retrieval and output length. Start on pay-as-you-go, instrument token usage, and move steady high-volume traffic to PTUs when the maths turns.

The Bill That Scales With Success

Most cloud costs scale with data or infrastructure. An Azure OpenAI bill scales with usage of the model — and usage grows exactly when the project is working. Every new user, every longer conversation, every larger document fed into a prompt, every retry on a failed call adds tokens, and tokens are what you pay for. A successful pilot that rolls out to the whole team can multiply its own bill without anyone changing a line of code.

That is not a reason to avoid Azure OpenAI; it is a reason to understand the cost model before the rollout, not after the invoice. The two ways to pay — pay-as-you-go per token, and provisioned throughput for reserved capacity — behave very differently as volume grows, and the right choice at pilot is usually the wrong choice at scale.

Exact prices change and vary by model and region, so this piece deals in the mechanics and the decision rather than quoting rates you should confirm against your own Azure agreement. The mechanics are what let you predict the bill; the rates only fill in the final number.

An Azure OpenAI bill scales with usage of the model — and usage grows exactly when the project is working. A successful pilot that rolls out to the whole team can multiply its own bill without a line of code changing.

Pay-As-You-Go, Explained

Pay-as-you-go (also called standard or consumption) bills per token processed. Input tokens — the prompt, the system message, the retrieved context you send — are priced separately from output tokens, which the model generates, and output is usually the dearer of the two. You pay only for what you actually use, with no reserved capacity and no commitment.

This is the right model for pilots and for spiky, low-volume or unpredictable workloads. If usage is intermittent, paying per call means you pay nothing when nothing is happening, which reserved capacity cannot match. It is also how you learn your real token profile — how many tokens a typical interaction actually costs — which is the information you need to make the PTU decision later.

The weaknesses of pay-as-you-go show up at scale. Cost rises linearly with volume with no volume break, so a high, steady workload pays full rate on every token. And because it runs on shared capacity, latency can vary and very high throughput can hit rate limits. For a heavy production workload, both the cost and the variability become arguments for the other model.

Provisioned Throughput, Explained

Provisioned Throughput Units (PTUs) reserve dedicated model capacity for a fixed cost per period. Instead of paying per token, you buy a guaranteed throughput — a capacity to process so many tokens per minute — and pay the same whether you use all of it or none. It is the reserved-capacity model applied to AI inference.

The two things PTUs buy are predictability of cost and predictability of performance. The bill is a known periodic number regardless of how busy the workload is, which finance teams value far more than a variable line. And because the capacity is dedicated, latency is consistent and you are not competing on shared infrastructure — which matters for a customer-facing or latency-sensitive application. There is usually a minimum commitment, so PTUs are a commitment, not a dial you nudge.

The risk with PTUs is the mirror of pay-as-you-go: you pay for the reserved capacity whether or not you use it, so under-utilised PTUs are pure waste. Buy capacity for a workload that is not yet steady, or size it for a peak you rarely hit, and you have bought an expensive idle. PTUs reward high, steady, predictable utilisation — and punish the opposite.

PTUs buy predictable cost and predictable latency by reserving dedicated capacity for a fixed periodic fee. They reward high, steady utilisation — and punish under-use, because you pay for the capacity whether or not you use it.

The Crossover — When PTUs Win

The choice between the two is arithmetic, not preference. Pay-as-you-go cost rises with every token; PTU cost is flat once bought. Below a certain steady volume, per-token is cheaper because you are not paying for idle capacity. Above it, reserved capacity is cheaper because per-token has no volume break and you are paying full rate on a lot of tokens. Where those lines cross is your decision point.

The crossover is driven by two variables: volume and steadiness. High volume alone is not enough — a workload that is high on average but wildly spiky may still favour pay-as-you-go, because PTUs sized for the peak sit idle in the troughs. It is sustained, predictable volume that tips the maths to PTUs: a production application serving steady traffic through the day, where the reserved capacity is genuinely used most of the time.

This is why you do not make the decision at pilot. At pilot, volume is low and unknown, so pay-as-you-go is right. You instrument the workload, learn the real token profile and the shape of demand, and revisit the choice when volume is high and steady enough that the PTU maths turns. Some estates end up hybrid — PTUs for the steady baseline, pay-as-you-go for the spikes above it — which is often the most economical answer of all.

The Levers That Cut Cost Either Way

Whichever billing model you are on, the cost is driven by tokens and by which model processes them — and those are levers you control. The first is context size. Every token you send in the prompt is paid for (on pay-as-you-go) or consumes reserved capacity (on PTUs), so sending a whole document when a retrieved paragraph would do is expensive. Tighten retrieval so the model gets the relevant chunk, not the whole knowledge base, and trim system prompts that have grown fat over time.

The second is model tiering. Not every task needs the frontier model. Routing easy, high-volume work — classification, extraction, short answers — to a smaller, cheaper model and reserving the larger model for the work that genuinely needs its reasoning can cut cost sharply, because the cheaper model is a fraction of the price per token. Match the model to the job rather than sending everything to the most capable, most expensive option.

The third is caching and output control. Cache and reuse responses where the same question recurs, so you are not paying to regenerate an identical answer. Cap output length — a model told to answer in a sentence generates far fewer (dearer) output tokens than one left to ramble. And control retries and agent loops, because an unbounded retry or a chatty multi-step agent can quietly multiply the token count of a single interaction. These levers apply on day one and matter more as volume grows.

Tokens and model choice are the levers, on either billing model: shrink the context you send, route easy work to a smaller cheaper model, cache repeated answers, cap output length, and bound retries and agent loops.

So What — How to Sequence It

Start on pay-as-you-go. At pilot and early rollout, volume is low and unpredictable, and per-token billing means you pay for exactly what you use while you learn the workload. Trying to size PTUs before you know your token profile is how you buy expensive idle capacity.

Instrument from the start. Track token usage per feature and per user, watch input versus output, and know your cost per interaction — this is the data that tells you both when to optimise and when the PTU maths turns. Apply the levers early: tight retrieval, model tiering, caching, output caps. They cost nothing to adopt and they lower the whole curve, which also pushes the PTU crossover further out.

Move to PTUs when volume is high and steady enough that reserved capacity is cheaper and the predictable bill and latency are worth the commitment — and consider a hybrid, PTUs for the steady baseline and pay-as-you-go for the spikes. The sequence is: pay-as-you-go and instrument, optimise the tokens and the model routing, then reserve capacity for the steady core once the numbers say so. Decide it on the data, not on a default.

Start pay-as-you-go and instrument token usage; apply the token and model levers early; move steady high-volume traffic to PTUs when the maths turns — often a hybrid of PTUs for the baseline and pay-as-you-go for the spikes.

If your Azure OpenAI spend is climbing with adoption and you are unsure whether to keep paying per token or reserve capacity, that is a maths question worth doing before the next rollout. 30 minutes with Amit on your token profile, the pay-as-you-go versus PTU crossover, and the levers that lower the whole curve. No slides. No pitch deck. No obligation to proceed.

Free Assessment

Where does your operation sit on the data maturity curve?

8 questions. 3 minutes. You get a scored breakdown across data infrastructure, analytics readiness, and automation potential — with a specific next step for your industry.

AI & AutomationAzure OpenAICost OptimisationAzureFinOpsLLM

Your Data · Our Technology · Our Automation

Get practical insights every fortnight

Amit writes about Microsoft Fabric, Power BI, AI in operations, and digital transformation for manufacturing and supply chain leaders. Practitioner perspective - no fluff, no vendor spin.

No spam. Unsubscribe any time. Also on Substack.

FAQ

Common questions

What is the difference between pay-as-you-go and PTUs in Azure OpenAI?

Pay-as-you-go bills per token processed — input and output priced separately, output usually dearer — so you pay only for what you use, with no commitment. Provisioned Throughput Units (PTUs) reserve dedicated capacity for a fixed periodic cost, giving predictable spend and consistent latency regardless of how busy the workload is. Pay-as-you-go suits pilots and spiky, low-volume workloads; PTUs suit high, steady production traffic. The choice is arithmetic — below a certain steady volume per-token is cheaper, above it reserved capacity is.

When should we switch from pay-as-you-go to PTUs?

When volume is high and steady enough that reserved capacity is cheaper than per-token, and the predictable cost and latency are worth the minimum commitment. High volume alone is not enough — a spiky workload may still favour pay-as-you-go because PTUs sized for the peak sit idle in the troughs. It is sustained, predictable demand that tips the maths. Do not decide at pilot; instrument the workload first, learn the real token profile and demand shape, then revisit. Many estates end up hybrid — PTUs for the steady baseline, pay-as-you-go for the spikes.

How do we reduce Azure OpenAI costs?

Cost is driven by tokens and by which model processes them, and both are controllable on either billing model. Shrink the context you send — tighten retrieval so the model gets the relevant chunk, not the whole knowledge base, and trim bloated system prompts. Route easy, high-volume tasks to a smaller, cheaper model and reserve the frontier model for what needs it. Cache and reuse responses for recurring questions, cap output length, and bound retries and agent loops so a single interaction cannot quietly multiply its token count. These lower the whole cost curve and push the PTU crossover further out.

Why did our Azure OpenAI bill grow so fast?

Because the cost scales with usage of the model, and usage grows exactly when the project succeeds. Every new user, longer conversation, larger document in a prompt and retry adds tokens, so a pilot that rolls out to the whole team can multiply its own bill with no code change. Common accelerants are sending too much context, using the most expensive model for every task, unbounded retries or chatty agent loops, and no output caps. Instrument token usage per feature, apply the context and model-tiering levers, and the growth becomes predictable rather than a shock.

Is this the challenge you're facing?

Book a 30-minute call. We'll look at your specific operation and tell you what's achievable - plainly and without slides.