KeMeT Tech
← All field notes

DeepSeek V4 Flash Pricing on Azure AI Foundry: Read the Meter, Not Just the Rate Card

October 1, 20264 min read
azureai-agentscost-managementdeepseekcloud-architecture

A two-day DeepSeek V4 Flash deployment on Azure AI Foundry Serverless ran up ₹161,648.62 (about $1,900 USD), and 98.4% of that came from a single line item: "V4 Flash Cached Global Tokens." That's the headline from a Microsoft Q&A thread filed in August 2026, and it's worth reading closely if you're pricing out DeepSeek V4 Flash for anything beyond a demo (learn-1).

We get asked to sanity-check model costs on client Azure bills often enough that this thread is a useful case study, not because DeepSeek V4 Flash is unusually expensive on paper, but because the gap between the rate card and the invoice is exactly the kind of thing that turns a cheap model into a five-figure surprise.

What the published rate card says

According to the thread, the user cross-checked their invoice against Microsoft's own Tech Community post, "Introducing DeepSeek V4 Flash and V4 Pro in Microsoft Foundry," and found these list rates for DeepSeek V4 Flash:

  • Input tokens: $0.19 per 1M tokens
  • Output tokens: $0.51 per 1M tokens
  • Cached input tokens: $0.028 per 1M tokens

Their input and output line items matched that documentation exactly: 145,840.449K input tokens billed at the equivalent of $0.19/M, and 890.597K output tokens billed at $0.51/M. No dispute there. These are list prices as cited in that Q&A thread; Microsoft doesn't break them out by region in the thread itself, so confirm the region-specific number on your own deployment's pricing page before you build a budget around it.

Where the meter diverges: cached global tokens

The problem is the third line. 168,432.896K tokens landed under "V4 Flash Cached Global Tokens," and Azure billed them at an effective $10.00 per 1M tokens, roughly 357 times the $0.028/1M documented rate for cached input. Azure Support reportedly confirmed the $10.00/1M figure was the rate actually applied, then redirected the user to Microsoft Q&A for model-specific clarification rather than resolving it directly.

A second commenter on the same thread, posting about two weeks later, reported something worse: all of their DeepSeek V4 Flash tokens, input, cache, and output, were charged at the flat $10/1M rate. Neither case has a public resolution in the thread as of this writing.

The open questions in the thread are the ones we'd ask too:

  1. Is "V4 Flash Cached Global Tokens" the same billing meter as the documented "Cached Input" line, just mislabeled or misconfigured, or is it a genuinely separate meter that was never disclosed?
  2. Was the $10.00/1M rate ever published anywhere before these invoices landed?
  3. Does the discrepancy only hit specific deployment configurations (region, SKU, serverless vs. provisioned), or is it systemic?

Nobody outside Microsoft can answer that from the thread alone, and we're not going to guess. What we can say is that a published per-token rate and a billed per-token rate disagreeing by two and a half orders of magnitude is not a rounding error. It's either a meter bug or a documentation gap, and either way it's your P&L that absorbs it until it's fixed.

How to catch this before the invoice does

If you're running DeepSeek V4 Flash (or any newly-GA'd Foundry model) in a prompt-caching workload, don't wait for month-end billing to tell you something's wrong. Set a budget action that fires on anomalous spend well before the invoice closes. A basic Azure budget with an action group gives you same-day notice instead of a six-week-old line item to dispute:

resource costBudget 'Microsoft.Consumption/budgets@2023-05-01' = {
  name: 'deepseek-v4-flash-guard'
  properties: {
    category: 'Cost'
    amount: 50
    timeGrain: 'Monthly'
    timePeriod: {
      startDate: '2026-10-01'
    }
    notifications: {
      actualSpend80: {
        enabled: true
        operator: 'GreaterThan'
        threshold: 80
        contactEmails: [
          '[email protected]'
        ]
      }
    }
  }
}

That's a blunt instrument, it caps on total spend, not on a specific meter. For this class of problem you also want a daily cost-by-meter check so a mislabeled line item surfaces in hours, not weeks. Pull Cost Management data and diff it against your own token accounting: if your app logs tokens sent to the cache versus tokens billed under a cache meter, and the two diverge by more than your expected cache-hit discount, that's your trigger to open a support ticket immediately, with screenshots, the way the original poster did.

Our read on deploying V4 Flash right now

This is judgement, not fact: we'd still consider DeepSeek V4 Flash for cost-sensitive, high-throughput workloads on Foundry, the documented input and output rates in this thread are genuinely cheap. But we would not turn on aggressive prompt caching against it in a production account without a day-one cost reconciliation job watching the cache meter specifically. A model billed per-token with a caching discount is a good deal exactly as long as the caching meter bills what it says it bills. Until this discrepancy has a public resolution, treat the cached-token line as unverified and budget your pilot as if caching might not save you anything. If it comes in at $0.028/1M as documented, that's upside. If it doesn't, you've already got a budget alert and a support ticket queued instead of a surprise invoice.

This is also a good moment to separate "which model is cheapest on paper" from "which model's billing we actually trust," because those are different questions and the second one only gets answered by watching real invoices against real token logs for a few weeks before you commit a production workload to it.

Next steps

If you're evaluating DeepSeek V4 Flash, Claude, or any Foundry model for a production agent and want a cost-reconciliation setup that catches meter discrepancies like this one before they hit your invoice, talk to our AI agents team or reach out through /contact.