Consumption-based AI billing (credits)
What is consumption-based AI billing?
Consumption-based AI billing is the way most software vendors now charge for the AI features inside products you already pay for. Your CRM, your helpdesk, your Microsoft 365 tenant and your wiki all sell AI on top of the licence, and almost none of them sell it as a flat price per user. They sell it in credits: an abstract unit that every question, every automated action and every resolved conversation draws down, at a rate the vendor decides.
A credit is not a token. When you call a model API directly, you pay per million tokens and the price list is public. A credit sits one layer above that. The vendor measures what your request cost it in tokens, model calls and compute, and converts that into a round number of credits from a table it publishes. Microsoft's Copilot Credits, Salesforce's Flex Credits, HubSpot Credits and Atlassian's Rovo credits all work this way: one currency for the whole product, so you never see the tokens. Even Anthropic does it when you buy through a cloud marketplace: your token usage is priced in dollars, converted to Claude Consumption Units at one cent each, and your AWS or Azure bill shows a single line of units.
Think of it as a prepaid phone card for AI. You buy a pool, the pool empties at different speeds depending on what you do with it, and the surprise is rarely the price per unit. The surprise is how many units a normal month turns out to need.
The billing shapes you meet in software you already license
Vendors mix these freely, and one product often offers three of them at once. Prices below come from the vendors' own pages on 4 September 2026, in US dollars, and they change often.
Per-seat add-on. A fixed amount per user per month on top of the base licence, with the AI usage included up to a fair-use limit. Microsoft 365 Copilot is the best-known example: the licensed user pays a flat price and, inside Microsoft's own apps, draws down no credits.
Credit packs. A prepaid monthly pool. Copilot Studio sells packs of 25,000 Copilot Credits at 200 dollars per pack per month, pooled across the tenant, and unused credits do not carry over. Salesforce sells Flex Credits at 500 dollars per 100,000, with no rollover into the next subscription term.
Pay-as-you-go. No pool, a price per credit billed in arrears. Microsoft's meter for Copilot Chat agents lists one cent per message, HubSpot charges one cent per credit, and Atlassian starts charging one cent per Rovo credit above the included allowance from 3 December 2026.
Outcome-based. You pay when the AI finishes the job. Fin, Intercom's support agent, bills 99 cents per outcome: a resolution, which it defines as the customer confirming the answer or leaving without asking for more help, or a hand-off to a human. HubSpot moved its Customer Agent to 50 credits per resolved conversation in April 2026, about 50 cents, and its Prospecting Agent to 100 credits per lead it recommends.
Hybrid. In practice nearly everything is a hybrid: a seat licence that includes some credits, a pack on top, and pay-as-you-go for the overage. Atlassian gives every Jira or Confluence seat a monthly credit allowance that pools across the organisation, then meters the extra usage.
The reasons vendors bill in credits
The first is that AI has a real marginal cost, unlike a report or a form. Every request costs the vendor tokens, and the amount varies per request: a short question about one document costs a fraction of a long conversation over your whole SharePoint. A flat price per user would overcharge the light users or lose money on the heavy ones. A credit lets the vendor pass that variance on to you without publishing its token bill.
The second is agents. A chat assistant answers once per question. An agent decides for itself how many steps a task needs, calls tools, reads results and loops. The vendor cannot predict that volume and neither can you, so it prices the unit it can count: an action, a message, a run. Bain's pricing partners looked at around 200 B2B software companies in mid-2026 and found that about 55 percent meter by output (messages, actions, documents), about 35 percent by effort (tokens or compute) and only about 10 percent by business outcome. Around 80 percent of the vendors with a meter sell it as a committed capacity rather than raw usage. Expect a pool, measured in something other than tokens.
Where the surprises on the invoice come from
A trigger that fires all day. An autonomous agent set to check a mailbox or a queue every ten minutes runs 144 times a day, whether or not there is anything to do. On Copilot Studio a trigger counts as an agent action at 5 credits, and each tool call it makes is another 5. At 15 credits per run that is about 65,000 credits a month, nearly three packs, for an agent that mostly finds an empty inbox.
Retries. An agent that fails a step and tries again pays again. The vendor counts actions, not successes, unless you are on an outcome-based plan.
Features that cost more than they look. The unit is the same, the rate is not. In Copilot Studio a classic answer is 1 credit, a generative answer 2, an agent action 5, and grounding a question on your tenant's Microsoft Graph adds 10 on top. One switch in the agent's settings can multiply the bill by six.
Expiry. Prepaid pools reset monthly and the remainder is gone. Microsoft, HubSpot and Atlassian all say so on their own pages; Salesforce says the same per subscription term. A pool sized for the busiest month is wasted eleven times a year.
Overage and minimums. When the pool runs out, either the feature stops or the meter starts. Copilot Studio lets agents run until 125 percent of the prepaid capacity and then disables them; Fin's base plan covers 50 resolutions a month whether you use them or not. Neither is wrong, but you should know which one you signed.
A new model, a new rate. The credit stays one credit; what it buys changes. When a vendor moves a feature to a reasoning model it usually adds a premium meter, as Copilot Studio does at 10 credits per thousand tokens on top of the normal feature rate. Nothing on your side changed and the bill doubled.
A worked example: 3,000 support conversations a month
A company gets 3,000 customer conversations a month through its website chat and wants an AI agent to take the first line. Two vendor models, same volume.
Credit pool (Copilot Studio rates). Say a typical conversation is two generative answers and one action, such as looking up an order status: 2 times 2 plus 5 is 9 credits. Three thousand conversations need 27,000 credits. One pack of 25,000 at 200 dollars is not enough; the remaining 2,000 come either from a second pack (another 200 dollars, of which 23,000 credits expire unused) or from a pay-as-you-go meter at one cent each, about 20 dollars. Roughly 220 dollars a month with the meter attached, 400 without it. Every conversation costs the same whether the agent solved it or handed it to a colleague.
Per outcome (Fin rates). Suppose the agent resolves 60 percent of conversations on its own. That is 1,800 resolutions at 99 cents, about 1,780 dollars a month, and nothing for the 1,200 it handed over. Roughly eight times the credit-pool figure.
The two numbers measure different things, so eight-to-one is not the verdict. The credit pool buys you a product you still have to build, tune and connect to your knowledge; the outcome price buys a finished support agent with the resolution rate already measured. What the arithmetic does show is what each model does to your risk. Under the credit pool your bill is flat whether the agent is good or bad. Under the outcome price your bill goes up exactly when the agent works, which is the month you would otherwise have paid a person.
Per-seat licence versus credit pool: who carries the volume risk
Put both models on one axis, who pays when usage turns out different from the forecast, and the choice becomes clearer.
With a per-seat licence the vendor carries the volume risk. You pay the same whether a colleague asks three questions a month or three hundred, and the vendor hopes the average lands under its cost. The price is set high enough to cover the heavy users, so if your people barely use the feature, you are subsidising someone else's. The upside is a bill your finance team can put in the budget in January and forget.
With a credit pool you carry the volume risk. The vendor gets paid per unit and does not care whether a run was useful. A quiet month leaves credits to expire; a busy month, a new trigger or a colleague who discovers the deep-research button pushes you into overage. The upside is that you pay in proportion to use, and that a pilot with five people costs almost nothing.
Outcome-based pricing moves part of the risk back to the vendor, who only earns when its agent finishes the job. Bain's partners put it plainly: under an output model the vendor is paid regardless of whether the output created value, under an outcome model payment depends on the result. The catch is who defines the result. Fin counts an assumed resolution when the customer simply leaves the chat, so a customer who gave up is billed the same as one who was helped. Read the definition before you read the price.
The practical rule for an SME: seats for tools that a known group uses daily, a capped pool for anything an agent triggers on its own, and outcome pricing only when you can audit the outcome yourself.
How to read a quote and keep the bill under control
Ask the vendor for five things before you sign, and put the answers in the contract.
The conversion table. How many credits each feature costs, and which features are on the table at all.
The overage price and the overage behaviour. Does the feature stop or does the meter start, and at what rate.
The expiry. Monthly, per term, or never.
The cap. Can you set a hard stop per agent, per team and per tenant, inside the product.
A sample bill for your volume. Give the vendor your real numbers, 3,000 conversations, 400 documents, a trigger every hour, and ask for the credit count. Microsoft publishes an agent usage estimator for Copilot Studio for exactly this.
Once it runs, four controls cover most of the risk.
Hard caps at the platform. Copilot Studio lets an admin set a monthly limit per agent with a hard stop that switches the agent off at the limit. Microsoft 365 spending policies do the same per user and per group, and HubSpot has an account-wide monthly maximum plus limits per tool. A cap you can raise in a minute is safer than a budget nobody checks.
Alerts at 50 and 80 percent. Most platforms only alert near the end: Atlassian at 80 and 100 percent, HubSpot from 75 percent. Set your own earlier warning, because half the pool gone in a third of the month is the signal.
One owner per pool. One person who sees the consumption report weekly, can turn an agent off and knows who to ask why a number moved. Without that person, the invoice is the first report anyone reads.
A monthly review. Which agents drew the most, whether the expensive features earned their rate, whether a pack or the meter would have been cheaper last month, and whether a new model changed the conversion table. Twenty minutes with the platform's consumption view open, the same discipline you apply to any other cloud bill.