
AI Pricing Models: How to Price an AI Product
A founder's first pricing page turns product value into a number a customer can understand. Three tiers and a per-seat price can look finished, but an artificial intelligence (AI) product must account for the model usage behind every customer action.
Five pricing models can account for usage costs and protect gross margin. They also support repricing once real usage data comes in.
What Are AI Pricing Models
An AI pricing model defines what a customer pays for and how payment scales with usage, team size or delivered results. It has two layers: the visible pricing architecture, including tiers, contract terms, add-ons and overage rules, and the value metric, such as a seat, token, action or resolved ticket, that billing counts. Software as a service (SaaS) products spent two decades treating the seat as the default metric because serving one more customer cost almost nothing. Buyers still expect a pricing page to look like those SaaS pricing models.
Customers budget around that metric, and your billing system counts it, so switching means rebuilding both at once. Most AI companies get one clean shot at the metric before every contract and every billing integration starts depending on it.
Why Seat-Based Pricing Breaks Down for AI Products
Seat pricing assumed the next customer cost nothing to serve, and every model call breaks that assumption. Inference sits in cost of goods sold (COGS) because it scales with usage, so AI-first products come in below the 70 to 80 percent gross margin range investors have priced into software for two decades. The median SaaS company still runs a 77 percent gross margin, and AI companies routinely land well under it once inference lands in COGS.
Customers vary enormously in how much they use an AI product, and you can't see that range in an average. In early 2023 GitHub Copilot charged $10 a month and reportedly lost more than $20 per user per month on average, with heavy users costing up to $80. Copilot request caps followed in 2025, and on June 1, 2026 GitHub moved to usage-based billing, calling the old model "no longer sustainable" because a quick question and an hours-long agent session cost the same. Seat counts also stop tracking value once the product does the work: a customer who replaces three analysts with one operator and an agent pays for fewer seats while getting more done.
The Five AI Pricing Models and When Each One Works
Seat-based, usage-based, credit-based, hybrid and outcome-based pricing cover nearly every AI pricing page. Each shifts usage risk between you and the customer differently. Fit depends on bounded customer activity, reliable metering and proof that the billed result came from your product.
Seat-Based Pricing
Seat pricing still fits when each customer's activity has a natural ceiling and the customer wants one predictable monthly number. Someone who opens a writing assistant a few times a day fits that description; an agent running unattended for hours does not. Seats only make sense while the inference one customer generates costs a fraction of what that seat bills.
Usage-Based Pricing
Usage pricing charges for tokens, model calls, actions or outputs, and it's the infrastructure default because the model providers bill that way. Claude Sonnet 5 costs $2.00 input and $10.00 output per million, and the cheapest GPT-5.6 tier costs $0.20 per million input tokens, with the flagship tier several times that. Cost and revenue move together, and a heavy customer can't sink you. The buyer pays for that with a bill they can't forecast, and your revenue swings with someone else's traffic.
Credit-Based Pricing
Credits hide the token math behind a unit the buyer understands and, because the customer prepays, fund your working capital. Conversion rates are yours to set: one summary equals two credits, say, and the customer prepays for a pool that finance can approve as a fixed budget. Customers resent credits that expire without warning, and retroactive repricing is worse. Cursor swapped 500 fast requests for $20 of usage at provider rates on June 16, 2025. Chief executive Michael Truell issued an apology and refunds on July 4.
Hybrid Pricing
Hybrid pricing pairs a subscription base with one variable metric, and most mature AI companies end up with some version of it. Among more than 30 established software vendors, roughly 65 percent have layered an AI meter on top of seats, and none had moved to pure usage or outcome pricing. CRV-backed CodeRabbit runs the pattern cleanly, pairing per-developer review tiers with pay-as-you-go usage and a configurable monthly spending cap. Every additional meter on the variable side is another thing sales has to explain.
Outcome-Based Pricing
Outcome pricing charges for the result: Fin starts at $0.99 per outcome, and Zendesk charges $1.50 per committed automated resolution or $2.00 pay-as-you-go. Unsuccessful interactions still consume compute even when they produce no billable result, so your price has to cover the full distribution of attempts behind each resolution. Attribution disputes are routine, which is why Zendesk counts a resolution only after a quiet period with no reopened ticket. Most companies are better off working toward this model than launching on it.
How to Choose the Right AI Pricing Model for Your Product
The choice comes down to a handful of questions about the product and the buyer. Those answers show where usage risk belongs and whether the unit you'd like to sell can be metered. In rough order of importance:
- Value tracking: A heavier customer should reliably get more out of the product. If ten times the volume delivers the same benefit, usage pricing feels like a tax.
- Cost tracking: Your inference bill should move with the unit you charge for. If cost follows actions and you bill seats, you absorb the difference yourself.
- Bill predictability: The customer needs to forecast next quarter's spend. Enterprise buyers who can't will negotiate caps, demand flat tiers or walk.
- Outcome attribution: You have to prove at renewal that the result was yours. If a support rep touched the ticket before it closed, expect the customer to dispute who earned the credit.
- Metering readiness: You can measure and bill the unit today rather than after a quarter of instrumentation work. Anything you can't meter yet is still a roadmap item.
Before product-market fit, the simplest model that generates usage data is usually the right answer to all of them.
How to Set Prices That Protect Your Gross Margin
Most AI pricing advice names unit economics and then skips the arithmetic. Founders tend to do it out of order and start from a competitor's price page. The order that works starts with what one action costs you, then how unevenly customers use the product, then the floor you need, then how often the number should move.
Cost per Action
Fully loaded cost per action covers inference, retrieval and vector storage, orchestration and logging plus any human review. A support agent session on Claude Opus 5 with 10,000 uncached input tokens, 40,000 cached tokens, 18,000 output tokens and an hour of runtime costs about $0.52 all in. Vector reads, tracing and storage add fractions of a cent, and hand review on five percent of sessions adds under a cent more. The loaded cost comes to about 54 cents. Against a $0.99 per outcome price, one successful attempt leaves you about 45 percent gross margin, and two attempts per success puts you underwater.
Usage Spread Across Customers
The median customer is the wrong planning number, because the customers who break your margin are at the far end of the distribution. Usage concentrates in a small group of heavy accounts, and the gap between a brief chat turn and a single agentic task can run orders of magnitude in tokens. You want the median customer priced to your target margin and the heaviest account still covering its own cost.
Margin Floors and Usage Caps
The margin floor comes first and the allowances come after. If you need 60 percent, size each tier's allowance so a customer who uses all of it still leaves you at 60, and make sure the overage rate clears cost on its own. The levers you have are included allowances, overage rates, hard caps and model routing. GitHub's Copilot caps paired a monthly request allowance with an overage rate.
Repricing Cadence
Tying your price one to one to provider token rates is a mistake, even though those rates fall fast enough to make it tempting. Token prices fell about 98 percent while enterprise AI bills tripled, because agentic workflows consume as much as 30 times the tokens per task that a chat completion did. Pricing should stay anchored to customer value. Your margin should absorb the cost movement, and the price itself belongs on a review schedule. About three in four software vendors changed pricing or packaging in the past year.
How to Change Your Pricing Without Losing Customers
New customers should use the new pricing for 60 to 90 days while you measure conversion and bill size, which protects core revenue until real billing data arrives. The next wave is existing customers who opt in for an incentive, and voluntary movers can become reference cases for the rest. After three to six months of data, the remaining customers should migrate cohort by cohort. Customers who are overpaying today are the right cohort to move first.
For the holdouts, three grandfathering policies cover how long the old price should hold:
- Time-limited grandfathering: The existing price holds for a defined window, often 12 months, and then aligns with the new pricing. Customers should get at least 60 days of notice before the change, and a cliff arrives when the window closes.
- Renewal-based grandfathering: Existing terms remain in place until the contract renews, so every account gets a natural transition point and you never break a signed commitment. Customers can face a steep jump when that renewal arrives.
- Permanent grandfathering for early customers: Early customers retain the legacy plan indefinitely as the cohort that carried the company through its first years. Someone has to maintain every legacy plan as a special case in billing, and that only pays off with a small cohort.
Whichever policy you pick, a migration that's purely a price rise moves nobody. Early switchers move when the new plan gives them something they wanted, such as a higher allowance or a bill that maps to their usage. Customers will also switch for access to a new model.
What Investors Read Into Your AI Pricing Model
Investors reviewing pricing in a diligence checklist usually start by asking whether the founder knows their per-customer inference cost. They go on to ask whether the value metric matches how customers describe the value and whether the heaviest customers are profitable.
Founders who can't answer the cost question are usually underpricing without knowing by how much. A customer who buys "fewer support tickets" and pays per token will eventually notice that the pitch and the invoice disagree. We've seen the strongest outcomes from teams that wire cost awareness into their architecture early, rather than after a surprise bill forces the conversation.
One useful example of usage-based pricing in developer infrastructure is CRV-backed Vercel. CRV led Vercel's Series A and backed the company through its B, C, D and E rounds. Its Pro plan charges $20 a month with $20 of usage credit included and meters everything above that in units a developer already counts. New teams also get a default $200 on-demand budget with a hard limit the customer can configure. None of that requires scale to build.
Where AI Pricing Goes Wrong and How to Get Ahead of It
The same pricing failures show up again and again in early stage AI companies. Founders price in a unit the customer doesn't understand, so the buyer can't forecast and finance blocks the deal. Others put a flat plan in front of wide usage variance, and one heavy account ends up setting the margin for everyone. Outcome pricing launches before anyone can prove attribution, and renewal becomes an argument. Underneath each of them is the assumption that the first price is permanent, when it's a hypothesis you'll revise several times before Series B.
Founders who instrument early keep those failures cheap to fix. If your billing system can absorb a new tier without a rewrite, a pricing change takes a week instead of a quarter. Per-customer cost should appear on a dashboard from the first paying customer, and you should meter the unit you'd like to charge for before you bill it. If you're an early stage founder looking for support pricing an AI product, reach out to us to see if we'd be a good fit.
Frequently Asked Questions About AI Pricing Models
How do you price an AI agent?
Most agents launch on a hybrid: a subscription base plus one metered unit, usually actions or credits, with a hard cap the customer controls. Outcome-based pricing is where many agent companies want to end up, but it requires provable attribution and a price high enough to cover the failed attempts you'll absorb.
Should you charge for AI features separately or bundle them into your plan?
Bundling works when the feature's inference cost per customer is small and predictable; a separate charge works when usage varies widely or the feature replaces paid human work. Large vendors have gone both ways, some selling AI as a per-user add-on and others folding it into existing plans while metering only agent features.
What gross margin should an AI startup target?
AI startups usually run gross margins well below the range software investors are used to after accounting for inference costs, and that's an acceptable start when the trend points up. Investors weigh the direction of margin more than the level. They would rather back a company moving from 45 to 55 percent with a model routing and caching plan than one stuck flat at 60 percent.
How often should you change your AI pricing?
Founders should revisit pricing at least once a year, and more often when AI costs are moving quickly. The internal cadence can be quarterly on per-customer cost and margin, while public changes should stay at once or twice a year so customers don't learn to expect surprises.