All posts

Research / VC diligence

What Does an AI Startup's Inference Bill Reveal?

Assess an AI startup's inference costs, gross margin, retries, human review, and customer concentration with a worked example and a practical VC diligence checklist.

By Waypoint ExponentialPublished Revised
A teal cube rests on layers of cream and terracotta blocks beneath a brass slab, representing the costs supporting an AI product

An AI startup's model bill shows part of the cost of serving its customers. To assess the business, connect that bill to completed work, the revenue that work earns, and the people who repair or review it. A cheap model call can still sit inside an expensive service.

Define the unit of completed customer work

Start with the job the customer buys: an accepted invoice extraction, a resolved support case, or an approved research report. Define its completion and quality criteria before calculating unit cost. A model response isn't automatically a completed job. The customer may reject it, staff may need to correct it, or the system may abandon the attempt.

For diligence, calculate delivery cost per accepted job as the total cost of serving the eligible workload divided by the number of jobs that meet the agreed criteria. Include the costs of failed attempts in the numerator. Keep unresolved cases and their age visible; a team can make its apparent completion rate look better by dropping difficult work from the denominator.

Track attempt count alongside accepted-job count. If 10,000 accepted jobs require 14,000 model attempts, dividing cost by 14,000 answers a different question from dividing it by 10,000. Also record the share of jobs that need staff intervention. An improving cost per attempt can conceal a worsening cost per accepted result.

Use the same definition across customers and reporting periods, or explain the differences. If a report takes longer because it includes more documents, segment by workload complexity. For a startup selling several workflows, separate their unit economics before calculating a blended total.

Reconcile the bill with a request ledger

Ask the startup to join provider usage to its application records using request and job identifiers. The ledger should connect each billable attempt to a customer, workflow, model version, and outcome. Record token categories, tool charges, cache behaviour, and the price or discount that actually applied. An estimate from the prompt text alone won't reconcile the invoice.

Anthropic's current pricing documentation, checked on 3 October 2026, separates input, output, cache writes, and cache reads, and describes additional charges for some tools. Its Batch API offers a 50% input and output discount for asynchronous processing. These are reasons to inspect the actual billing categories and service constraints before applying a single token rate.

Reconcile totals over the same period, currency, and billing basis. Check timing differences, credits, discounts, and shared accounts. Report any unallocated amount explicitly. A ledger that covers only the application's main model call may miss a fallback provider or a search service paid through another account.

Separate production delivery from product development and evaluation. Both consume cash, but they answer different questions. Then show actual paid cost alongside a normalised view that removes temporary credits or states the assumptions behind negotiated discounts. Don't silently treat an expiring credit as a permanent margin improvement.

For committed capacity or minimum-spend contracts, show the cash commitment and the allocation method. The startup may allocate a low cost to each busy request while still paying for idle capacity. Explain whether current volume uses the commitment and what another step in volume would require.

Count retries, tools, and human review

Trace a sample job through its full execution. A research assistant might retrieve documents, summarise each, draft an answer, check it, and retry after a reviewer rejects a citation. Attribute the associated model and tool usage to the same job. Check provider usage records to establish which failed or cancelled attempts incurred charges rather than assuming a billing rule.

Group retries by cause. A temporary service error needs a different repair from a missing source record or a prompt that produces invalid output. Put limits around attempts and elapsed time. A retry loop can increase cost without improving the chance of a correct result.

Measure cache hits on the real workload. A repeatable prompt prefix may reduce cost, while changing context, cache expiry, or cache writes alter the saving. Anthropic documents model-specific cache prices and write premiums; calculate the result from observed reads and writes. Don't apply a universal cache discount to all input tokens.

Include direct delivery labour. Record review minutes, correction time, and manual fallback against the workflow. Use a stated labour-cost basis and the startup's consistent accounting policy when calculating gross margin. A founder reviewing output without taking a salary still creates a staffing requirement; show a sensitivity for replacing that work with paid capacity.

Separate planned review from repair. An approval that the customer expects may belong in the product's design. Repeated manual corrections can point to an unresolved quality problem. For either case, report the time per accepted job and the backlog under current staffing. Our guide to turning time savings into financial effects explains why freed hours alone don't establish a cash saving.

Test gross margin as usage grows

Gross margin equals revenue less cost of revenue, divided by revenue. Ask finance to explain which delivery costs it includes and apply the policy consistently across periods. For an operating sensitivity, also show the direct costs that a narrow reported measure excludes. Gross profit still needs to fund product development, sales, and other operating expenses.

The following example uses illustrative monthly figures in US dollars. A fixed-price contract earns $100,000 and initially covers 20,000 accepted jobs. Model inference costs include all attempts. Other delivery costs cover infrastructure, retrieval, and data services; human review and direct support appear separately.

Illustrative monthly economics under a fixed-price contract, US dollars
Measure20,000 jobs40,000 jobs40,000 jobs, half the inference price
Revenue$100,000$100,000$100,000
Model inference$12,000$24,000$12,000
Other delivery costs$8,000$16,000$16,000
Human review$15,000$30,000$30,000
Direct support$5,000$5,000$5,000
Total cost of revenue$40,000$75,000$63,000
Gross margin60%25%37%

This sensitivity assumes unchanged revenue and quality, linear growth in inference, other delivery costs, and review, and unchanged direct support. It isn't a forecast or an industry benchmark. Counting only the initial $12,000 model bill produces an 88% figure, but including the stated delivery costs gives a 60% gross margin.

At twice the usage, cost per accepted job falls from $2 to about $1.88 because support stays flat. Revenue per job falls faster, from $5 to $2.50, so gross margin contracts. Halving the inference price improves that result to 37%; review and other delivery costs still consume revenue.

Replace these assumptions with the startup's evidence. Some costs stay fixed within a capacity band, others grow with each job, and a new customer may require an implementation step. Test both rising usage within existing contracts and new customers bringing new revenue. Keep changes in quality and review policy explicit.

Find the customers who change the economics

Calculate delivery cost and gross profit by customer and workflow. Show the share of revenue, accepted jobs, and total delivery cost for the largest accounts. A blended margin can hide a large customer whose workload costs more than its contract earns.

Look at the distribution within each account. The median job may be cheap while a small group of long-context or repeatedly rejected jobs consumes much of the budget. Inspect expensive jobs directly, then decide whether complexity bands, a product repair, or a different commercial term fits their cause.

Segment customers by launch date and deployment maturity. New implementations may need extra engineering and support that later cohorts avoid. Ask for evidence of that decline rather than accepting a plan to reduce it. Retained customers who use more of the product can also create a cost increase if their revenue stays fixed.

Customer concentration affects the next decision. Losing a high-revenue account may also release its variable delivery cost, but shared staff or committed capacity may stay. Model the gross profit and cash effects separately. For investor review, aggregate sensitive usage records and retain a controlled path to the supporting evidence.

Compare the cost meter with the pricing contract

Read the actual contracts alongside usage. Identify included work, overage terms, discounts, renewal dates, and the customer's right to expand usage. “Unlimited” access deserves a workload sensitivity. A published price page alone may miss the terms that govern the largest accounts.

In their August 2026 pricing commentary, a16z's Tugce Erten and Sarah Wang argue for pricing around work customers recognise rather than raw token consumption. Treat that as a commercial proposal to test: the chargeable unit must make sense to the buyer and have a clear measurement method.

An accepted report may suit outcome pricing if both parties can establish acceptance without a recurring dispute. A seat price with a stated allowance may fit predictable workloads. A consumption charge may fit variable work, but customers need a way to estimate spend and see what caused it.

Check who pays when a job fails, retries, or needs human correction. Make the rule clear in the product and contract. Then compare paid units with the internal cost ledger. Charging for a customer-visible outcome doesn't eliminate the cost of attempts that never reach it.

For a company rebuilding its own operations, use the same ledger against an internal budget rather than customer revenue. Compare total delivery cost and acceptable output with the current process. For a private equity portfolio company, link the saving to a named cost line or capacity constraint and retain implementation costs in the investment case.

Stress-test provider prices and model changes

Calculate current economics at the rates the startup pays, then test the end of credits, a higher effective rate, and a lower rate. Separate model price from tokens per job. Longer context, more output, or a shift towards a more expensive model can outweigh a reduction in the advertised unit price.

A proposed model change needs an evaluation on representative customer work. Measure accepted results, review time, cost, and latency together. A cheaper model that triggers more correction may increase delivery cost. A more expensive model may reduce it if the evidence shows enough avoided retries or review.

Routing simple jobs to a smaller model can improve economics if the team can identify those jobs reliably. Test misrouted difficult cases and the cost of fallback. For batch processing, verify that delayed completion fits the customer's service requirement before assigning a discount to the forecast.

If the startup runs its own inference, include idle capacity, serving infrastructure, operational labour, and the relevant hardware or rental cost. Compare cost per accepted job at observed utilisation. A full-capacity benchmark doesn't establish the cost of a service with uneven demand or a strict response-time commitment.

Ask what happens if the primary provider changes terms or becomes unavailable. A second-provider integration needs tested quality and operational coverage. Include the cost of maintaining it. Our build-or-buy guide covers the broader control and maintenance trade-offs.

Ask for an evidence pack before investing

Request a recent period long enough to include ordinary delivery and meaningful workload variation; 90 days is a practical starting point if the company has that history. A younger startup should state the shorter coverage. Match revenue and cost periods, and keep estimates separate from observed results.

  • Ask for provider invoices and effective rate agreements, with credits, commitments, and unallocated usage identified.
  • Request a customer and workflow ledger joining revenue, accepted jobs, all attempts, tool costs, and delivery labour.
  • Review rejection and fallback rates, review minutes, and cost distributions alongside the acceptance criteria.
  • Inspect the largest accounts' contract terms and contribution to revenue, delivery cost, and gross profit.
  • Ask for sensitivities covering more usage within existing contracts, new customers, expiring discounts, and changed provider rates.
  • Check the evaluation evidence behind planned routing, caching, or model changes, including any effect on customer outcomes.

Use the technical diligence questions to examine quality and operational dependencies behind the numbers. If finance can't reconcile invoices with production work yet, record that measurement gap and the owner of its repair. Present uncertain economics as uncertain.

A company with measured delivery cost, contract coverage for increased use, and a tested path to reduce review gives an investor evidence to assess growth. A company whose economics depend on unpaid founder work or an untested future model price needs a funded plan for those assumptions. Base the investment decision on what the current service costs and what the team can demonstrate about the next stage.

Put the work into practice

AI implementation and delivery

We help SMEs and scale-ups put AI into a specific business workflow. We define the problem, prepare the data, build the software, and help your team operate it in production.