What an AI Feature Really Costs
"How much will this cost to run?" is usually the first question a founder or CTO asks about an AI feature, and the most common answer is a shrug followed by "it depends on usage". That answer is how teams end up with a surprise invoice in month two, or with a feature that is priced into the product at a loss.
The running cost of an AI feature can be estimated with reasonable accuracy before you write production code. This post shows how: what you actually pay for, three worked examples from cheap to expensive, the costs that do not show up on the model provider's invoice and the levers that reduce the bill without hurting quality.
What you pay for
Language models are billed per token. A token is a piece of a word; in English, one token is roughly four characters or three quarters of a word, so a page of text is around 500 to 700 tokens. You pay separately for:
- Input tokens: everything you send in a request. The system prompt, instructions, tool definitions, retrieved documents, the conversation so far and the user's message.
- Output tokens: everything the model generates. Output is typically five times more expensive than input.
- Cached input: repeated input the provider has already processed recently, billed at a fraction of the normal input price.
Prices vary by model tier by an order of magnitude. As a concrete reference, these are Anthropic's Claude API list prices per million tokens as of September 2026:
- Claude Haiku 4.5 (small, fast): $1 input, $5 output, $0.10 cached input.
- Claude Sonnet 5.5 (mid-tier): $2 input, $10 output, $0.20 cached input.
- Claude Opus 5.5 (large): $4 input, $20 output, $0.20 cached input.
- Claude Fable 5.1 (most capable): $10 input, $50 output, $0.25 cached input.
Writing to the cache costs 1.25 times the input price for a five-minute cache and twice the input price for a one-hour cache. The Batch API, for work that can wait up to a day, halves every price. Other providers use the same structure with different numbers, so the method below applies regardless of which model you choose. Always check the current pricing page before you commit numbers to a business case, because prices change often and usually downwards.
Two things that make estimates wrong
Before the examples, two effects catch almost every first estimate.
Thinking tokens are output tokens. Current models reason before they answer, and that reasoning is billed as output even when you never show it to the user. A reply of 400 visible tokens can easily cost 1,000 output tokens. Most models let you control reasoning depth with an effort setting, and simple tasks like classification rarely need much of it.
The whole context is sent on every request. Models are stateless. In a conversation or an agent loop, every step resends the system prompt, the tool definitions and everything that happened so far. Costs grow with the length of the history, not with the length of the latest message. This is the main reason agent costs surprise people, as the third example shows.
Example 1: classifying support tickets
A single model call per ticket: read the ticket, return a category and a priority.
- Volume: 50,000 tickets per month.
- Input: 2,000 tokens (instructions, category definitions with examples, the ticket).
- Output: 100 tokens of structured JSON.
- Model: Claude Haiku 4.5. Classification is a good fit for a small model.
Per ticket, that is 2,000 input tokens at $1 per million plus 100 output tokens at $5 per million, which comes to $0.0025. At 50,000 tickets, the feature costs about $125 per month.
This is the typical profile of well-scoped single-call features: extraction, tagging, summarisation, routing. The model bill is small enough that engineering and evaluation time dominate the total cost.
Example 2: drafting support replies with your knowledge base
A retrieval-augmented feature: find the relevant help articles and past answers, then draft a reply for a human agent to review.
- Volume: 20,000 drafts per month.
- Input: 3,000 tokens of instructions and tone guidelines, 6,000 tokens of retrieved documents and 1,500 tokens of ticket thread.
- Output: a 400-token reply plus around 600 tokens of reasoning.
- Model: Claude Sonnet 5.5, because reply quality matters.
Without caching, each draft costs 10,500 input tokens at $2 per million and 1,000 output tokens at $10 per million, which is $0.031, or $620 per month. The 3,000-token instruction block is identical in every request, so with prompt caching it is billed at the cached price. That brings each draft down to $0.0256, or $512 per month.
Now compare it with the value. If a draft saves a support agent four minutes and their time costs $35 an hour, each draft is worth about $2.33. The model costs around one percent of the value it creates. Choosing the larger Claude Opus 5.5 for the same job would roughly double the bill to about $1,000 a month and would still be an easy decision if it measurably improved the drafts.
This is why the right question is rarely "which model is cheapest" but "which model is good enough, and what is a good outcome worth".
Example 3: an agent that investigates order problems
An agent loop: given a customer complaint, look up the order, check payments and shipping, compare with policies and propose a resolution.
- Volume: 5,000 investigations per month.
- An average of eight steps per run.
- 5,000 tokens of system prompt and tool definitions, plus about 1,500 tokens of new context per step (tool results and the model's previous output).
- 800 output tokens per step, most of it reasoning.
- Model: Claude Opus 5.5.
The final context is only 15,500 tokens, but because every step resends the history, one run processes 82,000 input tokens in total. Without caching, a run costs $0.46 and the feature costs $2,280 per month. With caching, each step only pays full price for the new part of the context and reads the rest from cache. The run drops to $0.19 and the feature to $974 per month, a 57% saving from a configuration change.
Two more effects matter for agents:
- Step count drives cost faster than linearly. If runs take sixteen steps instead of eight, total input tokens more than triple. Limiting steps, and designing tools that return what the agent needs in one call, is a cost decision as much as a quality one.
- Failed runs cost money too. If 15% of runs are retried, the monthly bill rises to about $1,120. Budget per completed task, not per request.
The same agent on Claude Sonnet 5.5 with caching would cost about $523 per month. Whether that is the right choice depends on how often each model resolves the case correctly, which is something you measure, not guess.
Measure your own numbers
The examples use assumptions. Your real numbers come from a prototype and a few dozen real inputs. Every API response includes a usage object with exact token counts, so logging the cost of each request takes a few lines:
import type Anthropic from '@anthropic-ai/sdk';
/** USD per million tokens, from the provider's pricing page (Claude API, September 2026). */
const PRICES = {
'claude-haiku-4-5': { input: 1, cacheRead: 0.1, output: 5 },
'claude-sonnet-5-5': { input: 2, cacheRead: 0.2, output: 10 },
'claude-opus-5-5': { input: 4, cacheRead: 0.2, output: 20 },
} as const;
export type PricedModel = keyof typeof PRICES;
/** Cost of one request in USD, from the usage the API returns with every response. */
export function requestCost(model: PricedModel, usage: Anthropic.Usage): number {
const price = PRICES[model];
const cacheWrites = usage.cache_creation?.ephemeral_5m_input_tokens ?? usage.cache_creation_input_tokens ?? 0;
const longCacheWrites = usage.cache_creation?.ephemeral_1h_input_tokens ?? 0;
const usd =
usage.input_tokens * price.input +
(usage.cache_read_input_tokens ?? 0) * price.cacheRead +
cacheWrites * price.input * 1.25 +
longCacheWrites * price.input * 2 +
// Thinking tokens are billed as output, whether or not you display them.
usage.output_tokens * price.output;
return usd / 1_000_000;
}Store the result next to the business entity it belongs to, such as the ticket or the agent run. After a week of real traffic you know the cost per task, its distribution and which inputs are expensive outliers. A simple estimation process looks like this:
- Define the unit of work: one ticket, one document, one agent run.
- Run the prototype on 20 to 50 real examples and record the cost per unit, including retries.
- Multiply by expected volume, and plan for growth and for heavy users.
- Add the overheads described in the next section.
- Compare the total with the value per unit: time saved, tickets deflected, revenue enabled.
Token counts also depend on the model's tokenizer. The same text can produce noticeably more tokens on one model generation than another, so measure on the model you plan to ship rather than converting estimates between models.
The costs that are not on the invoice
The model bill is often the smaller part of the total cost of an AI feature, especially in the first year:
- Engineering time: building the integration, the tools, the guardrails and the user interface around the model. For most features this is the largest cost by far.
- Evaluation: every prompt or model change should be tested against a set of real examples. Running 300 cases in three variants every week for the drafting feature above costs about $90 a month, or half that through the Batch API. Cheap in tokens, but someone has to build and maintain the evaluation set.
- Supporting infrastructure: embeddings and a vector database for retrieval, a queue for background runs and tracing to see what the model did.
- Human review: if people approve the model's work, their time is part of the cost and part of the benefit.
- Extras: server-side tools such as web search are billed per use ($10 per 1,000 searches on the Claude API), data residency options can add a premium, and every tool definition adds input tokens to every request.
If you want a structured way to plan the engineering side, the article on adding AI agents to existing apps covers the architecture, evaluation and rollout that production features need.
Levers that cut the bill
In rough order of impact and effort:
- Prompt caching: keep stable content such as instructions, tool definitions and reference documents at the start of the prompt, and mark it for caching. Cached input costs a fraction of the normal input price. It is usually the largest saving for the least work, as the agent example shows.
- Batch processing: anything nobody is waiting for, such as nightly enrichment, backfills and evaluation runs, can go through a batch API at half price.
- Right-sized models: use a small model for classification and routing, and a large one only where quality measurably improves. Test the larger model at lower effort too; newer models at low effort often match older ones at high effort.
- Smaller context: send summaries and IDs instead of full documents, retrieve fewer and better chunks, and let tools fetch details on demand.
- Limits in code: maximum output length, maximum agent steps, a cost budget per run and a spending limit per customer. Enforced in code, not discovered on the invoice.
Pricing the feature for your customers
Once you know the cost per unit, you can price the feature sensibly. Flat per-seat pricing works when usage per user is predictable and the cost per user is small compared with the seat price. Usage-based pricing or credits fit features where some customers use a hundred times more than others. In both cases, set fair-use limits, because a small number of heavy users or an automated integration can dominate the bill. Track cost per customer from the start, so pricing decisions are based on data rather than on the averages in a spreadsheet.
Conclusion
AI features are not expensive or cheap by nature. A well-scoped classification feature can cost less than a single software licence, while an agent with long histories and no limits can cost thousands a month. The difference is understood before you build: know the unit of work, measure tokens on real inputs, account for reasoning and context growth, use caching and batching and compare the cost with the value it creates.
If you are planning an AI feature and want a realistic estimate of what it will cost to build and run, get in touch . I help teams scope, estimate and build AI features on top of the products they already have.