Claude API pricing in US and Canadian dollars

The Claude API is how your own software sends work to Claude directly, without anyone typing into the Claude app. Claude API pricing is set per million tokens, the small pieces of text Claude reads and writes, and billed in US dollars. On Anthropic’s pricing page on September 27, 2026, Claude Opus 5.5 costs $4 per million input tokens and $20 per million output tokens, and Claude Sonnet 5 costs $2 and $10. Claude Haiku 4.5 is the cheapest at $1 and $5, and Claude Fable 5.1 the most expensive at $10 and $50. The Batch API, for work that can wait up to 24 hours, takes 50% off. Prompt caching stores text that goes out with every request, such as your instructions, and cuts the cost of sending it again. On the assumptions below, reading 500 two-page purchase orders a month costs about US$28 on Sonnet 5, or C$39.61 at the Bank of Canada rate.

Circular automated inspection line with orange test frames above a roller conveyor on a bright plant floor

Claude API pricing by model

A token is a small piece of text, often part of a word, that Claude reads or writes. Input tokens are what you send, such as instructions and documents. Output tokens are what Claude writes back, including its reasoning. MTok means one million tokens.

All prices on this page are in US dollars. Unless a line names another source, they come from Anthropic’s pricing page as it stood on September 27, 2026. The models overview lists these four as Anthropic’s current models:

ModelInput / MTokOutput / MTokCache read / MTokBatch input / outputAnthropic describes it as
Claude Fable 5.1$10$50$0.25$5 / $25For demanding reasoning and long-horizon agentic work
Claude Opus 5.5$4$20$0.20$2 / $10For long-running agentic coding and knowledge work
Claude Sonnet 5$2$10$0.20$1 / $5The best combination of speed and intelligence
Claude Haiku 4.5$1$5$0.10$0.50 / $2.50The fastest model

Agentic work means Claude carries out a long task in many steps on its own, such as using tools and editing files. A cache read is text Claude reads from the prompt cache, explained below with the cache write prices. Other models on the pricing page:

Claude Opus pricing

Anthropic launched Claude Opus 5.5 on September 22, 2026. Its launch page puts input and output at $4 and $20 per million tokens, 20% below Opus 5, and cache reads at $0.20, 60% below Opus 5. Anthropic also says its own tests show Opus 5.5 costing 40% less than Opus 5 on typical workloads at default settings. That figure comes from Anthropic’s tests, and your documents may give a different result.

Thinking, the reasoning Claude does before it answers, cannot be turned off on Opus 5.5, though Claude decides on each request whether to think and how much. Its default effort setting, which controls how much work goes into each answer, is medium. Fast mode, in research preview, returns output faster at a higher price: $8 per million input tokens and $40 per million output tokens on Opus 5.5. It works on Anthropic’s own API only and cannot be combined with the Batch API.

Claude Sonnet pricing

Claude Sonnet 5 costs $2 per million input tokens and $10 per million output tokens. A footnote on the pricing page says this was introductory pricing through August 31, 2026 and is now the standard price. The increase to $3 and $15 that had been planned for September 1 will not happen. Sonnet 4.6 and Sonnet 4.5 still cost $3 and $15.

Claude Haiku and Fable pricing

Claude Haiku 4.5 is the lowest-priced current model, at $1 and $5. Its context window, the amount of text it can take in one request, is 200,000 tokens, where the other three current models take 1 million. The model deprecations page lists Haiku 4.5 as active, with retirement not sooner than October 15, 2026. Check that page before you build a long-lived process on it.

Claude Fable 5.1 costs $10 and $50, with cache reads at $0.25. Anthropic’s data retention page says Fable 5.1, Fable 5, Mythos 5.1 and Mythos 5 require 30-day data retention. They are not available under zero data retention (ZDR) unless Anthropic expressly authorizes it. ZDR is an arrangement your organization requests from Anthropic’s sales team. Under it, Anthropic does not store your prompts or Claude’s answers after it sends the answer back, except where the law requires it or its automated safety systems flag a request.

How to estimate your tokens

Your monthly cost depends on how many tokens your documents use. Anthropic’s documentation gives these rules:

For a firmer number, send a few real documents to Anthropic’s token counting endpoint, a web address your software calls, which returns the input token count of a request before you run it. It is free to use, within rate limits, and Anthropic calls its result an estimate. Count with the model you plan to use, because the tokenizers differ.

How prompt caching lowers repeat costs

Prompt caching stores the start of a request, such as your instructions or a supplier price list, so later requests that begin with the same text read it from the cache. Writing to the cache costs more than normal input, and reading from it costs much less:

Anthropic says the 5-minute cache pays for itself after one read and the 1-hour cache after two. The caching multipliers stack with the batch discount and with the US-only multiplier.

Model5-minute write / MTok1-hour write / MTokCache read / MTokShortest prompt that caches
Claude Fable 5.1$12.50$20$0.25512 tokens
Claude Opus 5.5$5$8$0.20512 tokens
Claude Sonnet 5$2.50$4$0.201,024 tokens
Claude Haiku 4.5$1.25$2$0.104,096 tokens

A prompt shorter than the minimum is processed without caching, and the API returns no error. Check the cache fields in each response to confirm the cache is being used.

How the Batch API halves the price

The Batch API takes requests you do not need answered right away and charges 50% of the standard price for both input and output. Anthropic says most batches finish within an hour. Results are ready when every request is done or after 24 hours, whichever comes first, and requests still unprocessed at 24 hours expire without a charge. One batch holds up to 100,000 requests or 256 MB.

Batch data is kept. Results stay available for 29 days, and Anthropic’s retention page lists the Batch API as ineligible for zero data retention, with 29-day retention. If your company has a ZDR agreement, a batch request steps outside it for that data. What each AI provider keeps, and for how long, is covered in AI security.

What US-only processing adds, and where your requests run

Inference is the step where the model reads your request and writes the answer. On Anthropic’s API, the inference geography setting has two values. Global, the default, lets inference run in any available geography at the standard price. US keeps inference in the United States and costs 1.1 times the standard price on every token category, cache reads and writes included. Anthropic’s plans page puts it this way: “US-only inference is available at 1.1x pricing for input and output tokens.”

US-only processing works on Claude 4.6 and later models. A request that sends the setting to Haiku 4.5, Opus 4.5, Sonnet 4.5 or an earlier model returns an error. Workspace geo, which sets where Anthropic stores data at rest, offers the US only.

Neither setting keeps processing in Canada. If processing must stay in Canada, you need a model that runs on a server in Canada, which local LLMs and sovereign AI cover. Private AI for business explains the options in between.

Claude pricing on Amazon Bedrock and Google Cloud

Claude is also sold through Amazon Bedrock and Google Cloud. Anthropic’s pricing page says that on these platforms the cloud provider invoices you and sets its own regional prices. There, the endpoint your software calls also decides where the request is processed. A global endpoint routes it across regions for availability, and a regional endpoint keeps it within one geographic area, such as the US. The table below shows input and output prices per million tokens in US dollars. The global price matches Anthropic’s, and the regional or US price is 10% higher.

ModelBedrock globalBedrock regionalGoogle Cloud globalGoogle Cloud US multi-region
Claude Opus 5.5$4 / $20$4.40 / $22$4 / $20$4.40 / $22
Claude Sonnet 5$2 / $10$2.20 / $11$2 / $10$2.20 / $11
Claude Haiku 4.5$1 / $5$1.10 / $5.50$1 / $5Not offered
Claude Fable 5.1$10 / $50$11 / $55$10 / $50$11 / $55

The Bedrock pricing page loads its Claude prices with JavaScript, so the Bedrock columns come from AWS’s published price list for Canada (Central), dated September 25, 2026. The US East (N. Virginia) price list shows the same prices. The Google Cloud columns come from Google’s generative AI pricing page.

Amazon Bedrock

Google Cloud

Worked example: reading 500 purchase orders a month

Say you receive 500 purchase orders a month as PDFs. Claude reads each one and returns the customer, part numbers, quantities, prices and dates as structured data for the ERP, the way AI document extraction describes. The numbers below rest on our assumptions. We did not measure real orders.

Our assumptions
  • Each order is a 2-page PDF.
  • Each page uses about 8,000 input tokens: the top of Anthropic’s range of 3,000 text tokens, plus up to 4,784 image tokens on Claude 4.7 and later models, rounded up.
  • The instructions add 2,000 tokens, for 18,000 input tokens per order.
  • Claude writes 2,000 output tokens per order, covering the header fields, the line items and its thinking.
  • That makes 9 million input tokens and 1 million output tokens a month, at standard global prices with no caching and no batch.
ModelUS$ per monthC$ per month at 1.4145US$ per order
Claude Opus 5.5$56.00$79.21$0.112
Claude Sonnet 5$28.00$39.61$0.056
Claude Haiku 4.5$14.00$19.80$0.028
Claude Fable 5.1$140.00$198.03$0.280

Haiku 4.5 would likely count fewer tokens than shown, because it uses the older tokenizer and caps an image at 1,568 tokens.

How the options change the bill

Before you budget on these numbers, run 10 of your own orders through the free token counting endpoint to get your input count. Then run the same 10 through the model and read the output count in each response. Replace our assumptions with those numbers. Plan for a person to check each extracted order before it reaches the ERP, as human in the loop explains. The model cost sits beside the other costs of AI in a manufacturing business.

Estimate the cost for your own documents

Tell Derik which documents you want Claude to read and how many arrive each month. He will tell you which model fits the job and what the token bill is likely to be.

Start a conversation

Converting Claude API prices to Canadian dollars

Anthropic bills in US dollars. To budget in Canadian dollars, multiply the US price by an exchange rate, and note the rate and its date beside the budget. This page uses the Bank of Canada’s daily average rate for September 25, 2026, the last business day before we checked prices: 1.4145 Canadian dollars per US dollar. The rate moved from 1.4021 on September 21 to 1.4145 on September 25, so check it again when you budget.

Price in C$ = price in US$ × 1.4145

ModelInput, US$ / MTokInput, C$ / MTokOutput, US$ / MTokOutput, C$ / MTok
Claude Fable 5.1$10$14.15$50$70.73
Claude Opus 5.5$4$5.66$20$28.29
Claude Sonnet 5$2$2.83$10$14.15
Claude Haiku 4.5$1$1.41$5$7.07

The rate you actually pay is set by your card issuer or bank. Ask them how they convert US-dollar charges.

Claude API pricing compared with Claude plans

A search for Claude pricing can also mean the subscription plans for people who use Claude in the app. Anthropic’s plans page listed these prices on September 27, 2026:

The API fits software that sends work to Claude on its own, as in the purchase order example. How Claude compares with ChatGPT and Microsoft Copilot for a Canadian business is in ChatGPT alternatives, and enterprise AI platforms covers the tools a company standardizes on.

How to keep the monthly bill predictable

Anthropic API pricing FAQ

Is there a free tier for the Claude API?
New users receive a small amount of free credit to test the API, according to Anthropic’s pricing page. Anthropic’s sales team handles longer trials for enterprise evaluation. After the credit runs out, you pay for the tokens you use.
Is the Claude API billed in US dollars?
Yes. Anthropic says all payments are in US dollars and billing follows actual monthly usage. Standard accounts pay by credit card, and enterprise customers can arrange invoicing. At the Bank of Canada’s daily average rate for September 25, 2026, one US dollar was 1.4145 Canadian dollars.
Does a long prompt cost more per token?
No. Claude 4.6 and later models include the full 1-million-token context window at the standard price, and Anthropic says a 900,000-token request is billed at the same per-token rate as a 9,000-token request. Claude Haiku 4.5 has a 200,000-token context window.
Are thinking tokens billed?
Yes. Anthropic’s thinking guide says the tokens Claude spends reasoning are billed as output tokens, even when the reasoning text is not returned to you. Thinking cannot be turned off on Claude Opus 5.5 or Claude Fable 5.1, though Claude decides on each request whether to think and how much, and a lower effort setting reduces the cost.
Can batch and prompt caching be combined?
Yes. Anthropic says the caching multipliers stack with the batch discount and with the US-only multiplier. Batch data is kept for up to 29 days after the batch is created, unless you delete it sooner, and falls outside zero data retention.
Does Claude process requests in Canada?
No setting on Anthropic’s API, Amazon Bedrock or Google Cloud keeps Claude’s processing in Canada. Anthropic’s API offers US or global processing only. On Amazon Bedrock, Canada (Central) offers global or US routing and Canada West (Calgary) offers global only, and AWS says the processing happens in the region the request is routed to. Google Cloud lists no Canadian location for Claude. A model that runs on a server in Canada is covered in local LLMs.
Does web search cost extra?
Yes. According to Anthropic’s pricing page, web search on the Claude API costs $10 per 1,000 searches, plus the tokens its results add. Web fetch, which reads a page at a web address, has no charge beyond its tokens.

How ThriveAI helps

ThriveAI is an AI engineering company in Ottawa that builds private AI systems for manufacturers and distributors in Ontario and Quebec, on their own data. Derik Lawlis, the founder, leads every project and stays close to the build.

The platform ThriveAI builds on is designed to keep each client’s data on its own server in Canada. You choose the model that reads it: one that runs on that server, or a hosted model such as Claude, which runs on the vendor’s servers and is used under a written zero data retention agreement. A hosted model may process requests outside Canada. The agreement names the exact model and service, because retention rules differ by model, as the Fable 5.1 rules above show. ThriveAI also runs hands-on AI training on your own documents. For the company and how a project runs, see About ThriveAI.

Contact

Put Claude to work on one job in your business

If you already pay for Claude, tell Derik which documents or which job you want it to take on. He will tell you which model fits and whether the work belongs in the API or in the app your team uses.

Prefer to talk? Book a meeting.

Your message goes to Derik Lawlis, the founder.