Claude API pricing in US and Canadian dollars
The Claude API is how your own software sends work to Claude directly, without anyone typing into the Claude app. Claude API pricing is set per million tokens, the small pieces of text Claude reads and writes, and billed in US dollars. On Anthropic’s pricing page on September 27, 2026, Claude Opus 5.5 costs $4 per million input tokens and $20 per million output tokens, and Claude Sonnet 5 costs $2 and $10. Claude Haiku 4.5 is the cheapest at $1 and $5, and Claude Fable 5.1 the most expensive at $10 and $50. The Batch API, for work that can wait up to 24 hours, takes 50% off. Prompt caching stores text that goes out with every request, such as your instructions, and cuts the cost of sending it again. On the assumptions below, reading 500 two-page purchase orders a month costs about US$28 on Sonnet 5, or C$39.61 at the Bank of Canada rate.

Claude API pricing by model
A token is a small piece of text, often part of a word, that Claude reads or writes. Input tokens are what you send, such as instructions and documents. Output tokens are what Claude writes back, including its reasoning. MTok means one million tokens.
All prices on this page are in US dollars. Unless a line names another source, they come from Anthropic’s pricing page as it stood on September 27, 2026. The models overview lists these four as Anthropic’s current models:
| Model | Input / MTok | Output / MTok | Cache read / MTok | Batch input / output | Anthropic describes it as |
|---|---|---|---|---|---|
| Claude Fable 5.1 | $10 | $50 | $0.25 | $5 / $25 | For demanding reasoning and long-horizon agentic work |
| Claude Opus 5.5 | $4 | $20 | $0.20 | $2 / $10 | For long-running agentic coding and knowledge work |
| Claude Sonnet 5 | $2 | $10 | $0.20 | $1 / $5 | The best combination of speed and intelligence |
| Claude Haiku 4.5 | $1 | $5 | $0.10 | $0.50 / $2.50 | The fastest model |
Agentic work means Claude carries out a long task in many steps on its own, such as using tools and editing files. A cache read is text Claude reads from the prompt cache, explained below with the cache write prices. Other models on the pricing page:
- Claude Mythos 5.1 has the same prices as Fable 5.1. It is offered by invitation only, through Anthropic’s Project Glasswing.
- Older models still on the API. Fable 5 costs $10 and $50, with cache reads at $1. Opus 5, 4.8, 4.7, 4.6 and 4.5 cost $5 and $25. Sonnet 4.6 and 4.5 cost $3 and $15.
- Retired on the API. Opus 4.1, Sonnet 4 and Haiku 3.5 remain on Amazon Bedrock and Google Cloud, and Opus 4 on Google Cloud only.
Claude Opus pricing
Anthropic launched Claude Opus 5.5 on September 22, 2026. Its launch page puts input and output at $4 and $20 per million tokens, 20% below Opus 5, and cache reads at $0.20, 60% below Opus 5. Anthropic also says its own tests show Opus 5.5 costing 40% less than Opus 5 on typical workloads at default settings. That figure comes from Anthropic’s tests, and your documents may give a different result.
Thinking, the reasoning Claude does before it answers, cannot be turned off on Opus 5.5, though Claude decides on each request whether to think and how much. Its default effort setting, which controls how much work goes into each answer, is medium. Fast mode, in research preview, returns output faster at a higher price: $8 per million input tokens and $40 per million output tokens on Opus 5.5. It works on Anthropic’s own API only and cannot be combined with the Batch API.
Claude Sonnet pricing
Claude Sonnet 5 costs $2 per million input tokens and $10 per million output tokens. A footnote on the pricing page says this was introductory pricing through August 31, 2026 and is now the standard price. The increase to $3 and $15 that had been planned for September 1 will not happen. Sonnet 4.6 and Sonnet 4.5 still cost $3 and $15.
Claude Haiku and Fable pricing
Claude Haiku 4.5 is the lowest-priced current model, at $1 and $5. Its context window, the amount of text it can take in one request, is 200,000 tokens, where the other three current models take 1 million. The model deprecations page lists Haiku 4.5 as active, with retirement not sooner than October 15, 2026. Check that page before you build a long-lived process on it.
Claude Fable 5.1 costs $10 and $50, with cache reads at $0.25. Anthropic’s data retention page says Fable 5.1, Fable 5, Mythos 5.1 and Mythos 5 require 30-day data retention. They are not available under zero data retention (ZDR) unless Anthropic expressly authorizes it. ZDR is an arrangement your organization requests from Anthropic’s sales team. Under it, Anthropic does not store your prompts or Claude’s answers after it sends the answer back, except where the law requires it or its automated safety systems flag a request.
How to estimate your tokens
Your monthly cost depends on how many tokens your documents use. Anthropic’s documentation gives these rules:
- Newer models count more tokens. Claude 4.7 and later models, which include Opus 5.5, Sonnet 5 and Fable 5.1, use a newer tokenizer, the part that splits text into tokens. It produces about 30% more tokens for the same text. Haiku 4.5 uses the previous tokenizer.
- Words per token. On the newer tokenizer, 1 million tokens is roughly 555,000 English words or 2.5 million characters, which works out to about 1.8 tokens per word. Earlier models fit about 750,000 words in 1 million tokens. The pricing page’s FAQ still gives the older rule of thumb of 4 characters or 0.75 words per token, and says counts vary by language, so count French documents on their own.
- PDFs. Each page typically uses 1,500 to 3,000 text tokens. Claude also reads every page as an image, so image tokens come on top. There is no separate PDF fee. One request can hold up to 600 pages, or 100 pages on a model whose context window is under 1 million tokens.
- Images. An image costs its width divided by 28, rounded up, times its height divided by 28, rounded up. A 1,000 by 1,000 pixel scan is 1,296 tokens. On Claude 4.7 and later models, Claude shrinks any image longer than 2,576 pixels on its long edge or larger than 4,784 tokens, so one image costs at most 4,784 tokens. On other models, the limits are 1,568 pixels and 1,568 tokens.
- Thinking. The tokens Claude spends reasoning are billed as output tokens, even when the reasoning text is not returned to you.
For a firmer number, send a few real documents to Anthropic’s token counting endpoint, a web address your software calls, which returns the input token count of a request before you run it. It is free to use, within rate limits, and Anthropic calls its result an estimate. Count with the model you plan to use, because the tokenizers differ.
How prompt caching lowers repeat costs
Prompt caching stores the start of a request, such as your instructions or a supplier price list, so later requests that begin with the same text read it from the cache. Writing to the cache costs more than normal input, and reading from it costs much less:
- A 5-minute cache write costs 1.25 times the input price.
- A 1-hour cache write costs 2 times the input price.
- A cache read costs 0.1 times the input price, except 0.05 times on Opus 5.5 and 0.025 times on Fable 5.1 and Mythos 5.1.
Anthropic says the 5-minute cache pays for itself after one read and the 1-hour cache after two. The caching multipliers stack with the batch discount and with the US-only multiplier.
| Model | 5-minute write / MTok | 1-hour write / MTok | Cache read / MTok | Shortest prompt that caches |
|---|---|---|---|---|
| Claude Fable 5.1 | $12.50 | $20 | $0.25 | 512 tokens |
| Claude Opus 5.5 | $5 | $8 | $0.20 | 512 tokens |
| Claude Sonnet 5 | $2.50 | $4 | $0.20 | 1,024 tokens |
| Claude Haiku 4.5 | $1.25 | $2 | $0.10 | 4,096 tokens |
A prompt shorter than the minimum is processed without caching, and the API returns no error. Check the cache fields in each response to confirm the cache is being used.
How the Batch API halves the price
The Batch API takes requests you do not need answered right away and charges 50% of the standard price for both input and output. Anthropic says most batches finish within an hour. Results are ready when every request is done or after 24 hours, whichever comes first, and requests still unprocessed at 24 hours expire without a charge. One batch holds up to 100,000 requests or 256 MB.
Batch data is kept. Results stay available for 29 days, and Anthropic’s retention page lists the Batch API as ineligible for zero data retention, with 29-day retention. If your company has a ZDR agreement, a batch request steps outside it for that data. What each AI provider keeps, and for how long, is covered in AI security.
What US-only processing adds, and where your requests run
Inference is the step where the model reads your request and writes the answer. On Anthropic’s API, the inference geography setting has two values. Global, the default, lets inference run in any available geography at the standard price. US keeps inference in the United States and costs 1.1 times the standard price on every token category, cache reads and writes included. Anthropic’s plans page puts it this way: “US-only inference is available at 1.1x pricing for input and output tokens.”
US-only processing works on Claude 4.6 and later models. A request that sends the setting to Haiku 4.5, Opus 4.5, Sonnet 4.5 or an earlier model returns an error. Workspace geo, which sets where Anthropic stores data at rest, offers the US only.
Neither setting keeps processing in Canada. If processing must stay in Canada, you need a model that runs on a server in Canada, which local LLMs and sovereign AI cover. Private AI for business explains the options in between.
Claude pricing on Amazon Bedrock and Google Cloud
Claude is also sold through Amazon Bedrock and Google Cloud. Anthropic’s pricing page says that on these platforms the cloud provider invoices you and sets its own regional prices. There, the endpoint your software calls also decides where the request is processed. A global endpoint routes it across regions for availability, and a regional endpoint keeps it within one geographic area, such as the US. The table below shows input and output prices per million tokens in US dollars. The global price matches Anthropic’s, and the regional or US price is 10% higher.
| Model | Bedrock global | Bedrock regional | Google Cloud global | Google Cloud US multi-region |
|---|---|---|---|---|
| Claude Opus 5.5 | $4 / $20 | $4.40 / $22 | $4 / $20 | $4.40 / $22 |
| Claude Sonnet 5 | $2 / $10 | $2.20 / $11 | $2 / $10 | $2.20 / $11 |
| Claude Haiku 4.5 | $1 / $5 | $1.10 / $5.50 | $1 / $5 | Not offered |
| Claude Fable 5.1 | $10 / $50 | $11 / $55 | $10 / $50 | $11 / $55 |
The Bedrock pricing page loads its Claude prices with JavaScript, so the Bedrock columns come from AWS’s published price list for Canada (Central), dated September 25, 2026. The US East (N. Virginia) price list shows the same prices. The Google Cloud columns come from Google’s generative AI pricing page.
Amazon Bedrock
- Regional premium. Anthropic’s Bedrock page says regional endpoints carry a 10% premium over global ones, and that Fable 5.1 regional endpoints are in us-east-1 only.
- Canadian regions. The same page lists Canada (Central) with Global and US endpoint types, and Canada West (Calgary) with Global only. AWS says that from its Canadian region, data at rest stays in Canada while processing happens in the destination region.
- Batch. Among these four models, AWS’s price list has batch prices for Haiku 4.5 only, at $0.50 and $2.50 on the global endpoint. Anthropic’s Bedrock page lists Message Batches among the features its Bedrock integration does not support.
Google Cloud
- Locations. Google’s pricing page has tabs for the global endpoint, the US and EU multi-regions, us-east5, europe-west1, asia-southeast1 and asia-east1. None is in Canada. Google’s model pages, last updated September 25, 2026, list Opus 5.5 and Sonnet 5 for the US and EU multi-regions and global, and Haiku 4.5 for us-east5, europe-west1 and global.
- Batch. Google lists batch prices of $1 and $5 for Sonnet 5, $0.50 and $2.50 for Haiku 4.5 and $5 and $25 for Fable 5.1 on the global endpoint, and $2.20 and $11 for Opus 5.5 in the US multi-region. On September 27, Google’s global batch row for Opus 5.5 showed the Opus 5 price, which does not fit its other Opus 5.5 rows, so confirm that price with Google before you budget on it.
Worked example: reading 500 purchase orders a month
Say you receive 500 purchase orders a month as PDFs. Claude reads each one and returns the customer, part numbers, quantities, prices and dates as structured data for the ERP, the way AI document extraction describes. The numbers below rest on our assumptions. We did not measure real orders.
- Each order is a 2-page PDF.
- Each page uses about 8,000 input tokens: the top of Anthropic’s range of 3,000 text tokens, plus up to 4,784 image tokens on Claude 4.7 and later models, rounded up.
- The instructions add 2,000 tokens, for 18,000 input tokens per order.
- Claude writes 2,000 output tokens per order, covering the header fields, the line items and its thinking.
- That makes 9 million input tokens and 1 million output tokens a month, at standard global prices with no caching and no batch.
| Model | US$ per month | C$ per month at 1.4145 | US$ per order |
|---|---|---|---|
| Claude Opus 5.5 | $56.00 | $79.21 | $0.112 |
| Claude Sonnet 5 | $28.00 | $39.61 | $0.056 |
| Claude Haiku 4.5 | $14.00 | $19.80 | $0.028 |
| Claude Fable 5.1 | $140.00 | $198.03 | $0.280 |
Haiku 4.5 would likely count fewer tokens than shown, because it uses the older tokenizer and caps an image at 1,568 tokens.
How the options change the bill
- US-only processing (1.1 times). Opus 5.5 rises to $61.60 a month (C$87.13) and Sonnet 5 to $30.80 (C$43.57). Haiku 4.5 does not offer this setting.
- Batch (50% off). Opus 5.5 falls to $28.00 a month (C$39.61) and Sonnet 5 to $14.00 (C$19.80). Results can take up to 24 hours. Batch data is kept for up to 29 days after the batch is created, unless you delete it sooner, and falls outside zero data retention.
- Caching a reference list. Suppose each order is checked against a 50,000-token list of your part numbers and prices. Sent with every order, the list adds 25 million input tokens a month. On Opus 5.5 it costs $100 (C$141.45) without caching and $9.80 (C$13.86) with the 5-minute cache. On Sonnet 5 it costs $50 without caching and $7.30 with it. This assumes 20 sittings a month of 25 orders each, run back to back, for 20 cache writes and 480 cache reads.
Before you budget on these numbers, run 10 of your own orders through the free token counting endpoint to get your input count. Then run the same 10 through the model and read the output count in each response. Replace our assumptions with those numbers. Plan for a person to check each extracted order before it reaches the ERP, as human in the loop explains. The model cost sits beside the other costs of AI in a manufacturing business.
Estimate the cost for your own documents
Tell Derik which documents you want Claude to read and how many arrive each month. He will tell you which model fits the job and what the token bill is likely to be.
Start a conversationConverting Claude API prices to Canadian dollars
Anthropic bills in US dollars. To budget in Canadian dollars, multiply the US price by an exchange rate, and note the rate and its date beside the budget. This page uses the Bank of Canada’s daily average rate for September 25, 2026, the last business day before we checked prices: 1.4145 Canadian dollars per US dollar. The rate moved from 1.4021 on September 21 to 1.4145 on September 25, so check it again when you budget.
Price in C$ = price in US$ × 1.4145
| Model | Input, US$ / MTok | Input, C$ / MTok | Output, US$ / MTok | Output, C$ / MTok |
|---|---|---|---|---|
| Claude Fable 5.1 | $10 | $14.15 | $50 | $70.73 |
| Claude Opus 5.5 | $4 | $5.66 | $20 | $28.29 |
| Claude Sonnet 5 | $2 | $2.83 | $10 | $14.15 |
| Claude Haiku 4.5 | $1 | $1.41 | $5 | $7.07 |
The rate you actually pay is set by your card issuer or bank. Ask them how they convert US-dollar charges.
Claude API pricing compared with Claude plans
A search for Claude pricing can also mean the subscription plans for people who use Claude in the app. Anthropic’s plans page listed these prices on September 27, 2026:
- Free: $0.
- Pro: $20 a month billed monthly, or $17 a month billed annually ($200 up front).
- Max: from $100 a month, with 5 or 20 times the usage of Pro.
- Team: for 2 to 150 people, $25 per standard seat a month billed monthly or $20 billed annually. Premium seats, with 5 times the usage of a standard seat, cost $125 or $100.
- Enterprise: US$20 per seat a month billed annually, plus usage at API rates.
The API fits software that sends work to Claude on its own, as in the purchase order example. How Claude compares with ChatGPT and Microsoft Copilot for a Canadian business is in ChatGPT alternatives, and enterprise AI platforms covers the tools a company standardizes on.
How to keep the monthly bill predictable
- Spend caps. Each usage tier on Anthropic’s API has a monthly spend cap: $500 on Start, $1,000 on Build and $200,000 on Scale, in US dollars. The Custom tier has no cap, and its limits are arranged with Anthropic.
- Your own limit. You can set a lower limit on the Billing page of the Claude Console. Once usage reaches it, requests fail with an error until access resumes, so set it with room above your estimate.
- Effort and answer length. Anthropic’s thinking guide says to lower the effort setting first when you want to cut cost, and to use max_tokens, the most tokens Claude may write in one answer, thinking included, when you need a hard ceiling.
- A model per task. Sonnet 5 costs half as much as Opus 5.5 per input and output token, and Haiku 4.5 a quarter. Test the cheaper model on 10 real documents before you settle on the more capable one.
- Tracking and discounts. Usage is tracked in the Claude Console. Anthropic’s pricing page says volume discounts may be available for high-volume users and are negotiated case by case.
Anthropic API pricing FAQ
Is there a free tier for the Claude API?
Is the Claude API billed in US dollars?
Does a long prompt cost more per token?
Are thinking tokens billed?
Can batch and prompt caching be combined?
Does Claude process requests in Canada?
Does web search cost extra?
How ThriveAI helps
ThriveAI is an AI engineering company in Ottawa that builds private AI systems for manufacturers and distributors in Ontario and Quebec, on their own data. Derik Lawlis, the founder, leads every project and stays close to the build.
The platform ThriveAI builds on is designed to keep each client’s data on its own server in Canada. You choose the model that reads it: one that runs on that server, or a hosted model such as Claude, which runs on the vendor’s servers and is used under a written zero data retention agreement. A hosted model may process requests outside Canada. The agreement names the exact model and service, because retention rules differ by model, as the Fable 5.1 rules above show. ThriveAI also runs hands-on AI training on your own documents. For the company and how a project runs, see About ThriveAI.