Fine-tuning a language model: when it helps a business and what it costs
Fine-tuning is further training of a language model on your own examples, so it learns a format, a style or a narrow task. Anthropic’s glossary defines it as “the process of further training a pretrained language model using additional data.” For most business questions, better prompts or retrieval from your documents come first, and fine-tuning earns its place when a well-defined task keeps failing with them. This guide covers when to fine-tune, what data it needs, what OpenAI, Google, Amazon Bedrock and open-weight models offer and how they charge, checked on September 28, 2026, what RLHF is, and the risks.

What is fine-tuning?
A language model is built in stages. Anthropic’s glossary describes pretraining as “the initial process of training language models on a large unlabeled corpus of text,” in which the model learns to predict the next word. A model at that stage is poor at following instructions, and “Fine-tuning and RLHF are used to refine these pretrained models.” Assistants such as Claude come out of that process. In Anthropic’s words, “Claude is not a bare language model; it has already been fine-tuned to be a helpful assistant.”
When a business fine-tunes, it adds one more round of training on its own examples. The model’s weights, the numbers it learned in training, shift toward those examples. Google’s tuning guide says the process “adjusts the model’s weights to minimize the difference between its predictions and the actual labels.” LLM fine-tuning, or fine-tuning an LLM, means the same thing for a large language model such as GPT, Gemini or Llama.
Three kinds of fine-tuning
OpenAI’s model optimization guide describes the main methods, and Google and Amazon offer versions of each:
- Supervised fine-tuning: “Provide examples of correct responses to prompts to guide the model’s behavior.” OpenAI lists classification and “Generating content in a specific format” among its best uses.
- Preference tuning: you supply a better and a worse answer to the same prompt. OpenAI calls its version direct preference optimization, and Google calls its version preference tuning.
- Reinforcement fine-tuning: a grader scores the model’s answers and training favours the higher scores. On Amazon Bedrock, “you define reward functions that evaluate response quality” (AWS).
What is RLHF?
RLHF stands for reinforcement learning from human feedback. Anthropic defines it as “a technique used to train a pretrained language model to behave in ways that are consistent with human preferences.” People rank two or more answers to the same prompt, and the training “encourages the model to prefer outputs that are similar to the higher-ranked ones.”
OpenAI’s 2022 InstructGPT paper by Long Ouyang and colleagues showed what it can do. The team fine-tuned GPT-3 on answers written by people, then on people’s rankings of model outputs. In their tests, “outputs from the 1.3B parameter InstructGPT model are preferred to outputs from the 175B GPT-3, despite having 100x fewer parameters.” Model makers run RLHF to turn a pretrained model into an assistant, and Anthropic says Claude “has been trained using RLHF.” A business that fine-tunes starts from a model that has been through that step, and the nearest tools it can use are preference tuning and reinforcement fine-tuning.
When to fine-tune, and when to use RAG or better prompts
Start with prompts. OpenAI’s guide says “The prompt engineering process may be all you need in order to get great results for your use case,” and Google’s says foundation models, the general models before any tuning, work well “when the expected output or task can be clearly and concisely defined in a prompt and the prompt consistently produces the expected output.” Google adds that “Supervised fine-tuning is a good option when you have a well-defined task with available labeled data.” Microsoft’s guide to customizing LLMs lists good cases as “steering the model to output content in a specific and customized style, tone, or format,” or tasks where the instructions are too long for the prompt.
Facts that change belong in retrieval. Retrieval-augmented generation (RAG) has the model read the relevant passages from your documents for each question and cite them, and a corrected document changes the next answer. A fine-tuned model has neither property: to update what it learned, you train it again. The table sums up where each method fits in a plant or distribution business.
| The job | Start with | Why |
|---|---|---|
| Answer questions from drawings, procedures or quality records | RAG | The documents change, and each answer should cite its source |
| Look up a live value, such as stock on hand or an order’s status | A tool connection to the ERP, such as an MCP server | The value changes by the minute and lives in one system |
| Draft replies in the company’s tone | Prompts with a few good examples, then preference tuning if they fall short | Microsoft and OpenAI list tone and style among fine-tuning’s uses, and examples in the prompt cost less to try |
| Sort thousands of incoming emails into a fixed set of categories | Prompts first, then fine-tuning if accuracy or cost falls short | A well-defined task with past, labelled examples is the case Google and OpenAI describe |
| Return every supplier quote in one exact structure | Prompts first, then fine-tuning | A strict output format is one of the listed good cases |
Fine-tuning can also lower running costs. Microsoft notes it can cut costs “by using fewer tokens depending on the task” or “by using a smaller model” that can match a larger one on a particular task.
What data fine-tuning needs
A supervised fine-tuning dataset is a file of examples, each with an input and the output you want. OpenAI’s supervised fine-tuning guide requires a training file of at least 10 examples, one per line, and Microsoft’s example is a team that fine-tuned a model “with hundreds of requests and correct responses.” In a plant, the raw material is often already on file: past emails with the category someone gave them, or past quotes with the summary that went to the customer.
Quality matters more than volume. Microsoft says fine-tuning “requires the use of high-quality training data,” and warns that “Fine-tuning with bad data makes the base model worse, but without a baseline, it’s hard to detect regressions.” Before training, set aside a test set the model never sees, measure the base model on it with your best prompt, and compare the fine-tuned model on the same set. Remove personal information that the task does not need. Anthropic’s glossary adds that fine-tuning “requires careful consideration of the fine-tuning data and the potential impact on the model’s performance and biases.”
Providers bill training by tokens, the small pieces of text a model reads, multiplied by epochs, the number of passes over the data. Google and AWS both state that rule on their pricing pages.
Check whether a task needs fine-tuning
Tell Derik which task your team wants the model to do better. He will suggest a test that shows whether prompts, retrieval or fine-tuning is the right tool.
Start a conversationFine-tuning options and how they charge (September 28, 2026)
Each provider charges in US dollars and lists its rates on the pricing page linked in the table, checked on September 28, 2026. Training is charged per training token, which means tokens in the dataset times epochs, and a tuned model then costs money to store, host or call.
| Provider | What you can fine-tune | How training is charged | Other charges | Where it runs |
|---|---|---|---|---|
| OpenAI | GPT-4.1 and GPT-4.1 mini, only for organizations that have run a fine-tuned model in the past 60 days, until January 6, 2027. The platform is closed to new organizations (OpenAI) | Per million training tokens, by model (OpenAI) | A fine-tuned model is charged per million input and output tokens | On OpenAI’s platform |
| Google Cloud (Gemini Enterprise Agent Platform, formerly Vertex AI) | Gemini 3.5 Flash, Gemini 3.1 Flash-Lite, Gemini 2.5 Pro, Gemini 2.5 Flash and Gemini 2.5 Flash-Lite, plus open models such as Gemma 3 and Llama 3.3 (Google) | Per 1,000 training tokens for Gemini models and per million for open models such as Gemma 3 (Google) | From Gemini 3 onward, a tuned model’s prediction price is higher than the base model’s | Tuning for Gemini 3.5 Flash and 3.1 Flash-Lite in us-central1 and europe-west4, and serving from US and EU endpoints only |
| Amazon Bedrock | Amazon Nova models, Meta Llama 3.1, 3.2 and 3.3, and Anthropic’s Claude 3 Haiku (AWS) | Per 1,000 training tokens for Nova models and per million for Llama models (AWS) | A monthly fee to store each custom model. A custom Llama 3.1 model is served per model unit per hour without a commitment | US East (N. Virginia) for Nova, US West (Oregon) for Llama and Claude 3 Haiku |
| Anthropic | The Claude API does not offer fine-tuning. Anthropic asks interested customers to contact it (Anthropic). Claude 3 Haiku can be fine-tuned in Bedrock | Not listed | Not listed | Claude 3 Haiku fine-tuning runs in Bedrock’s US West (Oregon) Region |
| Open-weight model on your own server | Models published for download, such as OpenAI’s gpt-oss, which OpenAI describes as “Fine-tunable” (Hugging Face) | The cost of GPU time | You run the server that hosts the model | Wherever your server is, including Canada |
OpenAI’s deprecations page sets out the wind-down. Since May 7, 2026, organizations that had not run fine-tuning before cannot create fine-tuning jobs, and on January 6, 2027, active customers lose that ability too. Fine-tuned models keep working until their base model is retired, and fine-tuned GPT-4.1 nano models shut down on October 23, 2026.
Training is rarely the large cost. By our arithmetic, 500 examples of about 1,000 tokens each, trained for three epochs, come to 1.5 million training tokens, and the training bill is that count times the provider’s rate. Hosting can cost far more, because a custom model served by the hour is billed for every hour it runs, whether or not anyone uses it. AWS says custom Nova models trained with parameter-efficient techniques, which train only a small set of added weights, can also be billed on demand, and that “The prices are the same for custom models as base models.” For how these charges add up over a year, see the cost of AI.
Fine-tuning an open-weight model on your own server
An open-weight model is one whose weights are published so anyone can download and run it. OpenAI’s gpt-oss models are released under the Apache 2.0 licence, which OpenAI presents as suited to customization and commercial deployment. On your own server, the training data and the resulting model stay where the server is, which can be in Canada. Local LLM covers the hardware, and private AI for business and sovereign AI cover where data may go.
A method called LoRA, short for low-rank adaptation, cuts the hardware that fine-tuning needs. It comes from a 2021 Microsoft paper by Edward Hu and colleagues. LoRA keeps the original weights frozen and trains a small set of added weights. Compared with full fine-tuning of GPT-3 175B, the authors report that “LoRA can reduce the number of trainable parameters by 10,000 times and the GPU memory requirement by 3 times.”
Risks of fine-tuning
- Bad examples make a worse model. Microsoft’s warning applies: without a baseline measured on held-out questions, a regression can go unnoticed.
- Safety behaviour can weaken. A 2023 study by Xiangyu Qi and colleagues removed GPT-3.5 Turbo’s safety guardrails “by fine-tuning it on only 10 such examples at a cost of less than $0.20,” using harmful examples written for the test. The authors found that fine-tuning on ordinary, harmless datasets can also weaken safety behaviour, to a lesser extent.
- The base model gets retired. A fine-tuned model depends on its base model. OpenAI’s fine-tuned GPT-4.1 nano models stop working on October 23, 2026, and their training has to be redone on another model.
- Knowledge goes stale and cannot be cited. A fine-tuned model cannot point to the document behind an answer, and a price or specification it learned stays fixed until you train it again. Keep facts in RAG.
- Training data leaves your building. With a hosted provider, the training file goes to the provider’s region, which for the Google and Bedrock options above is in the United States or Europe. Quebec’s private-sector privacy act requires a privacy impact assessment before personal information is communicated outside the province, so remove personal information from training data or keep training on a server in Canada. AI governance covers the legal side.
Write 50 to 100 real test questions with correct answers before you fine-tune anything. Score your best prompt on them, then score the fine-tuned model on the same set. If the fine-tuned model does not win clearly, keep the prompt.
Questions people ask
What is fine-tuning in AI?
What is RLHF?
Is fine-tuning better than RAG?
How much data do you need to fine-tune an LLM?
How much does fine-tuning cost?
Can you fine-tune Claude or ChatGPT?
How ThriveAI helps
ThriveAI is an AI engineering company in Ottawa. It builds private AI systems on the client’s own data for manufacturers and distributors in Ontario and Quebec. Derik Lawlis, the founder, leads every project and stays close to the build.
The platform is designed to keep each client’s data on its own server in Canada. You choose the model that reads it: one on that server, or a hosted model under a written zero data retention agreement, under which the provider keeps no copy of a request or its answer. A hosted model may process requests outside Canada, so the contract names the model. A named person at your company approves every action before anything is sent or saved. ThriveAI also runs hands-on AI training, Claude API pricing covers the cost of a hosted model, and About ThriveAI covers the company.