Fine-tuning a language model: when it helps a business and what it costs

Fine-tuning is further training of a language model on your own examples, so it learns a format, a style or a narrow task. Anthropic’s glossary defines it as “the process of further training a pretrained language model using additional data.” For most business questions, better prompts or retrieval from your documents come first, and fine-tuning earns its place when a well-defined task keeps failing with them. This guide covers when to fine-tune, what data it needs, what OpenAI, Google, Amazon Bedrock and open-weight models offer and how they charge, checked on September 28, 2026, what RLHF is, and the risks.

A fine engraving tool held a hair above a stainless steel rule under a CNC spindle

What is fine-tuning?

A language model is built in stages. Anthropic’s glossary describes pretraining as “the initial process of training language models on a large unlabeled corpus of text,” in which the model learns to predict the next word. A model at that stage is poor at following instructions, and “Fine-tuning and RLHF are used to refine these pretrained models.” Assistants such as Claude come out of that process. In Anthropic’s words, “Claude is not a bare language model; it has already been fine-tuned to be a helpful assistant.”

When a business fine-tunes, it adds one more round of training on its own examples. The model’s weights, the numbers it learned in training, shift toward those examples. Google’s tuning guide says the process “adjusts the model’s weights to minimize the difference between its predictions and the actual labels.” LLM fine-tuning, or fine-tuning an LLM, means the same thing for a large language model such as GPT, Gemini or Llama.

Three kinds of fine-tuning

OpenAI’s model optimization guide describes the main methods, and Google and Amazon offer versions of each:

What is RLHF?

RLHF stands for reinforcement learning from human feedback. Anthropic defines it as “a technique used to train a pretrained language model to behave in ways that are consistent with human preferences.” People rank two or more answers to the same prompt, and the training “encourages the model to prefer outputs that are similar to the higher-ranked ones.”

OpenAI’s 2022 InstructGPT paper by Long Ouyang and colleagues showed what it can do. The team fine-tuned GPT-3 on answers written by people, then on people’s rankings of model outputs. In their tests, “outputs from the 1.3B parameter InstructGPT model are preferred to outputs from the 175B GPT-3, despite having 100x fewer parameters.” Model makers run RLHF to turn a pretrained model into an assistant, and Anthropic says Claude “has been trained using RLHF.” A business that fine-tunes starts from a model that has been through that step, and the nearest tools it can use are preference tuning and reinforcement fine-tuning.

When to fine-tune, and when to use RAG or better prompts

Start with prompts. OpenAI’s guide says “The prompt engineering process may be all you need in order to get great results for your use case,” and Google’s says foundation models, the general models before any tuning, work well “when the expected output or task can be clearly and concisely defined in a prompt and the prompt consistently produces the expected output.” Google adds that “Supervised fine-tuning is a good option when you have a well-defined task with available labeled data.” Microsoft’s guide to customizing LLMs lists good cases as “steering the model to output content in a specific and customized style, tone, or format,” or tasks where the instructions are too long for the prompt.

Facts that change belong in retrieval. Retrieval-augmented generation (RAG) has the model read the relevant passages from your documents for each question and cite them, and a corrected document changes the next answer. A fine-tuned model has neither property: to update what it learned, you train it again. The table sums up where each method fits in a plant or distribution business.

The jobStart withWhy
Answer questions from drawings, procedures or quality recordsRAGThe documents change, and each answer should cite its source
Look up a live value, such as stock on hand or an order’s statusA tool connection to the ERP, such as an MCP serverThe value changes by the minute and lives in one system
Draft replies in the company’s tonePrompts with a few good examples, then preference tuning if they fall shortMicrosoft and OpenAI list tone and style among fine-tuning’s uses, and examples in the prompt cost less to try
Sort thousands of incoming emails into a fixed set of categoriesPrompts first, then fine-tuning if accuracy or cost falls shortA well-defined task with past, labelled examples is the case Google and OpenAI describe
Return every supplier quote in one exact structurePrompts first, then fine-tuningA strict output format is one of the listed good cases

Fine-tuning can also lower running costs. Microsoft notes it can cut costs “by using fewer tokens depending on the task” or “by using a smaller model” that can match a larger one on a particular task.

What data fine-tuning needs

A supervised fine-tuning dataset is a file of examples, each with an input and the output you want. OpenAI’s supervised fine-tuning guide requires a training file of at least 10 examples, one per line, and Microsoft’s example is a team that fine-tuned a model “with hundreds of requests and correct responses.” In a plant, the raw material is often already on file: past emails with the category someone gave them, or past quotes with the summary that went to the customer.

Quality matters more than volume. Microsoft says fine-tuning “requires the use of high-quality training data,” and warns that “Fine-tuning with bad data makes the base model worse, but without a baseline, it’s hard to detect regressions.” Before training, set aside a test set the model never sees, measure the base model on it with your best prompt, and compare the fine-tuned model on the same set. Remove personal information that the task does not need. Anthropic’s glossary adds that fine-tuning “requires careful consideration of the fine-tuning data and the potential impact on the model’s performance and biases.”

Providers bill training by tokens, the small pieces of text a model reads, multiplied by epochs, the number of passes over the data. Google and AWS both state that rule on their pricing pages.

Check whether a task needs fine-tuning

Tell Derik which task your team wants the model to do better. He will suggest a test that shows whether prompts, retrieval or fine-tuning is the right tool.

Start a conversation

Fine-tuning options and how they charge (September 28, 2026)

Each provider charges in US dollars and lists its rates on the pricing page linked in the table, checked on September 28, 2026. Training is charged per training token, which means tokens in the dataset times epochs, and a tuned model then costs money to store, host or call.

ProviderWhat you can fine-tuneHow training is chargedOther chargesWhere it runs
OpenAIGPT-4.1 and GPT-4.1 mini, only for organizations that have run a fine-tuned model in the past 60 days, until January 6, 2027. The platform is closed to new organizations (OpenAI)Per million training tokens, by model (OpenAI)A fine-tuned model is charged per million input and output tokensOn OpenAI’s platform
Google Cloud (Gemini Enterprise Agent Platform, formerly Vertex AI)Gemini 3.5 Flash, Gemini 3.1 Flash-Lite, Gemini 2.5 Pro, Gemini 2.5 Flash and Gemini 2.5 Flash-Lite, plus open models such as Gemma 3 and Llama 3.3 (Google)Per 1,000 training tokens for Gemini models and per million for open models such as Gemma 3 (Google)From Gemini 3 onward, a tuned model’s prediction price is higher than the base model’sTuning for Gemini 3.5 Flash and 3.1 Flash-Lite in us-central1 and europe-west4, and serving from US and EU endpoints only
Amazon BedrockAmazon Nova models, Meta Llama 3.1, 3.2 and 3.3, and Anthropic’s Claude 3 Haiku (AWS)Per 1,000 training tokens for Nova models and per million for Llama models (AWS)A monthly fee to store each custom model. A custom Llama 3.1 model is served per model unit per hour without a commitmentUS East (N. Virginia) for Nova, US West (Oregon) for Llama and Claude 3 Haiku
AnthropicThe Claude API does not offer fine-tuning. Anthropic asks interested customers to contact it (Anthropic). Claude 3 Haiku can be fine-tuned in BedrockNot listedNot listedClaude 3 Haiku fine-tuning runs in Bedrock’s US West (Oregon) Region
Open-weight model on your own serverModels published for download, such as OpenAI’s gpt-oss, which OpenAI describes as “Fine-tunable” (Hugging Face)The cost of GPU timeYou run the server that hosts the modelWherever your server is, including Canada

OpenAI’s deprecations page sets out the wind-down. Since May 7, 2026, organizations that had not run fine-tuning before cannot create fine-tuning jobs, and on January 6, 2027, active customers lose that ability too. Fine-tuned models keep working until their base model is retired, and fine-tuned GPT-4.1 nano models shut down on October 23, 2026.

Training is rarely the large cost. By our arithmetic, 500 examples of about 1,000 tokens each, trained for three epochs, come to 1.5 million training tokens, and the training bill is that count times the provider’s rate. Hosting can cost far more, because a custom model served by the hour is billed for every hour it runs, whether or not anyone uses it. AWS says custom Nova models trained with parameter-efficient techniques, which train only a small set of added weights, can also be billed on demand, and that “The prices are the same for custom models as base models.” For how these charges add up over a year, see the cost of AI.

Fine-tuning an open-weight model on your own server

An open-weight model is one whose weights are published so anyone can download and run it. OpenAI’s gpt-oss models are released under the Apache 2.0 licence, which OpenAI presents as suited to customization and commercial deployment. On your own server, the training data and the resulting model stay where the server is, which can be in Canada. Local LLM covers the hardware, and private AI for business and sovereign AI cover where data may go.

A method called LoRA, short for low-rank adaptation, cuts the hardware that fine-tuning needs. It comes from a 2021 Microsoft paper by Edward Hu and colleagues. LoRA keeps the original weights frozen and trains a small set of added weights. Compared with full fine-tuning of GPT-3 175B, the authors report that “LoRA can reduce the number of trainable parameters by 10,000 times and the GPU memory requirement by 3 times.”

Risks of fine-tuning

Test before you train

Write 50 to 100 real test questions with correct answers before you fine-tune anything. Score your best prompt on them, then score the fine-tuned model on the same set. If the fine-tuned model does not win clearly, keep the prompt.

Questions people ask

What is fine-tuning in AI?
Fine-tuning is further training of a pretrained language model on additional examples, so it learns a format, a style or a narrow task. The model's weights change to match the examples. Assistants such as Claude have already been fine-tuned by their maker, and a business fine-tuning a model adds one more round on its own data.
What is RLHF?
RLHF, or reinforcement learning from human feedback, trains a pretrained model to behave the way people prefer. People rank two or more answers to the same prompt, and training pushes the model toward answers like the higher-ranked ones. OpenAI's 2022 InstructGPT paper used it, and Anthropic says Claude has been trained using RLHF.
Is fine-tuning better than RAG?
They do different jobs. RAG has the model read passages from your documents for each question and cite them, which suits facts that change. Fine-tuning changes the model itself, which suits a consistent format, style or narrow task. Many teams try better prompts first, add RAG for documents, and fine-tune only when a well-defined task still falls short.
How much data do you need to fine-tune an LLM?
OpenAI's supervised fine-tuning required a file of at least 10 examples, and Microsoft describes a team that fine-tuned a model with hundreds of requests and correct responses. Quality matters more than volume: Microsoft warns that fine-tuning with bad data makes the base model worse. Keep a separate test set to prove the fine-tuned model beats your best prompt.
How much does fine-tuning cost?
Training is usually the small part. Providers charge for training per training token, which means the tokens in your dataset times the number of epochs. Hosting can cost more: Bedrock serves a custom Llama 3.1 model per model unit per hour without a commitment, and a tuned Gemini 3 model's prediction price is higher than the base model's. Each provider's pricing page lists the rates.
Can you fine-tune Claude or ChatGPT?
Most businesses cannot fine-tune them directly. Anthropic's glossary says the Claude API does not offer fine-tuning, though Claude 3 Haiku can be fine-tuned in Amazon Bedrock. OpenAI closed its fine-tuning platform to new organizations on May 7, 2026, and active customers can create jobs until January 6, 2027. Gemini models on Google Cloud and open-weight models remain options.

How ThriveAI helps

ThriveAI is an AI engineering company in Ottawa. It builds private AI systems on the client’s own data for manufacturers and distributors in Ontario and Quebec. Derik Lawlis, the founder, leads every project and stays close to the build.

The platform is designed to keep each client’s data on its own server in Canada. You choose the model that reads it: one on that server, or a hosted model under a written zero data retention agreement, under which the provider keeps no copy of a request or its answer. A hosted model may process requests outside Canada, so the contract names the model. A named person at your company approves every action before anything is sent or saved. ThriveAI also runs hands-on AI training, Claude API pricing covers the cost of a hosted model, and About ThriveAI covers the company.

Contact

Check whether your job needs fine-tuning at all

Tell Derik which task you want a model to do better and how your team does it today. He will tell you whether fine-tuning would beat better prompts or retrieval for it.

Prefer to talk? Book a meeting.

Your message goes to Derik Lawlis, the founder.