Local LLM: running a language model on your own hardware

A local LLM is a large language model that runs on a computer or server you control. The model reads your prompts and files on that machine. This page shows what it takes to run one at a manufacturer or distributor, and when a hosted model fits better. You need a model whose licence suits your business, a machine with enough memory to hold it, a program to run it and a person to look after it. Every vendor fact below links to the vendor’s own page, checked on September 27, 2026.

Roller conveyors run through a bright, clean automated line past a control station with a touchscreen and an emergency stop

What a local LLM is

A large language model (LLM) is the kind of AI model behind ChatGPT and Claude. The model itself is a large file of numbers, called weights or parameters, learned from text. Sizes are counted in them: a “27B” model has about 27 billion. An open-weight model is one whose maker publishes those weights, so anyone can download and run it. Running a model to answer a request is called inference, and with a local LLM, inference happens on hardware you control. Local AI is the wider term for any AI model you run this way.

Models read and write text in tokens, which are pieces of words. Anthropic’s pricing page gives a rough rule: “1 token is approximately 4 characters or 0.75 words in English.”

What stays on your machine

The program that runs the model decides what leaves the machine. Ollama’s FAQ says: “Ollama runs locally. We don’t see your prompts or data when you run locally.” Ollama also offers cloud-hosted models, and for those, the same page says Ollama processes your prompts and responses to provide the service, without storing or logging them. The OLLAMA_NO_CLOUD=1 setting puts Ollama in local-only mode, with its cloud features turned off.

LM Studio’s home page promotes Bionic, its agent (a program that uses a model to carry out tasks step by step), and cloud services that it describes as “Zero Data Retention (ZDR) across the board.” Zero data retention means the provider does not keep your prompts or its answers after it replies. Bionic’s model page says to “choose a Cloud model when you want a hosted model without using your computer’s memory for inference,” and those run in “LM Studio Secure Cloud, under Zero Data Retention.” Check which models your team can choose. In Ollama, the OLLAMA_NO_CLOUD=1 setting above turns the cloud options off.

The machine also needs an internet connection to download models. Ollama’s FAQ says it “pulls models from the Internet and may require a proxy server to access the models.”

How to run an LLM locally

To run an LLM locally, you need a model file and a program that loads it and answers requests. The table describes four of them and who each one suits. A GPU (graphics processing unit) is the chip on a graphics card that does most of a model’s arithmetic, and a CPU is the computer’s main processor. An API (application programming interface) is a way for other software to send the model a request and get the answer back.

ProgramWhat it isLicence or termsWho it suits
OllamaA program that downloads and runs models on your computer, with an API that other software can call.MIT licenceAn IT lead or developer setting up one computer or a first shared server.
LM StudioA desktop app for Mac, Windows and Linux to download models and chat with them. It runs models with “MLX and llama.cpp under the hood,” where MLX is Apple’s machine-learning framework for Apple silicon. It offers OpenAI and Anthropic-compatible endpoints for other software.Free to use at work since July 8, 2025. An Enterprise plan adds single sign-on and control over which models teams can use.One person trying models on a laptop or desktop, or a small team testing before it buys a server.
llama.cppAn inference engine written in C and C++, built to run models “on a wide range of hardware.” It supports quantization from 1.5-bit to 8-bit, which shrinks a model so it needs less memory (explained under hardware below).MIT licenceDevelopers who want direct control. LM Studio uses it inside.
vLLMA serving engine with “continuous batching of incoming requests,” which handles many users at once, and an “OpenAI-compatible API server.” It supports NVIDIA, AMD and Intel GPUs, and several kinds of CPU.Apache 2.0A GPU server that many people or programs use at the same time.

Ollama “supports a subset of the OpenAI API,” and LM Studio and vLLM offer OpenAI-compatible endpoints, the addresses other software sends its requests to.

Check the hardware each program supports. Ollama supports NVIDIA GPUs “with compute capability 5.0+ and driver version 550 and newer,” and Apple’s own GPUs (Ollama GPU page). Compute capability is NVIDIA’s label for the features each of its chips supports, and NVIDIA lists it for each card. LM Studio on a Mac needs Apple Silicon and macOS 14 or newer, with “16GB+ RAM recommended.” On Windows, Intel and AMD (x64) processors need AVX2, a set of processor instructions, and ARM machines with a Snapdragon X Elite chip are also supported. LM Studio says “at least 16GB of RAM is recommended.”

Hardware a local LLM needs

Memory comes first: the model has to fit in the machine’s memory. Ollama’s FAQ shows a model “loaded entirely into the GPU,” “entirely in system memory,” or “partially onto both.” A graphics card has its own memory, called VRAM. On machines such as the NVIDIA DGX Spark and the Apple Mac Studio, the processor and the graphics chip share one pool, called unified memory.

Quantization shrinks a model by storing each weight with fewer bits, for example 4 bits instead of 16. A 4-bit model needs about a quarter of the memory, with some loss of quality. Google publishes the memory each Gemma 4 model needs just to load, including “20% overhead of loading additional things”:

Gemma 4 model16-bit (BF16)8-bit4-bit
E2B11.4 GB5.7 GB2.9 GB
E4B17.9 GB8.9 GB4.5 GB
12B26.7 GB13.4 GB6.7 GB
26B A4B57.7 GB28.8 GB14.4 GB
31B69.9 GB34.9 GB17.5 GB

The same page adds two warnings. First, “Larger context windows require significantly more VRAM on top of the base model weights.” The context window is the amount of text the model holds at once, such as a long supplier contract and your question about it. Second, 26B A4B is a mixture-of-experts model, which activates only about 4 billion of its parameters for each token, yet “all 26 billion parameters must be loaded into memory.”

Other makers publish their own sizes. OpenAI’s gpt-oss model card says its quantization makes “gpt-oss-120b run on a single 80GB GPU (like NVIDIA H100 or AMD MI300X) and the gpt-oss-20b model run within 16GB of memory.” Meta says Llama 4 Scout fits on a single H100 GPU with Int4 (4-bit) quantization, and Llama 4 Maverick fits on “a single H100 host,” meaning a whole server.

These machines show the range, from one graphics card to a rented server in Toronto:

MachineMemory for the modelVendor detail
NVIDIA GeForce RTX 509032 GB of VRAMA graphics card for a desktop PC.
NVIDIA RTX PRO 600096 GB of VRAM with error-correcting code (ECC)A workstation graphics card. ECC memory detects and corrects memory errors.
NVIDIA DGX Spark128 GB unified memoryBuilt for “AI development and testing workloads with AI models up to 200 billion parameters.” Memory bandwidth of 273 GB per second (NVIDIA blog).
Apple Mac Studio with M5 Ultra96 GB unified memory, or 256 GB or 512 GB on the version with the 80-core GPU1.2 TB per second of memory bandwidth.
DigitalOcean server with one RTX 4000 Ada20 GBCharged per GPU per hour on demand, in Toronto.
DigitalOcean server with one H10080 GBCharged per GPU per hour on demand, in Toronto.

In February 2026, NVIDIA raised the suggested retail price of the DGX Spark Founders Edition “due to memory supply constraints,” according to its developer forum. DigitalOcean’s availability page lists both GPU servers on demand in TOR1, its Toronto data centre, and DigitalOcean bills in US dollars. Vendor details were checked on September 27, 2026.

What sets the speed

Memory bandwidth is how fast a chip reads its memory, and the speed at which a model writes its answer rises with it. In llama.cpp’s benchmark on Apple chips, running Llama 2 7B at 4 bits, an M1 Pro with 200 GB per second of bandwidth generated 36.41 tokens per second. An M4 Max with 546 GB per second generated 83.06, and an M2 Ultra with 800 GB per second generated 94.27.

NVIDIA measured the DGX Spark with one user sending 2,048 tokens and receiving 128. On that test, gpt-oss-20b generated 82.74 tokens per second and gpt-oss-120b generated 55.37. Qwen3 235B, split across two DGX Sparks, generated 11.73.

Which model fits your hardware

Tell Derik what machine you have or plan to buy, and what the model would read. He will tell you which models fit it and what to test first.

Start a conversation

Open-weight models and their licences

The licence decides what your company may do with a model. Open-weight means the weights are published for download, and several popular licences still add conditions a business has to check. These models show the range, from the model pages checked on September 27, 2026:

ModelMakerLicenceWhat to check
gpt-oss-20b and gpt-oss-120bOpenAIApache 2.0gpt-oss-120b has 117 billion parameters, with 5.1 billion active for each token.
Gemma 4 (E2B, E4B, 12B, 26B A4B, 31B)GoogleApache 2.0“Out-of-the-box support for 35+ languages.”
Qwen3.8-27BQwenApache 2.0A context window of 262,144 tokens natively.
Qwen3.8-Flash-NextQwenQwen Community License 1.0A company that runs a “Model as a Service or AI Work Assistant business” needs a separate licence. Internal use is exempt if nothing is made available to a third party.
Llama 4 Scout and MaverickMetaLlama 4 Community LicenseIf your products or services had more than 700 million monthly active users in the month before the Llama 4 release date, you must request a licence from Meta. If you distribute or make available the model, or a product or service that contains it, you must display “Built with Llama.” Downloads on Hugging Face are gated: you give your full legal name, date of birth and organization first.
Mistral Small 4MistralApache 2.0119 billion parameters, 6.5 billion active. French is among its listed languages.
Mistral Medium 3.5MistralModified MITNo rights under the licence if your company’s “global consolidated monthly revenue” exceeds $20 million.
DeepSeek-V4.1-FlashDeepSeekMIT“552B backbone parameters,” far beyond a single graphics card.
Granite 4.2 (3B, 8B, 30B)IBMApache 2.0Released August 25, 2026. French is among its tested languages.
Phi-4-reasoning-vision-15BMicrosoftMIT15 billion parameters.

The Apache 2.0 licence lets a business use and change a model without paying a royalty. Read any other licence in full, with your lawyer, before you build on a model. An earlier licence table and DigitalOcean server costs are in ChatGPT alternatives for business.

How to pick the best local LLM for your work

The best local LLM for a business is the one that fits its hardware and does best on a test built from its own documents. We check in this order:

  1. Fit the memory. Pick sizes that fit your machine at 4 or 8 bits, with room left for the context your documents need.
  2. Check the licence. Models under the standard Apache 2.0 or MIT licence carry none of the extra conditions in the table above. Mistral Medium 3.5’s modified MIT licence adds a revenue limit.
  3. Check French. If your team works in French, Mistral Small 4 and Granite 4.2 list French by name, and Gemma 4 lists support for more than 35 languages.
  4. Test on your own work. Run 30 to 50 real questions through two or three candidates, and compare the answers with ones your staff know are right.

Quality and speed against hosted models

The best hosted models still lead the best open-weight models, and the size of the lead depends on the measure.

Some open-weight models are very large. DeepSeek-V4.1-Flash has 552 billion backbone parameters, plus 196 billion in a separate memory component called Engram. At 4 bits per weight, the backbone alone comes to about 276 GB, more than any single graphics card in the table above holds.

Quantization costs some quality. NVIDIA says its 4-bit NVFP4 format gives “near-FP8 accuracy (<1% degradation),” where FP8 is an 8-bit format. That is a vendor’s claim about its own format. OpenAI’s gpt-oss models ship quantized in a format called MXFP4, and its model card says: “All evals were performed with the same MXFP4 quantization.” Test the exact file you plan to run, at the bit count you plan to run it.

The speed figures above are for one user at a time, so time your own requests on both options, with documents as long as your team’s.

A self-hosted LLM for your whole team

A self-hosted LLM is a local model on a server your team shares. Moving from one computer to a shared server changes the defaults you rely on.

Someone also has to run the server. It needs security patches, backups, monitoring and access control, and each new model needs testing on your documents before it replaces the old one. Name that person before you buy hardware. If the model drafts anything that leaves the company, keep a person approving each draft, as human in the loop explains.

When your business should run a local LLM

Confidential data. If the model will read drawings, costing, customer lists or employee records, where it processes them matters. Under PIPEDA, the federal privacy law for businesses, the Privacy Commissioner’s guidelines say: “An organization is responsible for personal information in its possession or custody, including information that has been transferred to a third party for processing.” In Quebec, section 17 of the private-sector privacy act requires a privacy impact assessment and a written agreement before personal information is communicated outside Quebec.

A model on hardware your company owns, with its cloud features off, processes those records without sending them to an AI vendor. A rented GPU server runs in the host’s data centre, so the host is a third party too. Who controls the server also matters, and sovereign AI covers that question. This page is not legal advice.

Cost at volume. A hosted model charges for every token. A machine costs the same each month whether it answers ten requests or ten thousand, up to what it can handle. Your monthly volume decides which costs less.

How to work out the break-even

1. Count a month of tokens. Multiply your requests per month by the tokens each one sends and receives. The example here is a request that sends 4,000 tokens, about 3,000 English words, and gets 600 back.

2. Price it on the hosted model. Multiply the input tokens by the model’s input rate per million tokens, and the output tokens by its output rate. Anthropic’s pricing page lists the rates for each Claude model. Anthropic says Claude 4.7 and later models produce about 30% more tokens for the same text, so count tokens with the model you plan to use.

3. Price the machine per month. A rented server running all month is about 730 hours, so multiply its hourly rate by 730. For bought hardware, spread the price over the months you expect to use it, and add power.

4. Add the people. Count the hours to set up the server and keep it patched, and the hours to test each new model.

5. Divide. The break-even is the machine’s monthly cost, people included, divided by the hosted cost per request. Above your break-even, check that one server keeps up with your busiest hour, and compare answer quality on your own test before you compare prices.

Full per-model rates are in Claude API pricing.

When a hosted model is the better choice

A hosted model fits better in these cases:

If you choose a hosted model, get two things in writing. The first is retention. Under a zero data retention (ZDR) agreement, “Anthropic does not store customer prompts or responses at rest after the API response is returned,” according to its data retention page, and it enables ZDR per organization. Some models are left out: Anthropic says Claude Fable 5 and 5.1 and Claude Mythos 5 and 5.1 “require 30-day data retention and are not available under ZDR unless expressly authorized by Anthropic.”

The second is where requests are processed. Anthropic’s data residency page says its default “global” setting means “Inference may run in any available geography for optimal performance and availability,” and the “us” setting means “Inference runs only in US-based infrastructure.” On Amazon Bedrock, the model card for Claude Opus 5.5 shows no in-region processing in the Canada (Central) and Canada West (Calgary) regions. Its US geographic profile “Keeps data within US and Canada regions.” Other providers publish their own terms, and secure AI at work lists the settings to check.

How ThriveAI sets it up

ThriveAI is an AI engineering company in Ottawa that builds private AI systems for manufacturers and distributors in Ontario and Quebec, on their own data. Derik Lawlis leads every project and stays close to the build.

The platform is designed to keep each client’s data on its own server in Canada. You choose the model: an open-weight model on that server, or a hosted model under a written zero data retention agreement. A hosted model may process requests outside Canada, so the contract names the model and the provider plan it runs under. Private AI for business explains how the pieces fit, and AI training covers hands-on sessions for your team on your own documents. The company and how a project runs are on About ThriveAI.

Questions people ask

Does a local LLM send my prompts anywhere?
Not while the model runs on your machine with cloud features off. Ollama says it does not see your prompts or data when you run locally, and the OLLAMA_NO_CLOUD=1 setting turns its cloud features off. Ollama and LM Studio also offer cloud services, which process prompts on the vendor's servers. The program still downloads models over the internet.
Do I need a GPU to run a local LLM?
Not to try one. Ollama’s FAQ describes a model running “100% CPU,” meaning “loaded entirely in system memory,” and llama.cpp is built for “a wide range of hardware.” Speed rises with memory bandwidth, though. In llama.cpp’s benchmark, an Apple M2 Ultra with 800 GB per second generated about 2.6 times as many tokens per second as an M1 Pro with 200 GB per second.
What is the best local LLM for a business?
It is the model that fits your hardware and does best on a test built from your own documents. Start with Apache 2.0 or MIT models such as gpt-oss, Gemma 4, Qwen3.8-27B, Mistral Small 4 or Granite 4.2. Pick sizes that fit your memory, then compare two or three on 30 to 50 real questions whose answers your staff know.
Can a business use a local LLM commercially?
Usually, and the licence sets the terms. The Apache 2.0 licence lets a business use and change a model without paying a royalty. Llama 4 requires a licence from Meta if your products or services had more than 700 million monthly active users in the month before its release, and if you distribute or make available a product or service that contains it, you must display “Built with Llama.” Mistral Medium 3.5 gives no rights to a company whose global consolidated monthly revenue exceeds $20 million. The Qwen Community License requires a separate licence for a Model as a Service or AI Work Assistant business. This is not legal advice.
Is a local LLM cheaper than a hosted model?
It depends on your volume. A hosted model such as Claude Haiku 4.5 charges per million input and output tokens. A machine costs the same each month however much you use it, such as a 20 GB GPU server in Toronto rented by the hour. Price a month of your own requests both ways, and add the time someone spends running the server.
How many people can share one local LLM?
It depends on the hardware and the program. Ollama handles one request per model at a time by default, and the memory it needs grows with each parallel request you allow. vLLM batches requests from many users together on a GPU server. Test with your busiest hour before you buy.
Does a local LLM keep data in Canada?
Yes, if the machine is in Canada and its cloud features are off, because the model processes prompts on that machine. Hosted models differ by provider. Anthropic’s API offers global or US processing. On Amazon Bedrock, Claude Opus 5.5 has no in-region processing in AWS's two Canadian regions. Check each provider's page for the model you plan to use.

Contact

See whether a model on your own server fits the job

Tell Derik which documents a model would read and where they have to stay. He will tell you whether a model on your own hardware fits the job or a hosted one fits better.

Prefer to talk? Book a meeting.

Your message goes to Derik Lawlis, the founder.