Local LLM: running a language model on your own hardware
A local LLM is a large language model that runs on a computer or server you control. The model reads your prompts and files on that machine. This page shows what it takes to run one at a manufacturer or distributor, and when a hosted model fits better. You need a model whose licence suits your business, a machine with enough memory to hold it, a program to run it and a person to look after it. Every vendor fact below links to the vendor’s own page, checked on September 27, 2026.

What a local LLM is
A large language model (LLM) is the kind of AI model behind ChatGPT and Claude. The model itself is a large file of numbers, called weights or parameters, learned from text. Sizes are counted in them: a “27B” model has about 27 billion. An open-weight model is one whose maker publishes those weights, so anyone can download and run it. Running a model to answer a request is called inference, and with a local LLM, inference happens on hardware you control. Local AI is the wider term for any AI model you run this way.
Models read and write text in tokens, which are pieces of words. Anthropic’s pricing page gives a rough rule: “1 token is approximately 4 characters or 0.75 words in English.”
What stays on your machine
The program that runs the model decides what leaves the machine. Ollama’s FAQ says: “Ollama runs locally. We don’t see your prompts or data when you run locally.” Ollama also offers cloud-hosted models, and for those, the same page says Ollama processes your prompts and responses to provide the service, without storing or logging them. The OLLAMA_NO_CLOUD=1 setting puts Ollama in local-only mode, with its cloud features turned off.
LM Studio’s home page promotes Bionic, its agent (a program that uses a model to carry out tasks step by step), and cloud services that it describes as “Zero Data Retention (ZDR) across the board.” Zero data retention means the provider does not keep your prompts or its answers after it replies. Bionic’s model page says to “choose a Cloud model when you want a hosted model without using your computer’s memory for inference,” and those run in “LM Studio Secure Cloud, under Zero Data Retention.” Check which models your team can choose. In Ollama, the OLLAMA_NO_CLOUD=1 setting above turns the cloud options off.
The machine also needs an internet connection to download models. Ollama’s FAQ says it “pulls models from the Internet and may require a proxy server to access the models.”
How to run an LLM locally
To run an LLM locally, you need a model file and a program that loads it and answers requests. The table describes four of them and who each one suits. A GPU (graphics processing unit) is the chip on a graphics card that does most of a model’s arithmetic, and a CPU is the computer’s main processor. An API (application programming interface) is a way for other software to send the model a request and get the answer back.
| Program | What it is | Licence or terms | Who it suits |
|---|---|---|---|
| Ollama | A program that downloads and runs models on your computer, with an API that other software can call. | MIT licence | An IT lead or developer setting up one computer or a first shared server. |
| LM Studio | A desktop app for Mac, Windows and Linux to download models and chat with them. It runs models with “MLX and llama.cpp under the hood,” where MLX is Apple’s machine-learning framework for Apple silicon. It offers OpenAI and Anthropic-compatible endpoints for other software. | Free to use at work since July 8, 2025. An Enterprise plan adds single sign-on and control over which models teams can use. | One person trying models on a laptop or desktop, or a small team testing before it buys a server. |
| llama.cpp | An inference engine written in C and C++, built to run models “on a wide range of hardware.” It supports quantization from 1.5-bit to 8-bit, which shrinks a model so it needs less memory (explained under hardware below). | MIT licence | Developers who want direct control. LM Studio uses it inside. |
| vLLM | A serving engine with “continuous batching of incoming requests,” which handles many users at once, and an “OpenAI-compatible API server.” It supports NVIDIA, AMD and Intel GPUs, and several kinds of CPU. | Apache 2.0 | A GPU server that many people or programs use at the same time. |
Ollama “supports a subset of the OpenAI API,” and LM Studio and vLLM offer OpenAI-compatible endpoints, the addresses other software sends its requests to.
Check the hardware each program supports. Ollama supports NVIDIA GPUs “with compute capability 5.0+ and driver version 550 and newer,” and Apple’s own GPUs (Ollama GPU page). Compute capability is NVIDIA’s label for the features each of its chips supports, and NVIDIA lists it for each card. LM Studio on a Mac needs Apple Silicon and macOS 14 or newer, with “16GB+ RAM recommended.” On Windows, Intel and AMD (x64) processors need AVX2, a set of processor instructions, and ARM machines with a Snapdragon X Elite chip are also supported. LM Studio says “at least 16GB of RAM is recommended.”
Hardware a local LLM needs
Memory comes first: the model has to fit in the machine’s memory. Ollama’s FAQ shows a model “loaded entirely into the GPU,” “entirely in system memory,” or “partially onto both.” A graphics card has its own memory, called VRAM. On machines such as the NVIDIA DGX Spark and the Apple Mac Studio, the processor and the graphics chip share one pool, called unified memory.
Quantization shrinks a model by storing each weight with fewer bits, for example 4 bits instead of 16. A 4-bit model needs about a quarter of the memory, with some loss of quality. Google publishes the memory each Gemma 4 model needs just to load, including “20% overhead of loading additional things”:
| Gemma 4 model | 16-bit (BF16) | 8-bit | 4-bit |
|---|---|---|---|
| E2B | 11.4 GB | 5.7 GB | 2.9 GB |
| E4B | 17.9 GB | 8.9 GB | 4.5 GB |
| 12B | 26.7 GB | 13.4 GB | 6.7 GB |
| 26B A4B | 57.7 GB | 28.8 GB | 14.4 GB |
| 31B | 69.9 GB | 34.9 GB | 17.5 GB |
The same page adds two warnings. First, “Larger context windows require significantly more VRAM on top of the base model weights.” The context window is the amount of text the model holds at once, such as a long supplier contract and your question about it. Second, 26B A4B is a mixture-of-experts model, which activates only about 4 billion of its parameters for each token, yet “all 26 billion parameters must be loaded into memory.”
Other makers publish their own sizes. OpenAI’s gpt-oss model card says its quantization makes “gpt-oss-120b run on a single 80GB GPU (like NVIDIA H100 or AMD MI300X) and the gpt-oss-20b model run within 16GB of memory.” Meta says Llama 4 Scout fits on a single H100 GPU with Int4 (4-bit) quantization, and Llama 4 Maverick fits on “a single H100 host,” meaning a whole server.
These machines show the range, from one graphics card to a rented server in Toronto:
| Machine | Memory for the model | Vendor detail |
|---|---|---|
| NVIDIA GeForce RTX 5090 | 32 GB of VRAM | A graphics card for a desktop PC. |
| NVIDIA RTX PRO 6000 | 96 GB of VRAM with error-correcting code (ECC) | A workstation graphics card. ECC memory detects and corrects memory errors. |
| NVIDIA DGX Spark | 128 GB unified memory | Built for “AI development and testing workloads with AI models up to 200 billion parameters.” Memory bandwidth of 273 GB per second (NVIDIA blog). |
| Apple Mac Studio with M5 Ultra | 96 GB unified memory, or 256 GB or 512 GB on the version with the 80-core GPU | 1.2 TB per second of memory bandwidth. |
| DigitalOcean server with one RTX 4000 Ada | 20 GB | Charged per GPU per hour on demand, in Toronto. |
| DigitalOcean server with one H100 | 80 GB | Charged per GPU per hour on demand, in Toronto. |
In February 2026, NVIDIA raised the suggested retail price of the DGX Spark Founders Edition “due to memory supply constraints,” according to its developer forum. DigitalOcean’s availability page lists both GPU servers on demand in TOR1, its Toronto data centre, and DigitalOcean bills in US dollars. Vendor details were checked on September 27, 2026.
What sets the speed
Memory bandwidth is how fast a chip reads its memory, and the speed at which a model writes its answer rises with it. In llama.cpp’s benchmark on Apple chips, running Llama 2 7B at 4 bits, an M1 Pro with 200 GB per second of bandwidth generated 36.41 tokens per second. An M4 Max with 546 GB per second generated 83.06, and an M2 Ultra with 800 GB per second generated 94.27.
NVIDIA measured the DGX Spark with one user sending 2,048 tokens and receiving 128. On that test, gpt-oss-20b generated 82.74 tokens per second and gpt-oss-120b generated 55.37. Qwen3 235B, split across two DGX Sparks, generated 11.73.
Which model fits your hardware
Tell Derik what machine you have or plan to buy, and what the model would read. He will tell you which models fit it and what to test first.
Start a conversationOpen-weight models and their licences
The licence decides what your company may do with a model. Open-weight means the weights are published for download, and several popular licences still add conditions a business has to check. These models show the range, from the model pages checked on September 27, 2026:
| Model | Maker | Licence | What to check |
|---|---|---|---|
| gpt-oss-20b and gpt-oss-120b | OpenAI | Apache 2.0 | gpt-oss-120b has 117 billion parameters, with 5.1 billion active for each token. |
| Gemma 4 (E2B, E4B, 12B, 26B A4B, 31B) | Apache 2.0 | “Out-of-the-box support for 35+ languages.” | |
| Qwen3.8-27B | Qwen | Apache 2.0 | A context window of 262,144 tokens natively. |
| Qwen3.8-Flash-Next | Qwen | Qwen Community License 1.0 | A company that runs a “Model as a Service or AI Work Assistant business” needs a separate licence. Internal use is exempt if nothing is made available to a third party. |
| Llama 4 Scout and Maverick | Meta | Llama 4 Community License | If your products or services had more than 700 million monthly active users in the month before the Llama 4 release date, you must request a licence from Meta. If you distribute or make available the model, or a product or service that contains it, you must display “Built with Llama.” Downloads on Hugging Face are gated: you give your full legal name, date of birth and organization first. |
| Mistral Small 4 | Mistral | Apache 2.0 | 119 billion parameters, 6.5 billion active. French is among its listed languages. |
| Mistral Medium 3.5 | Mistral | Modified MIT | No rights under the licence if your company’s “global consolidated monthly revenue” exceeds $20 million. |
| DeepSeek-V4.1-Flash | DeepSeek | MIT | “552B backbone parameters,” far beyond a single graphics card. |
| Granite 4.2 (3B, 8B, 30B) | IBM | Apache 2.0 | Released August 25, 2026. French is among its tested languages. |
| Phi-4-reasoning-vision-15B | Microsoft | MIT | 15 billion parameters. |
The Apache 2.0 licence lets a business use and change a model without paying a royalty. Read any other licence in full, with your lawyer, before you build on a model. An earlier licence table and DigitalOcean server costs are in ChatGPT alternatives for business.
How to pick the best local LLM for your work
The best local LLM for a business is the one that fits its hardware and does best on a test built from its own documents. We check in this order:
- Fit the memory. Pick sizes that fit your machine at 4 or 8 bits, with room left for the context your documents need.
- Check the licence. Models under the standard Apache 2.0 or MIT licence carry none of the extra conditions in the table above. Mistral Medium 3.5’s modified MIT licence adds a revenue limit.
- Check French. If your team works in French, Mistral Small 4 and Granite 4.2 list French by name, and Gemma 4 lists support for more than 35 languages.
- Test on your own work. Run 30 to 50 real questions through two or three candidates, and compare the answers with ones your staff know are right.
Quality and speed against hosted models
The best hosted models still lead the best open-weight models, and the size of the lead depends on the measure.
- Arena ranking. Stanford’s 2026 AI Index says: “As of March 2026, the top closed model leads the top open model by 3.3%, up from 0.5% in August 2024.” It adds that six of the top ten models on the Arena Leaderboard, a public ranking of models, are closed.
- Months behind. Epoch AI reported on May 29, 2026 that since January 2026, “the most capable open-weight models have lagged frontier closed models by an average of four months” on its Epoch Capabilities Index, which Epoch describes as its aggregate measure of model capability. A frontier model is one of the newest and most capable models.
- One graphics card. In August 2025, Epoch AI found that with a single RTX 5090 (“under $2500”), anyone “can locally run models matching the absolute frontier of LLM performance from just 6 to 12 months ago.”
Some open-weight models are very large. DeepSeek-V4.1-Flash has 552 billion backbone parameters, plus 196 billion in a separate memory component called Engram. At 4 bits per weight, the backbone alone comes to about 276 GB, more than any single graphics card in the table above holds.
Quantization costs some quality. NVIDIA says its 4-bit NVFP4 format gives “near-FP8 accuracy (<1% degradation),” where FP8 is an 8-bit format. That is a vendor’s claim about its own format. OpenAI’s gpt-oss models ship quantized in a format called MXFP4, and its model card says: “All evals were performed with the same MXFP4 quantization.” Test the exact file you plan to run, at the bit count you plan to run it.
The speed figures above are for one user at a time, so time your own requests on both options, with documents as long as your team’s.
A self-hosted LLM for your whole team
A self-hosted LLM is a local model on a server your team shares. Moving from one computer to a shared server changes the defaults you rely on.
- Who can reach it. Ollama’s FAQ says: “Ollama binds 127.0.0.1 port 11434 by default,” which means only the machine it runs on can connect. The OLLAMA_HOST setting opens it to your network, so decide who may connect before you change it.
- How many at once. Ollama’s OLLAMA_NUM_PARALLEL setting allows one request per model at a time by default, and the FAQ warns: “Required RAM will scale by OLLAMA_NUM_PARALLEL * OLLAMA_CONTEXT_LENGTH.” Its default context window is 4,096 tokens, so long documents need a larger setting and more memory.
- A serving engine. vLLM offers “continuous batching of incoming requests,” meaning the GPU works on several people’s requests together as they arrive. It suits a GPU server that many people or programs call.
- A chat window. Staff need an interface, and Open WebUI is one. Since version 0.6.6, released April 19, 2025, its licence says you may not alter or remove its branding unless an exemption applies, such as having “50 or fewer users” in a 30-day period.
Someone also has to run the server. It needs security patches, backups, monitoring and access control, and each new model needs testing on your documents before it replaces the old one. Name that person before you buy hardware. If the model drafts anything that leaves the company, keep a person approving each draft, as human in the loop explains.
When your business should run a local LLM
Confidential data. If the model will read drawings, costing, customer lists or employee records, where it processes them matters. Under PIPEDA, the federal privacy law for businesses, the Privacy Commissioner’s guidelines say: “An organization is responsible for personal information in its possession or custody, including information that has been transferred to a third party for processing.” In Quebec, section 17 of the private-sector privacy act requires a privacy impact assessment and a written agreement before personal information is communicated outside Quebec.
A model on hardware your company owns, with its cloud features off, processes those records without sending them to an AI vendor. A rented GPU server runs in the host’s data centre, so the host is a third party too. Who controls the server also matters, and sovereign AI covers that question. This page is not legal advice.
Cost at volume. A hosted model charges for every token. A machine costs the same each month whether it answers ten requests or ten thousand, up to what it can handle. Your monthly volume decides which costs less.
1. Count a month of tokens. Multiply your requests per month by the tokens each one sends and receives. The example here is a request that sends 4,000 tokens, about 3,000 English words, and gets 600 back.
2. Price it on the hosted model. Multiply the input tokens by the model’s input rate per million tokens, and the output tokens by its output rate. Anthropic’s pricing page lists the rates for each Claude model. Anthropic says Claude 4.7 and later models produce about 30% more tokens for the same text, so count tokens with the model you plan to use.
3. Price the machine per month. A rented server running all month is about 730 hours, so multiply its hourly rate by 730. For bought hardware, spread the price over the months you expect to use it, and add power.
4. Add the people. Count the hours to set up the server and keep it patched, and the hours to test each new model.
5. Divide. The break-even is the machine’s monthly cost, people included, divided by the hosted cost per request. Above your break-even, check that one server keeps up with your busiest hour, and compare answer quality on your own test before you compare prices.
Full per-model rates are in Claude API pricing.
When a hosted model is the better choice
A hosted model fits better in these cases:
- The hardest tasks. The best closed models lead on both measures above. Where the best answer matters most, test the leading hosted model against your local candidate.
- Low or uneven volume. A hosted model charges only for the requests you send, while a rented server charges for every hour it runs, used or not.
- No one to run the machine. A hosted model leaves no server for you to patch or back up.
- A model larger than your hardware. If the model that passes your test needs more memory than you can buy or rent, a hosted version may be the practical way to use it.
If you choose a hosted model, get two things in writing. The first is retention. Under a zero data retention (ZDR) agreement, “Anthropic does not store customer prompts or responses at rest after the API response is returned,” according to its data retention page, and it enables ZDR per organization. Some models are left out: Anthropic says Claude Fable 5 and 5.1 and Claude Mythos 5 and 5.1 “require 30-day data retention and are not available under ZDR unless expressly authorized by Anthropic.”
The second is where requests are processed. Anthropic’s data residency page says its default “global” setting means “Inference may run in any available geography for optimal performance and availability,” and the “us” setting means “Inference runs only in US-based infrastructure.” On Amazon Bedrock, the model card for Claude Opus 5.5 shows no in-region processing in the Canada (Central) and Canada West (Calgary) regions. Its US geographic profile “Keeps data within US and Canada regions.” Other providers publish their own terms, and secure AI at work lists the settings to check.
How ThriveAI sets it up
ThriveAI is an AI engineering company in Ottawa that builds private AI systems for manufacturers and distributors in Ontario and Quebec, on their own data. Derik Lawlis leads every project and stays close to the build.
The platform is designed to keep each client’s data on its own server in Canada. You choose the model: an open-weight model on that server, or a hosted model under a written zero data retention agreement. A hosted model may process requests outside Canada, so the contract names the model and the provider plan it runs under. Private AI for business explains how the pieces fit, and AI training covers hands-on sessions for your team on your own documents. The company and how a project runs are on About ThriveAI.