"Private LLM" has become one of those phrases that means something different to everyone who uses it. To some people it means a ChatGPT-style assistant with the company logo on it. To others it means training a model from scratch on the company's documents. Neither is what most businesses need, and the second is almost never sensible.
What it usually should mean is simpler: a language model that runs on infrastructure you control, answering questions using your own data, without that data ever leaving your network. No customer records sent to a third-party API. No contracts in someone else's logs. That is the requirement we hear most, and it is usually a contractual or regulatory one rather than a preference.
Here is what building one actually involves.
You are not training a model
This is the most common misconception, so it is worth clearing up first.
Almost nobody needs to train or even fine-tune a model to get useful answers about their own business. The approach that works is called retrieval-augmented generation, or RAG. When a user asks a question, the system first searches your data for the relevant passages, then hands those passages to the model along with the question, and the model answers from what it was given.
The model supplies the language. Your data supplies the facts. That separation matters: when a document changes, the answers change the moment the search index does, with no retraining. And every answer can show exactly which records it came from, which is what makes people trust it.
The four parts
A private RAG system has four moving parts, and only one of them is the model.
The model server. Open-weight models — published by Meta, Mistral, Google, Alibaba and others — can be downloaded and run on your own hardware. Serving software such as vLLM, Ollama or llama.cpp exposes them over HTTP, using the same API shape as the hosted services. Your application talks to it exactly as it would talk to a cloud provider, which means you can switch between the two without rewriting anything.
The embedding model. A second, much smaller model that turns text into numbers so that passages with similar meaning end up close together. It is what lets a question about "late payments" find a document that only ever says "overdue invoices".
The vector store. Somewhere to keep those numbers and search them quickly. This does not need to be exotic. SQL Server 2025 has a native vector type, PostgreSQL has pgvector, and for many businesses the right answer is the database they already run and already back up.
The ingestion pipeline. The unglamorous part that decides whether the whole thing is any good. Documents are split into sensible passages, cleaned, tagged with where they came from and who may see them, embedded and stored — and kept in step as the source changes.
Hardware, honestly
For an internal assistant used by a team or a department, the hardware is more modest than people expect. A single server with one modern data-centre GPU, or even a well-specified workstation-class card, comfortably runs a mid-sized open model for dozens of concurrent users. Models are usually run in a compressed form that trades a small amount of quality for a large reduction in memory, and for summarisation and question-answering over your own documents the difference is rarely noticeable.
What drives the size is not the number of documents. It is how many people are asking questions at the same moment and how long their answers are. Start with one machine, measure, and grow if the usage justifies it.
The real costs are elsewhere: someone has to patch it, monitor it, update models when better ones are released, and keep the ingestion jobs healthy. Budget for operating it, not just buying it.
The model matters less than you think
We have said this before about AI in business software generally, and it is doubly true here: when a private assistant gives poor answers, the cause is almost always retrieval, not the model.
If the right passage is not found, no model can answer correctly. If the passages are split mid-sentence, stripped of their headings, or drawn from three superseded versions of the same policy, the answer will be confident and wrong. Most of the engineering effort in a good system goes into the ingestion pipeline: choosing what to include, removing duplicates and outdated versions, keeping the structure of a document intact, and attaching metadata such as dates, owners and record types so the search can filter on them.
A modest open model over well-prepared data will beat the best model available over a pile of exported PDFs. Every time.
Permissions are the hard part
An assistant that can read "the company's documents" can also repeat them to anyone who asks. The HR policy is fine. The salary review spreadsheet is not.
The rule is simple to state and harder to build: retrieval must only ever return passages the person asking is allowed to see, and that check must happen during the search, not afterwards. Filtering answers after the model has read restricted material is too late — it has already used it.
In practice that means every stored passage carries the access rules of its source, and every search is filtered by the identity of the person asking, using the same permission model as the rest of your application. If those permissions are complicated — and in operational software they always are — this is where most of the design time goes. It is also the strongest argument for starting with a narrow scope.
How to know it works
Before anyone uses it, write down fifty real questions that real people ask, along with the answer and the document that answer comes from. Run them through the system. Check whether the right passages were found and whether the answers were correct.
That list becomes your test suite. Every change to the chunking, the embedding model, the prompt or the model itself gets measured against it. Without it, you are tuning by impression, and impressions of AI systems are unreliable in both directions.
Where to start
The sequence that works for most businesses:
- Pick one body of knowledge with simple access rules. Product documentation, internal procedures, or the history of one type of record. Not "everything".
- Prove it with a hosted model first, if your data rules allow it for a test set, or go straight to a private model if they do not. The goal is to learn whether the workflow saves time before buying hardware.
- Build the ingestion and the evaluation set properly. This is the part that carries over to everything that follows.
- Move the model in-house once the value is proven. Because the application speaks the same API to both, this is a configuration change, not a rebuild.
- Widen the scope one source at a time, adding the permissions each new source needs.
A private LLM is not a research project. It is an ordinary piece of business infrastructure: a database, a search index, a model server and some careful plumbing. Built that way, it answers questions from your own records, shows its sources, respects who is allowed to see what, and never sends a byte of it outside your walls.
If you are weighing whether it makes sense for your systems, our AI integration page sets out how we approach it.