RAG or fine-tuning? A straight answer for small businesses
The two ways to make an AI model work with your own information, what each is actually for, and why retrieval is almost always the right first move for a smaller company.
If you have asked anyone about putting AI to work on your company’s own information, you will have heard both terms thrown around, usually as if they are alternatives. They are not really competing options — they solve different problems, and most businesses only need one of them.
The short version
Retrieval-augmented generation (RAG) changes what the model knows. You keep your documents in a searchable store; when someone asks a question, the system finds the relevant passages and hands them to the model along with the question.
Fine-tuning changes how the model behaves. You show it hundreds or thousands of examples of the kind of output you want, and it learns the pattern — the tone, the format, the house style.
The question is not “which is better”. It is “is my problem about knowledge or about behaviour?” Almost always, for a smaller company, it is knowledge.
Why retrieval usually wins
Your information changes. A price list, a policy, a product catalogue. With retrieval, you update the document and the next answer is correct. With fine-tuning, the outdated information is baked into the model and you have to train again.
It can show its working. A retrieval system can cite the document each claim came from, so a person can check it. This matters more than it sounds. It is the difference between a tool your team trusts and one they quietly stop using after it confidently gets something wrong.
Access control still works. You can filter what gets retrieved based on who is asking, so the system never surfaces the salary spreadsheet to someone who should not see it. Once information is fine-tuned into a model, it is in there for everyone.
It is far cheaper to start and to change. Setting up retrieval is configuration and engineering. Fine-tuning is a training run plus the work of building a good dataset, and you repeat it every time things change meaningfully.
When fine-tuning is genuinely the answer
It earns its place when the problem really is about behaviour:
- You need output in a rigid format every single time and prompting keeps drifting.
- You need a specific voice — a regulated tone, a particular style of technical writing — that is hard to describe but easy to demonstrate with examples.
- You are running a very high volume of a narrow task, and a smaller fine-tuned model would be cheaper per call than a large general one.
- You work in a specialist domain with vocabulary the general models handle badly.
Notice that none of those are “the model needs to know our data”. That is the mistake people make.
What a retrieval system actually involves
It is more than dropping files into a vector database, and the difference between a demo and something usable is mostly in these details:
Chunking. Splitting documents so each piece is self-contained. Do it badly and you retrieve half a sentence, or a table divorced from its heading.
Hybrid search. Semantic search understands meaning but misses exact strings — part numbers, error codes, names. Keyword search does the opposite. Production systems run both and merge the results.
Re-ranking. Retrieve a generous set of candidates, then use a second, more careful pass to order them. This is one of the cheapest large improvements available.
Evaluation. A set of real questions with known good answers, so you can tell whether a change made things better or worse. Without this you are guessing, and most systems that quietly underperform are ones nobody measured.
A sensible first step
Pick one question your team asks repeatedly and currently answers by digging through documents. Build retrieval over just that set of documents. Get it answering well with citations, put it in front of the people who asked, and see what they actually do with it.
That is a small, contained project. It will teach you more about whether AI helps your business than any amount of strategy work — and if the answer turns out to be “not much, here”, you have found that out cheaply.
Where we land
For most of the small and medium businesses we talk to, retrieval is the right answer and fine-tuning is a distraction. We will say so, and we will tell you if yours is one of the genuine exceptions.