RAG or fine-tuning: the question is almost never framed correctly
It is one of the first questions asked on an AI project. It pits two techniques against each other that answer different needs, and in most cases the answer is “neither, not yet”.
Christophe Bellec ·
Short answer
RAG gives the model knowledge it does not have: your documents, your data, your context. Fine-tuning changes its behaviour: output format, tone, adherence to a taxonomy, one very repetitive task. A knowledge problem is not solved by fine-tuning, and a format problem is not solved by RAG. In practice, more than nine projects out of ten are a RAG problem or simply a prompting problem.
Two different problems
A language model has two things: what it knows, and how it behaves. RAG acts on the first, fine-tuning on the second. Confusing them means paying a lot for a result that does not improve.
- RAG: supplying knowledge
- You retrieve the right excerpts at question time and hand them to the model. The knowledge stays outside: you update it by updating the source, with nothing to retrain. This is what you need when the model knows nothing about your products, your procedures or your customers.
- Fine-tuning: shaping behaviour
- You partially retrain the model on input/output examples so it consistently adopts a form: always produce this JSON, always classify against this taxonomy, always write in this register. Whatever knowledge gets baked in stays marginal and, above all, frozen.
The simplest test: if the right answer changes when your data changes, it is RAG. If the right answer keeps the same shape whatever the data, it might be fine-tuning.
Why RAG wins almost every time
- Updates are immediate: a document corrected this morning is used in this afternoon’s answer.
- Answers are traceable: you can cite the source, so you can verify, so you can correct.
- Permissions stay enforceable: you filter at retrieval, which is impossible on a trained model.
- The provider stays replaceable: your value sits in retrieval, not in model weights.
- Errors are diagnosable: a bad answer comes from a bad retrieved excerpt, which you can see and fix.
That last point is underrated. A fine-tuned model that answers badly is a black box: the only available move is to rebuild a dataset and retrain. A RAG system that answers badly shows you exactly which documents it used.
When fine-tuning genuinely earns its place
There are legitimate cases, usually downstream of a system that already works:
- A single, repetitive, high-volume task: classifying, extracting fields, normalising. A small specialised model can be faster and markedly cheaper than a large general one.
- A strict output format that instructions do not hold reliably, even with structured output.
- A very particular register or domain terminology that is hard to obtain by instruction alone.
- A latency or cost constraint that forces a smaller model, which then has to be specialised to compensate.
In all four, you already know the task, you already have real examples, and you are optimising a system that works. That is very different from “we don’t know what to do, so let’s fine-tune”.
What fine-tuning really costs
Compute is not the issue. The cost sits everywhere else:
- Building a clean dataset: hundreds to thousands of correct examples, annotated by people who know the domain.
- Building a separate evaluation set, without which you cannot say whether the model improved.
- Doing it again every time the task changes meaningfully.
- Doing it again every time the base model changes, and base models change fast.
- Living with a model frozen on dated knowledge, which usually brings you back to adding RAG on top anyway.
In other words, fine-tuning turns an engineering problem into an annotated-data problem. That is a defensible choice, but it should be made knowingly.
What to try first
In order, cheapest first. Stopping at step 2 or 3 is common.
- Sharpen the instruction: role, constraints, edge cases, what not to do. There is often a lot of headroom here.
- Force structured output instead of hoping for a format in free text.
- Add a few representative examples to the context.
- Split the task: two simple, checkable calls beat one call that has to get everything right at once.
- Add retrieval if the model is missing knowledge.
- Only then consider fine-tuning, on a task that is stable and measured.
Both together, in the right order
When both are useful, order matters: build retrieval first, measure, then possibly specialise the model on one precise step of the pipeline. For example a small fine-tuned model to rewrite the query or classify intent, and a general model to compose the answer from the retrieved excerpts.
Starting with fine-tuning means optimising one step of a system that does not exist yet.
Deciding with a single question
“If I hand the right documents to a competent person who has never worked here, do they answer correctly?” If yes, your problem is a retrieval problem: that is RAG. If no (because it takes an implicit internal convention, an exact format, a jargon), then, and only then, fine-tuning enters the conversation.
Frequently asked questions
Can fine-tuning teach the model our data?
Badly, and in frozen form. Fine-tuning mostly adjusts behaviour; whatever knowledge it absorbs is hard to verify, impossible to cite and out of date as soon as your data changes. For knowledge, retrieval is more reliable and far cheaper.
How many examples does a useful fine-tune need?
It depends on the task, but the order of magnitude is hundreds to thousands of correctly annotated examples, plus a separate evaluation set. If you do not already have those examples coming out of production, that is usually the sign it is too early.
Is RAG more expensive to run?
Each request sends more context, so more tokens. In exchange there is no training, no retraining and no dataset to maintain. Over the life of a project, RAG is almost always the cheaper of the two.
How do we know which one our case calls for?
By starting from the use case and the available data, not from the technique. That is what a short audit is for: identifying what the model must know, what it must produce, and how fast each of those changes.