Answer · updated

Fine-tuning or RAG: which one do you need?

They solve different problems. Retrieval changes what the model can see, so it fixes wrong or out-of-date facts. Fine-tuning changes the model, so it fixes tone, format and task behaviour. Most production systems need retrieval first, and fine-tune only once they can show what retrieval alone cannot fix.

The two are often presented as competing options. They solve different problems, and confusing them is one of the more expensive mistakes in an AI programme.

What each one actually changes

Retrieval-augmented generation, introduced by Lewis and colleagues at NeurIPS 2020, leaves the model alone and changes what it can see. At the moment of the question, the system finds relevant passages in your documents and asks the model to answer from them. The knowledge lives in the index, so updating it is a matter of re-indexing a document, not retraining anything.

Fine-tuning changes the model itself. It teaches a format, a tone, a vocabulary or a task — the shape of the output rather than the facts in it. Work published at ACL 2020 by Gururangan and colleagues showed that continued training on domain text improved performance across every domain and task tested, so this is well-established for specialised language.

How to choose

Ask what is wrong with the current output. If the model gives a well-formed answer that is factually wrong or out of date, the problem is knowledge, and retrieval is the answer. If the model knows the facts but answers in the wrong register, ignores your house format, or cannot follow your domain’s conventions, the problem is behaviour, and fine-tuning is the answer.

One more test: how often does the underlying information change? Facts that change weekly belong in a retrieval index. Facts that never change, and conventions that are stable, can be trained in.

Usually the answer is both, in order

Most production systems end up doing retrieval for knowledge and a light fine-tune for behaviour. The order matters: build retrieval first, measure it, and only fine-tune once you can show what retrieval alone cannot fix. Fine-tuning first is how teams end up with an expensive model that still cites last year’s policy.

What neither one fixes

Retrieval reduces hallucination; it does not remove it. A preregistered study published in the Journal of Empirical Legal Studies in 2025 found commercial retrieval-based research tools still hallucinating between 17% and 33% of the time. And simply stuffing more documents into the prompt does not help: research in the Transactions of the ACL in 2024 found models used information best at the start and end of their input, with performance dropping significantly when the relevant passage sat in the middle. Ranking and evaluation still do the real work.

Related questions

Is fine-tuning a way to teach a model our company’s facts?

It is a poor way. Facts that change belong in a retrieval index, where updating one means re-indexing a document rather than retraining a model. Fine-tuning suits stable behaviour: format, tone and task conventions.

Does retrieval remove hallucination?

No. A 2025 Journal of Empirical Legal Studies evaluation found commercial retrieval-based tools still hallucinating between 17% and 33% of the time. Retrieval lowers the rate; evaluation and human review handle what is left.

Can we just put all our documents in the prompt instead?

Long context is not a substitute for ranking. Research in the Transactions of the ACL in 2024 found performance dropped significantly when the relevant information sat in the middle of a long input.

Related Qylis capability

AI Applications

Build generative and agentic AI applications that integrate with the systems you already run.

Discuss your requirement