Prompting, RAG or Fine-Tuning? Pick in This Order
A team asks a language model to answer customer questions and the answers come back wrong. The instinct is to say the model needs training. Sometimes that is true. More often the model was never shown the right information, or was never told clearly what a good answer looks like. Those are two different problems, and neither is solved by training.
Prompting, retrieval-augmented generation (RAG) and fine-tuning each fix a different kind of failure. Choosing the wrong one costs weeks.
Name the Failure Before the Fix
Ask what exactly is going wrong. The answer usually falls into one of three buckets.
Wrong format or tone
Prompting, with examples and clear instructions
Missing or outdated facts
RAG, so the model reads your documents at answer time
Right facts, wrong behavior every time
Fine-tuning, to change how it responds
That table is the whole idea. Everything below is detail.
Start With the Prompt
Prompting is the cheapest experiment you can run, so it goes first. Rewrite the instruction, add two or three worked examples, say what the output must look like, and test again on a dozen real inputs. Many failures disappear here. Plenty of "the model is bad at this" complaints turn out to be "the instruction was vague."
Guides across the industry give the same advice: begin with prompting for every new use case, because it ships in hours and needs no infrastructure. Treat that as a rule of thumb, not a measured law, but it matches how most working teams behave.
Move to RAG When the Problem Is Knowledge
A model only knows its training data plus whatever sits in its context window. If answers depend on a handbook that changed last month, a price list or a ticket history, no amount of prompt polish helps. The model cannot quote what it never saw.
RAG fixes that by fetching relevant passages from your own documents and placing them in front of the model with the question. The knowledge stays in a store you control, so updating it means editing a document, not retraining anything. It also lets you show sources, which matters when someone asks why the bot said what it said. If you want the mechanics, our first RAG app walkthrough covers chunking and retrieval step by step.
Reach for Fine-Tuning Last
Fine-tuning changes the model's weights. That makes it the right tool when the facts are fine but the behavior is not: a strict JSON schema the model keeps breaking, a house writing style, a classification task where you have thousands of labeled examples.
Cost used to be the argument against it. Parameter-efficient methods such as LoRA changed that. Per Together AI's documentation and similar summaries, LoRA trains only a small fraction of the weights, roughly 0.1 to 1 percent, which is why it has become the usual way to fine-tune. You still need clean training data, an evaluation set and someone to own the result when the base model is upgraded.
- 1
Prompt
Instructions, examples, output format
- 2
Retrieve
Add RAG when the model lacks your facts
- 3
Adapt
Fine-tune when behavior still misses
- 4
Combine
Most production systems end up using more than one
The Mistake Worth Calling Out
The expensive mistake is fine-tuning to teach a model facts. It feels logical, since training is how models learn. But facts baked into weights go stale, cannot be cited and are hard to correct. A policy changes on Monday and the model confidently repeats the old one until you retrain. Put facts in retrieval and keep fine-tuning for style and structure.
The reverse mistake is rarer but real: stuffing a long style guide into every prompt and retrieved context when a small adapter would produce the format reliably at lower per-request cost.
Measure Before You Escalate
Whichever rung you are on, build a small test set first. Twenty to fifty real questions with known good answers is enough to start. Run every change against it. Without one, you cannot tell whether RAG helped or only changed the wording, and teams end up arguing from impressions.
20-50
Real questions with known good answers
1
Change at a time between test runs
3
Failure types: format, knowledge, behavior
Where the Course Fits
SkyTrainings' Generative AI Training course lists prompt engineering, LangChain retrieval, vector databases and fine-tuning as separate modules, which is the same ladder this article describes. You learn to try the cheap option first and to recognise when it has stopped working. If that is the order you want to learn them in, see the Generative AI Training course and its current batch dates.