“We want an AI model trained on our business” is a reasonable starting point. But it can mean several different things, and those differences matter.
You might want a model to answer questions using current internal documents. You might want more consistent classification. You might want a specific writing style or a reliable output format.
Those goals do not all point to the same solution. Before choosing fine-tuning, separate what the model needs to know from how it needs to behave.
Start with a baseline.
Choose a small set of representative tasks and define what a good result looks like. Include difficult cases, incomplete inputs, and examples where the appropriate response is to ask a person for help.
Test an existing model with clear instructions and a few good examples. Measure usefulness, error types, cost, and response time. If a simpler approach meets the need, there may be no business case for training a model.
Keep some examples separate from development. If you repeatedly adjust the system against the same cases, it becomes easy to mistake familiarity with the test for genuine improvement.
Use retrieval for information that needs to stay current.
Retrieval-augmented generation, or RAG, finds relevant information and gives it to a model as context for a response. It can be a useful fit for questions about policies, product documentation, or other material that changes over time.
Instead of expecting the model to memorize your documents, you maintain a searchable knowledge source. Updating a policy can then mean updating that source rather than running another training job.
Retrieval is not automatic truth. The system can find the wrong passage, miss a key detail, or produce an unsupported answer. Evaluate source selection and response quality separately. Show useful references and build a clear path for “I don’t have enough information.”
Access control matters too. A retrieval system must not expose a document to someone who could not access it through the original system.
Changing what a model can look up is different from changing how it behaves.
Consider fine-tuning for a repeatable behavior.
Fine-tuning updates a model using examples. Depending on the model and training method, it can help with tasks such as applying specialized categories, following a consistent response style, or producing a domain-specific format.
A good candidate has a clear task, enough high-quality examples, and a measurable gap that prompting alone has not addressed. The desired behavior should be relatively stable; otherwise, you may be repeatedly training around a moving target.
Fine-tuning is not a reliable substitute for a live database. It is also not a guarantee of accuracy, confidentiality, or instruction-following. Treat it as a candidate improvement that needs evidence.
- Review your data rights. Confirm you can use the examples for the intended training arrangement.
- Remove unnecessary sensitive information. Training data should contain what the task needs, not everything you happen to have.
- Check consistency. Conflicting labels or weak examples teach the wrong behavior.
- Hold out evaluation cases. Test generalization on material the model did not train on.
Sometimes the answer is both—or neither.
A system can use retrieval for current information and a fine-tuned model for consistent behavior. But combining approaches adds operational work, so each part should earn its place.
Some tasks are better solved with ordinary software. If a rule is deterministic, a calculation must be exact, or a database query can answer the question reliably, a language model may introduce uncertainty without adding value.
Judge the whole workflow.
A model score is only one part of the result. Consider what happens when it is wrong, who reviews uncertain outputs, how sensitive data moves through the system, and whether the overall process actually saves effort.
Plan for monitoring after launch. Your documents, customers, and inputs will change. An evaluation that passed once is not a permanent certificate of reliability.
The useful question is not “Can we fine-tune a model?” It’s “Which approach delivers a dependable improvement for this job, at a cost and risk we can support?”
Start there. The technology choice becomes much clearer.