Skip to content
HJ ConsultHJ Consult
esc
Start a brief
Services
Work
Field notes
About
Brand
Norsk
Englishen
Switch mode: Nordlys
Switch mode: Systems
↑↓ select, ↵ open, esc close⌘K
Mode

AI, search, knowledge base

What is RAG, and does a small business need it?

RAG, retrieval-augmented generation, means the model looks up the relevant parts of your own documents before it answers, and answers from those. You need it when the answers have to come from content that is too large to hand the model every time, or that changes often: manuals, contracts, product sheets, years of email. You do not need it for a one-page FAQ, which fits in the instructions as it is. The hard part is keeping the documents current, and making sure nobody finds what they are not allowed to see.

Published: September 27, 2026

What happens when somebody asks

  1. The question is searched against your documents. The documents were split into passages in advance, and each passage stored with a representation of what it means.
  2. The best passages are fetched. Usually a handful, chosen by meaning, often mixed with plain keyword search.
  3. The model answers from them. The instructions tell it to use only what it was given, to cite the passages, and to say when they do not answer the question.

None of the three steps is the model's knowledge. That is the point: the answer comes from your content, as it is today.

When you need it, and when you do not

You do not need it when everything the assistant should know fits on a few pages. Opening hours, prices and rules go straight into the instructions, and the answer is just as good without a search in front of it.

You need it when the content is large, scattered or changing: hundreds of product sheets, the manuals for every machine you service, contracts, the support history. Then no set of instructions can hold it, and the search decides which part the model sees.

Where it goes wrong

  • Stale documents. The index is a copy. If nobody updates it when a document changes, the assistant confidently quotes last year’s terms.
  • Bad splitting. A table cut in half, or a passage without the heading that explains it, gives the model half an answer to work with.
  • Search that misses. Names, codes and abbreviations are where meaning-based search is weakest. That is why keyword search stays in the mix.
  • Permissions after the fact. Hiding a passage in the answer is too late. It has to be filtered out of the search.
  • No test set. Without real questions to run, every change is a guess about whether it got better or worse.

How to find out in an hour

Whether a knowledge base needs RAG depends on how much there is and how often it changes, which is quicker to see in your case than to explain in general.

Describe it in the Living Brief and you get an architecture sketch, phases and a range in weeks straight back, without anyone phoning you. If you are not sure where AI would help at all, the free AI review starts from the business instead.

Field notes

Questions we get

The same questions, every time. These are the answers we give across a table.

Does RAG stop the model from making things up?

It makes it much less likely, and it makes it checkable, but it does not make it impossible. A model that answers from retrieved text can still misread it, or answer anyway when the search found nothing relevant. The fix is to have it cite the passages it used and say so plainly when it found nothing, and to test it on real questions before customers do.

Is RAG the same as training a model on our data?

No, and it is usually the better choice for a business. Training bakes knowledge into the model, where it is hard to update and hard to trace. RAG leaves your documents where they are and fetches from them at the moment of the question, so a change to a document is a change to the answer the same day, and every answer can point at its source.

Do we need a vector database?

You need some way to search by meaning, and a vector index is the usual one. It does not have to be a separate product: Postgres with pgvector does the job for most small businesses, next to the data they already have. Plain keyword search alongside it still helps, because product codes and names are exactly what meaning-based search gets wrong.

How do we know it answers well?

Collect fifty real questions with the answers you would accept, and run them every time something changes: a document, the instructions, the model. It is the cheapest insurance there is, and it turns "it seems fine" into a number you can watch.

Can it respect who is allowed to see what?

It has to, and it has to happen in the search, not in the answer. If a document is found, the model will use it, so a salary sheet in the index is a salary sheet in somebody’s answer. Filter the search by the permissions of the person asking before anything reaches the model.

Next step

Want to know what it means for you?

Describe your operation and get an architecture sketch, phases and a range in weeks straight back. Nobody phones you.