Skip to main content

How to connect an LLM to internal company data without reindexing everything

The reflex, when a company wants to “plug AI into our data”, is to copy everything into a vector database. That is usually the longest, most expensive and most fragile route. Here are the three real options and how to choose between them.

Christophe Bellec ·

Short answer

In most cases you reindex nothing. You let the model call the systems that already hold the data (through the existing application layer, with its permissions) and you only build a vector index for unstructured content that no API can query. Wholesale reindexing is a default decision, rarely a motivated one.

The question is almost always framed wrong

“How do we put our data into the AI?” assumes the model has to own the data. It does not. At the moment it answers, an LLM needs the right excerpt in the right format, not a copy of your information system.

That distinction changes everything. Owning the data means copying it, synchronising it, replaying your permission rules inside a second system, and living with a permanent gap from the source. Reaching it when needed means none of that.

So the real question is: for each kind of information the model needs, is there already a way to get it? In a company that has been running for a few years, the answer is yes more often than people expect.

Three ways to give access to your data

1. Tool calls on your existing APIs
The model calls a function you expose: “find a customer”, “list this month’s orders”, “read product sheet X”. Your backend answers, with your business rules and your permissions. No copy, no synchronisation lag, and exact data rather than approximate data.
2. The search you already have
A full-text engine, an Elasticsearch index, the search built into your document management system or your wiki: if it gives a human decent results, it gives the model decent results. You expose it as a tool and let the model phrase the query and read what comes back.
3. A vector index (RAG)
Useful when the information is locked inside unstructured text (PDFs, minutes, contracts, documentation) and no existing search retrieves it properly. It is the only option that involves indexing, therefore synchronising, therefore maintaining.

These options combine. A genuinely useful internal assistant typically queries one or two business APIs and a single document index, not a vector database holding the entire information estate.

When a vector index is genuinely necessary

Three conditions have to hold at the same time. If even one is missing, there is almost always a simpler answer.

  1. The information only exists as free text: it is neither in a database nor exposed by an API.
  2. The question does not translate into a filter. “Which orders exceed €10,000?” is a query, not a semantic search.
  3. Users’ vocabulary differs from the documents’ vocabulary. That gap is exactly what vector search adds over keyword search.

The deciding test: if a colleague would find the answer in thirty seconds using the existing search bar, a vector index solves no problem. It creates one: maintaining it.

Access rights decide the architecture

This is where projects fail latest, and therefore most expensively. Not all your documents are readable by everyone. A vector index, by nature, is flat: it knows nothing about roles, departments or exceptions.

So before indexing anything, decide how permissions travel with the data:

  • Filtering at retrieval: each chunk carries the identifiers allowed to read it, and the query filters on them. The only approach that holds at scale.
  • Partitioned indexes: one index per confidentiality boundary. Simple, but unmanageable as soon as boundaries overlap.
  • Filtering after generation: let the model answer, then redact. Never do this: the information has already leaked into the answer.

The tool-call option sidesteps the whole subject: your application answers, so your authorisation model applies, the one you already maintain.

Freshness is the hidden cost

An index is never up to date; it is up to date as of the last run. As long as the content is documentation that changes quarterly, nobody notices. The moment it covers stock, pricing, availability or order status, the lag becomes a wrong answer delivered confidently.

The practical rule: the faster a piece of data changes, the more it belongs to a live read rather than an index. Transactional data goes through tool calls; stable content can be indexed.

A reference architecture

This split covers the large majority of internal use cases, and it can be built up incrementally.

  1. A tool layer exposed to the model, defined by explicit contracts: name, parameters, return format, possible errors.
  2. Those tools call your existing services, carrying the current user’s identity. Never the database directly.
  3. A separate vector index, fed only by the unstructured content that justifies it, with permissions attached to each chunk.
  4. A composition step that assembles the context, cites its sources, and is able to say it found nothing.
  5. Logging: which question, which tools called, which chunks used, what it cost. Without that, you can neither improve nor audit.

The mistakes that show up every time

  • Indexing everything “to see”: you get a huge, expensive index nobody can describe and nobody knows who may read.
  • Giving the model direct database access: you bypass ten years of business rules written into the application layer.
  • Leaving permissions until last: the single most common reason a project never reaches production.
  • Forgetting the “I don’t know” case. An assistant that invents because it found nothing is worse than an empty search.
  • Measuring only satisfaction. Without a reference set of questions, you cannot tell whether a change improved or degraded answers.

Where to actually start

Take one question colleagues ask every week and follow the data: where it lives, who may see it, how fast it changes, whether a way to fetch it already exists. You will then know which of the three options applies, and you will have a reliable answer in days rather than a months-long indexing programme.

Frequently asked questions

  • Do you always need a vector database to do RAG?

    No. RAG means retrieving information before generating an answer; retrieval can come from an API, a full-text engine or a vector index. The vector database is one means among several, useful mainly when the question’s vocabulary differs from the documents’ vocabulary.

  • Does our data end up with the model provider?

    The excerpts sent in the context do pass through the provider. That is a design criterion: depending on sensitivity, you pick suitable hosting, limit what is sent, anonymise, or use a model deployed inside your own environment. It is decided before implementation, not after.

  • How long before a usable first result?

    On a single, well-bounded use case, a few days are enough to know whether the approach holds: you wire one tool onto an existing API and measure answer quality on real questions. Industrialisation comes afterwards, and only if those results justify it.

  • What if our data is poorly structured?

    That is the most common situation, and AI does not fix it. A preliminary audit exists precisely to separate what is usable as-is from what needs tidying up first, work that benefits the whole company, regardless of the AI project.

How I can help on this

Get in touch

More articles