06 52 57 47 69
Français — non disponible sur cette pageEnglish
Free audit

Artificial intelligence

Querying your internal documents with AI without sending them to the United States

Document RAG: where your documents travel, what a local model, a European provider or a US provider's EU region guarantees, and how to choose.

Woman consulting digital documents projected in an archive room

What a document RAG is, without the jargon

An assistant that answers from your own documents almost always rests on the same mechanism, known as RAG (retrieval-augmented generation). It happens in two stages. First, indexing: your procedures, contracts, technical notes or product sheets are cut into passages a few paragraphs long, then each passage is converted into a sequence of numbers representing its meaning, called an embedding. Those vectors are stored in a dedicated database. Then, the question: when a member of staff asks “what is the notice period in supplier X’s contract?”, the question is itself converted into a vector, the database finds the closest passages, and those passages are sent to the language model with the instruction to answer only from them, citing its sources.

The model is not trained on your documents: it reads them at the moment of the question. In exchange, extracts from your documents travel with every question. Knowing where they travel is the real sovereignty question.

Where your data actually travels

Four stages, and each one can take place somewhere different.

Indexing. The documents are read, cut up and stored somewhere: on your server, with a hosting provider, or in an online vector database service. A working copy of your documents sits there.

Computing the embeddings. Turning each passage into a vector requires a model. If it is called through an API at a provider, the full text of every passage is sent to that provider, once at indexing and again at every update. This is easily overlooked: even if the model that writes the answers is hosted on your premises, an external embedding model sends the entire corpus out. The vectors themselves are derived from the text; treat them with the same care as the original documents.

The query to the model. With every question, the provider receives the member of staff’s question and the retrieved passages. The volume is lower than at indexing, but it is continuous.

The logs. On the provider’s side, requests may be kept for abuse monitoring. On the application side, conversation history, technical logs and monitoring tools also keep copies. At OpenAI, the documentation states that abuse monitoring logs are kept for up to 30 days by default; at Mistral AI, API inputs and outputs are kept unless zero data retention has been granted, on a reasoned request and depending on the plan. These periods depend on the plan you subscribe to and they change: check them when the project starts.

Three hosting options, three levels of guarantee

The model running on your own server

An open model runs on a machine you control, on your premises or with a hosting provider of your choosing, with the embedding model and the vector database in the same place. What this guarantees: no data leaves your perimeter, no model provider is involved, and no transfer question arises for this part of the processing. What it does not guarantee: the security of the machine itself, which becomes your responsibility. If the server sits with a hosting provider, the jurisdiction questions shift to that provider.

A European provider hosted in the Union

A vendor established in the European Union, running its models on servers located in the Union. Mistral AI, for instance, states in its documentation that data is hosted in the European Union by default, unless you explicitly use its US endpoint. The same page specifies that data may be temporarily transferred outside the Union to certain listed sub-processors, under the safeguards of Article 46 of the GDPR. Another point to check: depending on the plan, API data may be used for training unless you have opted out (Mistral documentation). What this guarantees: a contracting party subject to European law, and European hosting by default. What it does not guarantee: the absence of any non-European sub-processor, nor the training settings, which depend on the plan and must be checked.

A US provider with a European region

OpenAI offers data residency in Europe for eligible API projects: requests are processed in the region, through a dedicated endpoint. Its documentation specifies that this regional processing requires zero data retention or a modified abuse monitoring regime to have been granted, that the region is chosen when a project is created and does not apply to existing projects, and that some features remain limited. It also states that data sent to the API is not used for training unless you explicitly agree to it.

What this guarantees: storage and processing located in Europe, written into the configuration and the contract. What it does not guarantee: freedom from exposure to US law. The CLOUD Act, passed in 2018, allows the US authorities to require a provider subject to their jurisdiction to hand over data it controls, whether it is stored in the United States or elsewhere. Article 48 of the GDPR, for its part, provides that a decision of an authority of a third country requiring a transfer is only recognised or enforceable if it is based on an international agreement, without prejudice to other grounds for transfer. The two texts can therefore conflict, and it is the provider that finds itself in the middle. The risk is real; its scale depends on the nature of your data.

What about the Data Privacy Framework?

Transfers to certified US companies currently rest on the adequacy decision (EU) 2023/1795 of 10 July 2023. The General Court of the European Union dismissed the action for annulment on 3 September 2025 (case T-553/23), but an appeal lodged on 31 October 2025 is pending before the Court of Justice (case C-703/25 P). The two previous frameworks, Safe Harbor and then the Privacy Shield, were both struck down by the Court: do not make this one the sole basis of your compliance.

For the most sensitive data, ANSSI, the French national cybersecurity agency, has created the SecNumCloud qualification, which aims in particular to protect data against the application of extraterritorial laws. Check the list of qualified providers rather than marketing claims.

The contract, in every case

Whichever option you choose for an external service, the provider acts as a processor within the meaning of Article 28 of the GDPR. The contract must cover, among other things, instructions, security, the use of other sub-processors and what happens to the data at the end of the contract. Ask for the list of sub-processors and their locations. If a transfer outside the Union remains possible, the CNIL, the French data protection authority, sets out the approach in its guidance on how to identify and handle data transfers outside the EU and in its transfer impact assessment guide (both in French).

How to choose

The right choice depends on the documents, not on a matter of principle. Ask yourself four questions, in this order.

  1. What do the indexed documents contain? Public product documentation, internal procedures, personal customer data, files covered by professional secrecy or health data do not call for the same hosting.
  2. Who could legitimately require on-premise execution? A large corporate customer, a principal contractor, a professional obligation: these requirements often appear in contracts you have already signed.
  3. What volume do you expect? A few dozen questions a day does not carry the same constraints as hundreds of simultaneous users.
  4. Who will handle maintenance? An on-premise model with nobody to update it quickly becomes a security risk.

A mixed architecture is often the most reasonable: embeddings and vector database on your side, a European model for writing answers on everyday corpora, a local model for the most sensitive perimeter. Designing the integration so that the model can be swapped out lets you change option without rewriting everything.

The limits to know before deciding

Quality. Open models that can run on a reasonably sized server have come a long way, but they generally remain behind the largest commercial models on long chains of reasoning or poorly structured documents. Within a RAG the gap narrows, because quality depends first of all on the retrieval. Test on your own questions before drawing conclusions.

Hardware cost. Running a model locally requires a graphics card sized for it, or a rented server suitably equipped. The investment is fixed and is paid even in months when nobody asks a question, whereas an API is paid for only as you use it. At the scale of a small or mid-sized company the gap weighs heavily in the API’s favour: a local model is chosen for what must not leave the building, not to save money.

Maintenance. Updates, monitoring, backing up the vector database, re-indexing: an on-premise assistant is a service to run, not a piece of software installed once.

The documents themselves. No hosting arrangement fixes a corpus in which three versions of the same procedure coexist. Sorting out which sources are authoritative remains the first step, whichever option you choose.

Key points

  • In a RAG, your documents leave at two moments: when the embeddings are computed, and with every question. Check both.
  • A European region at a US provider localises the data, but does not put it beyond the reach of the CLOUD Act.
  • The Data Privacy Framework is valid today, with an appeal pending: do not make it the sole basis of your compliance.
  • The processing contract, the list of sub-processors and the retention settings matter as much as where the hosting is.
  • On-premise execution offers the most control, at the price of hardware and maintenance.

For the method we apply when scoping each project, see our sovereign, GDPR-compliant AI page.

Does this subject concern your business?

This article belongs to our artificial intelligence pillar. The first conversation is free and without obligation, in Nice and the surrounding area.

Read the case study: LLM Monitor: making brand visibility inside AI answers measurable