AI knowledge base agent: answering from your own documents
The information exists in your company — but it is scattered across procedures, contracts, an intranet and the memory of a few people. Finding it wastes a considerable amount of time, and when someone leaves, their knowledge leaves with them. Your agent turns that scattered documentation into an instant, reliable answer. Hosted in France — on local inference or an isolated resource — it cites its sources and respects access rights.
Updated on
The consumer procedure is different (14 days) — I am setting it aside here.
⛓ Source · Returns_procedure_v4.pdf, §2 — Intranet
I have invented nothing: everything comes from the procedure cited — to be reviewed before sending.
✎ Action · draft ready for review — you approve the sending
An AI knowledge base (RAG) agent answers questions based on your own documents — procedures, contracts, product sheets, intranet — rather than on generic knowledge. It finds the relevant passages, formulates a clear answer and cites its sources, without inventing. Hosted on local inference or in France, it respects confidentiality and access rights: everyone only queries what they are entitled to see. It also serves as the knowledge foundation for your other agents (support, legal, accounting). Live in one to two weeks.
Illustrative reference points describing our offer — to be confirmed by a pilot on your own documentation.
What RAG is, explained simply
RAG stands for “retrieval-augmented generation”. Instead of answering from memory like a general-purpose AI, the agent starts by finding the relevant passages in your documents, then writes an answer based on those passages alone.
! The issue
The information is there, but scattered: procedures, contracts, product sheets, intranet and the know-how of a few people. Finding it costs time, and when someone leaves they take part of the company's memory with them. Consumer AI solutions, for their part, paraphrase the web, ignore your internal truth, sometimes invent — and send your documents to a third party often hosted outside Europe and subject to the Cloud Act.
✓ Our answer
A RAG agent is only of interest if it is faithful to your sources and sovereign by design. It relies solely on your files, cites its passages and says honestly when the information does not exist. Local inference or an isolated resource hosted in France, access rights respected, human oversight: the time saved on searching is never paid for in lost confidentiality. The aim is not to replace your experts, but to make their knowledge reachable — the agent assists, the human decides.
Your internal documents: sovereignty, access rights & compliance
Your internal documents are among your most sensitive assets. Here is how the architecture of our agents protects them, source by source.
Local inference
The agent can run on a machine belonging to the company: no document leaves the network, nothing passes through a cloud.
Hosting in France
Otherwise, a dedicated and isolated resource, hosted in France under French law — your documents: processing and access within the European Union targeted by the architecture.
Reduced extraterritorial exposure
Exposure of your documents to the Cloud Act and FISA 702 is reduced by design; location alone does not guarantee immunity.
Access rights respected
Everyone only queries the documents they are entitled to see: the agent never reveals a source to an unauthorised person.
Sourced answers, no invention
Every answer cites the passages used; where there is no source, the agent says so rather than inventing.
AI Act: governed deployment
An agent strictly in support; no answer is binding on its own; traceability and human oversight from end to end.
What depends on the architecture chosen These points are not general guarantees: they are settled deployment by deployment, in the quotation.
- The applicable location is that of the architecture set out in the quotation and verified before commissioning.
- Local execution is announced only for the configuration explicitly described and accepted in the quotation.
- The applicable isolation depends on the deployment mode set out in the quotation; no dedicated isolation is presumed.
- The events logged, their content, their retention period and who may access them are defined for the deployment chosen.
See the agent at work
4 real situations, taken from those that come up most often. Pick one: the exchange unfolds as it would in your organisation.
A scripted demonstration. These exchanges show how the agent behaves — its sources, its refusals, what it leaves to your teams. Nothing is sent from this page, no model is queried here, and the matters named are fictional. That is precisely what we promise your data.
The behaviours shown here — monitoring, automation rules, routing and reminders — are configured with you during deployment, from your tools, your rules and your thresholds.
The architecture points named in these exchanges — location, local execution, isolation, encryption, role-based access, logging — are not a guarantee attached to the demonstration: they are those of the architecture set out in your quotation, and verified before commissioning.
· Seven pairs of documents contradict each other on the same question. All seven were served in an answer this month.
· Forty-one documents have no identifiable author. When one of them becomes wrong, there is nobody to tell.
· A document last updated in 2019 underpinned 340 answers this month. It is the most consulted in the base.
· Two hundred and fourteen documents were never served in twelve months. Twelve documents carry 61% of the answers. morning-watch_4-flags.pdf7 contradictions · all served
⛓ Source · 1,480 documents, answer log, author and date metadata
What I record: it is the most consulted document in your base. Last modified: 2019. It underpinned 340 answers this month.
What I do with it: I keep using it, and I serve it dated. An old document is not a wrong document, and a procedure stable for seven years is a good procedure — withdrawing it would cost 340 answers a month and correct none of them. If it is to leave the base, it leaves the day you say so, and I hand you first the list of the 340 answers to be re-grounded.
What I have always done, and it nearly suffices: I give its date every time I use it. "According to [document], modified in 2019…" The reader decides what to make of it — and that is very different from an answer that does not say where it comes from.
What I also propose: that it be reread. Not because it is old: because it is the most used. A document read 340 times a month deserves an annual reread; a document read twice deserves none, and that is what your base does at the moment — the reverse.
What I supply: the list of documents ranked by answers underpinned, with their dates. 1-document_340-answers.pdfReread the most read, not the oldest
⛓ Source · answer log, modification dates, 1,480 documents
Routing follows who can settle it: a contradiction goes to both documents' authors, together — separately, each assumes the other is current; a document with no author to the base's owner, with the answers it underpins; a heavily used old document to its author, once a year; a document never served to nobody — it is information, not a problem.
With a chase: 7 days on a contradiction, monthly on the rest. Then a monthly summary: by document and by unanswered question, never by the person querying.
What this morning has already brought to light: 7 contradictions, every one of them served at least once this month, 41 documents with no author to write to, and above all 12 documents carrying 61% of your answers while 214 sleep. From tomorrow: those twelve re-read once a year instead of never, 340 answers a month resting on a text that is certain again, and the search that used to cross three folders handed back as one sourced, dated sentence. The access is yours: opened by role, logged, withdrawn with a word — your documents stay inside your walls, and each person sees only what they are entitled to see. Every answer carries the document, its date and the exact passage: it can be checked in ten seconds, and that is what makes it possible to contradict it when it is wrong — the net tightens every month, contradiction by contradiction, and today's seven will not come back. The next step is ready: your documents ranked by the number of answers they underpin, with their date and their author where there is one — name an owner for the first twelve, the rest can wait.
✎ Framework · source and date on every answer, no merging of contradictory sources
Why "more recent" is not the right rule: a recent note may address a special case without replacing the general rule. An old document may be the reference, updated elsewhere without its date moving. And a file date is not a content date — a document resaved changes date without changing a word.
What I do: I give both, with their dates and titles, and I say explicitly that they diverge. "Two documents answer your question and they do not say the same thing. Here are both passages."
What I build instead of a smoothed answer: the divergence sheet — both passages side by side, their dates, their authors, and the exact sentence that sets them against each other. A smoothed answer is more dangerous than two contradictory ones: it removes the reader's most useful piece of information, that the question is unsettled. The sheet goes to both authors together — the one act that makes two documents converge without a third party settling it for them.
What that gives: seven contradictions found, seven served as they stand this month to 61 people in total. All seven were reported to both authors concerned, together.
What I note in passing: none of them was known. 7-contradictions_61-answers.pdfA smoothed answer hides that nothing is settled
⛓ Source · 7 pairs of documents, 61 answers served
What I record: 41 documents have no author recorded. They underpinned 612 answers this month — these are not marginal documents.
What that changes concretely: when I find a document contradicting another, or that should be reread, the flag goes to the base's owner, who first has to find out who knows.
What I bring to shorten that search: for each of the 41, what the document says about itself — the department it describes, the person it names as the point of reference, the previous version and who uploaded it. Those three clues point to three different people more often than to one: whoever uploads a document is not whoever answers for it. The name is set, not guessed — and set by me, it would be indistinguishable from a verified one.
What I propose: a field "who answers for this document", distinct from author and uploader. They are not the same person in 9 cases out of 10, and it is that one I need.
What I supply to start: the 41, ranked by answers underpinned. The first eight carry 400 between them — eight names to set settles two thirds of the subject. 41-documents_612-answers.pdf8 documents carry 400 answers
⛓ Source · 41 documents with no author, 612 answers underpinned
Ingestion: 1,480 documents taken from your six shared directories and your document management tool — word processing, spreadsheets, slide decks, plus 212 scanned PDFs put through character recognition. Each document enters with its source path, its content date where one exists, and its author where one is recorded.
Indexing: each document is split into passages, and it is the passage that is indexed, not the file — which is what allows citing the exact line rather than pointing at an 80-page document. 1,480 documents yield 41,700 passages.
What ingestion choked on, and I publish it: 37 documents out of 1,480 came in degraded — 29 scans too faint to be read reliably, 8 spreadsheets whose data lives in formulas rather than cells. 2.5 %. They are marked as such and I do not use them to found an answer: I point to the source file, with its path.
The accounting doctrine and procedures in particular: that corpus has a property the others lack — it is dated by nature, each application note carrying its financial year. So I index it with its year of application, and a question about 2024 does not surface a 2026 procedure. Of your 214 accounting doctrine and procedure documents, 31 carry no year: I serve them saying so.
Summaries of long documents: 68 of your documents run past 40 pages. For each I produce a one-page summary in which every sentence carries the page number it came from — a summary of a long document with no page references is a text nobody can check, and therefore one that gets copied unchecked.
Case memory: when a question bears on an open case — an audit under way, a year-end close, a dispute —, I attach the question to the case and keep what was served, so the next question starts where the last one stopped without the context being restated. Case memory is about the case, never about the person asking.
⛓ Source · 1,480 documents ingested, 41,700 passages indexed, 37 degraded entries, 68 documents over 40 pages
The employee FAQ: I took 4,900 questions asked over twelve months, rephrased with nothing that identifies them, and drew out 62 questions that recur at least ten times. The first six carry 31 % of the volume: expense claims, leave days, equipment, remote work, sick leave, employment certificates. Every answer cites the document and its date, like any other answer.
Onboarding: new joiners account for 38 % of the questions among those 62 during their first six weeks. So I build an onboarding path of 22 entries, ordered by what a person actually meets in order — badge and access, expense claim, leave, equipment — and not by org chart. It is not pushed: it is available, and it closes when the person stops opening it.
The support chatbot's foundation: those 62 answers and the 41,700 passages are one and the same foundation. The chatbot has no base of its own: it queries this one, and so inherits the central rule — two sources that diverge are both returned, with their dates. A support chatbot with a separate corpus always ends up answering something other than the base, and nobody knows which of the two is wrong.
The figure that does not suit me: among the 62 frequent questions, 9 have no answer in the base — they bear on rules that were never written down, only practised. I serve them as unanswered and give the department to ask, and I hand you the 9 questions ranked by volume: that is the list of the 9 notes missing from your base, and it is worth more than the other 53 combined.
⛓ Source · 4,900 rephrased questions, 62 frequent questions, 22 onboarding entries, 9 unanswered
What I answer: "I find no answer to this question in the base. The subject sits with [department]."
What I do not answer: a likely answer, an answer deduced from a neighbouring document, nor an answer prefaced with "I think that". "I think" is read as an answer, and it gets repeated without the caveat.
What I also refuse, and am often asked for: showing a confidence percentage. A "78% confidence" helps nobody decide — it gives the reader permission to accept, with the feeling of having been careful.
What I do with unanswered questions: I report them by subject and as counts. Over the month, 1,240 questions asked, 380 with no answer in the base, and the 380 group into 46 subjects.
What that gives: forty-six subjects to document, ranked by frequency. It is your base's writing plan, and it is written by the people querying it. 380-questions_46-subjects.pdfThe writing plan written by the readers
⛓ Source · 1,240 questions, 380 unanswered, grouping by subject
What I record: 214 documents of 1,480 underpinned no answer in twelve months. Twelve documents carry 61%.
What that does not mean: that they are useless. A document never served may cover a rare and serious case — an incident procedure, an annual obligation, a case that arises every three years. Those are exactly the documents nobody consults and that are needed one day.
What I propose instead of deletion: that they be looked at once, and sorted into two piles: those covering a rare case, which must stay and be dated, and those that are out of date, which are a decision and not an automatic tidy-up.
What I change in the corpus myself: nothing. No document removed, none demoted in my answers — a document I stopped serving would become invisible without anyone having decided, and that is a deletion that does not say its name. What I hand you instead, and it is ready: the 214, sorted into two piles, each with the date of its last real change — not the file date — and the department that filed it. The review fits in half a day with two people, and every decision stays yours, document by document. 214-never-served_12-carry-61.pdfA document never served is not useless
⛓ Source · 1,480 documents, 12-month service log
What is kept: the question reworded without what identifies it, the document served and its date, whether an answer was found or not, and contradictions spotted.
What is not kept: who asked which question, no history per person, no count per department, and no profile of subjects consulted.
Why the first line matters more here than elsewhere: a knowledge base is queried when you do not know. An employee who thinks their questions are tracked looks elsewhere — they ask a colleague, or they guess. This month's 380 unanswered questions would never have been asked, and the writing plan would not exist.
And no count per department: "department X asks twice as many questions" reads as a judgement on department X, when it describes missing documentation on its scope.
What the monthly summary contains: contradictions, subjects with no answer, heavily used old documents, and documents with no author. Four indicators about the base. what-is-kept.pdf4 items kept · 4 impossible
✎ Framework · retention periods to be set by the company
What I record: 96 documents in your base have no content date. The system displays a file date, which is the last time somebody opened and resaved it.
What I refuse: presenting that date as the content's. A document resaved in 2026 and written in 2017 would look current, and that is the worst of both worlds: old and reassuring.
What I say: "according to [document], undated…". Two words, and the reader knows what they are holding.
What it produces: it is not comfortable, and that is intended. Of 96 undated documents, 31 were dated by their authors within two months of going live — simply because the word "undated" appeared in answers those authors were reading themselves.
What I do to go and get the other 65: in each one I surface the strongest internal clue — a budget year cited, a regulation mentioned, a department since dissolved, a named tool version — and I put it to the author: "this document cites a department abolished in 2019: date it, or tell me it still stands." The date itself is applied by the person who answers for the document: applied by me, it would be indistinguishable from a real one, and that is exactly what you asked me to avoid. 96-undated_31-fixed.pdf"Undated" in two words
⛓ Source · 96 undated documents, 31 dated since
Your case is not here? That is exactly what a 15-minute conversation is for. Book the free audit →
What the agent does — and what it feeds
Each use corresponds to an agent we deploy. The knowledge base also serves as the foundation for your other agents, subject to your approval.
Ingestion & indexing
Ingests and indexes your documents (PDF, Word, web pages, intranet) and updates itself when they change.
Unified sourced search
Answers in plain language from several sources at once, citing the passages used.
Summarising long documents
Summarises a contract, a procedure or a report and extracts the key points, with the source to back them.
Foundation of the support chatbot
Feeds the support chatbot with reliable answers based on your customer documentation.
Employee FAQ & onboarding
Answers internal questions and speeds up new joiners' familiarity with your procedures.
Memory of the case files
Builds up procedures, doctrine and case files: the knowledge stays in the base even after someone leaves.
Accounting doctrine & procedures
Instantly find a rule, a procedure or a precedent in the firm's files.
Need to go further?
These agents handle a different business process, with their own owner and their own price. They are added to this one.
FAQ & product advice
Answering from the catalogue: availability, specifications, sourced purchasing advice.
Product FAQ agent for e-commerce from 527 € excl. VAT / month AI for e-commerce →Internal FAQ agent
The same questions come round: where to find a given form, what the rule is on a given subject, who to contact for what.
Internal FAQ agent (staff intranet) from 521 € excl. VAT / month Discover the agent →Client accounting support
A practice spends a great deal of time on two things: answering its clients' everyday questions and asking for the documents that are missing.
Client accounting support agent (practice) from 653 € excl. VAT / month Discover the agent →In 15 minutes we identify the most relevant agent — without oversizing the project.
How much time can a team win back?
By making information instantly reachable, the agent cuts the time lost searching, speeds up the onboarding of new joiners and limits the repetitive questions put to experts. Gains to be validated according to the volume of your documentation and the size of your teams.
The stages of your AI agent project
Audit & scoping
15 minutes to target the use case with the best return.
Quote or direct sign-up
A catalogue offer is bought online; a specific need gets a costed quote.
Design
We design the agent and its guardrails.
Integration & testing
We connect your tools to the agent, which is itself hosted in France.
Rollout
Going live and training your team.
Operation
Continuous supervision and improvement.
One knowledge base, one agent
A document agent (RAG) that answers from your sources and cites its passages, installed and operated for you. Prices exclude VAT — annual subscription, the time it takes for the gains to settle in.
Setup + controlled subscription
- Installation, configuration and training for your teams
- Operation, human oversight, updates and support
- Sovereign hosting in France, a dedicated and isolated resource
All inclusive, no setup fee
- Setup included (installation, configuration, training)
- Operation, human oversight, updates and support
- Sovereign hosting in France, managed end to end
On site, you own it
- Hardware installed on your premises (you own it)
- French / European AI models run locally
- Secure remote maintenance (Pro support included)
Four guarantees that matter to your documentation
Related resources
Your questions, our answers
What is a knowledge base (RAG) agent?
Can the agent invent answers?
Which documents does it work on?
Are my confidential documents protected?
How does it differ from a customer support chatbot?
Do we have to learn a new tool to use it?
Does the agent state that it is an artificial intelligence?
How long does it take to deploy this agent?
The agents your knowledge base feeds
Let us estimate the potential of your documentation
15 minutes to identify the priority sources — hosted in France, supervised, sources cited, with no commitment.