+33 (0)1 87 66 00 65 · Monday to Friday, 9am–6pm Free audit (15 min)
● Business offer — Documents & data

Document processing: read, extract, link back to the source

Documents arrive in piles — scans, PDFs, attachments — and the data they contain has to end up in your systems. Your agent reads them, extracts the fields you care about, indexes everything for search, and links each value back to the place in the document it came from. Hosted in France: nothing is uploaded to a foreign service. The officer approves the extractions, and the agent shows them first the ones it is least sure about.

Hosted in France Documents never uploaded GDPR & AI Act: governed deployment Human oversight

Updated on

Deployed in a few weeks
Document processing · hosted in France
Process the batch of documents received this morning.
Batch processed: documents read, type recognised and fields extracted according to your configuration.
Every value is linked to its position in the document — one click and you see where it came from.
Four extractions are flagged as uncertain: poor-quality scan, handwritten note. They are placed at the top for checking.
🔗 Sourced · every value points back to its position
Illegible documents — what do you do with them?
I set them aside without extracting anything, with the reason — missing page, insufficient resolution, truncated document.
Extracting an uncertain value from an illegible document would mean injecting a silent error into your systems.
✎ Action · doubtful documents set aside, never guessed
Local inference · no data outside the EU
Documents hosted in France
Sovereign by designLocal inference or hosting in France
GDPR & AI Act: governed deploymentTraceability & human oversight
TurnkeyDesigned, installed and operated for you
The officer decidesThe agent prepares, never rules
✦ In brief

A Blue Lemon Agent document processing agent reads your scanned documents, recognises their type, extracts the fields you have defined and indexes them for search. Every value stays linked to its position in the original document, and uncertain extractions are flagged first rather than buried in the batch. It runs on local inference or is hosted in France: your documents are never uploaded to a foreign service, architecture designed to reduce exposure to extraterritorial legislation, location alone not being enough to guarantee immunity.

100%
hosted in France in the target architecture
0
transfer outside the EU in the target architecture
6
document uses ready to deploy
0
decision taken without human approval

Reference points describing our offer, not results measured at a client. The scale of the gain, given your volume and the quality of your scans, is confirmed by a pilot.

The context

What does an AI agent bring to your document flows?

The useful data already exists in your documents. Extracting it in a verifiable way, keeping the link to the original, turns a pile into a usable database.

! The issue

A document extraction is only worth something if it is verifiable: a value taken out of a document with no link to its origin is impossible to check, and an error in it spreads silently. The agent links each value to its position in the document, ranks the extractions by its own confidence and sets aside what is illegible — so that human checking goes where it counts.

Our answer

You get usable data without losing traceability: every field points back to the place in the document it came from, which makes verification immediate and correction targeted. Local inference or an isolated resource hosted in France: contracts, invoices and supporting documents never leave your organisation.

The decisive point

The content of your scanned documents: sovereignty & compliance

Your scanned documents often hold your organisation's most sensitive material. Here is how the architecture of our agents protects it.

Local inference

The agent can run on a machine belonging to your organisation: no document leaves the network, no scan is uploaded to a third-party service.

Hosting in France

Otherwise, a dedicated and isolated resource hosted in France, under French law — your documents and the data extracted from them: processing and access within the European Union targeted by the architecture.

Reduced extraterritorial exposure

For the content of your scanned documents, the architecture aims to reduce exposure to the Cloud Act and FISA 702; being located in France or in the European Union does not, on its own, guarantee immunity.

Isolated resource

No pooling whatsoever: an environment strictly dedicated to your organisation and its documents.

Every value linked to its source

Any extracted data points back to its position in the original document; encryption, role-based access (RBAC) and logging of every processing run.

AI Act: governed deployment

An agent strictly in support; no value extracted from an illegible document, no data injected without approval; traceability and human oversight from end to end.

What depends on the architecture chosen These points are not general guarantees: they are settled deployment by deployment, in the quotation.

  • The applicable location is that of the architecture set out in the quotation and verified before commissioning.
  • Local execution is announced only for the configuration explicitly described and accepted in the quotation.
  • The applicable isolation depends on the deployment mode set out in the quotation; no dedicated isolation is presumed.
  • Roles and permissions are configured and accepted for the identities and systems actually connected.
  • The events logged, their content, their retention period and who may access them are defined for the deployment chosen.
For documents covering health data, identity papers or banking details, SecNumCloud and reinforced hosting options are available depending on your requirements. A single architecture is designed to answer both the GDPR and extraterritorial exposure.
Demonstration

See the agent at work

5 real situations, taken from those that come up most often. Pick one: the exchange unfolds as it would in your organisation.

A scripted demonstration. These exchanges show how the agent behaves — its sources, its refusals, what it leaves to your teams. Nothing is sent from this page, no model is queried here, and the matters named are fictional. That is precisely what we promise your data.
The behaviours shown here — monitoring, automation rules, routing and reminders — are configured with you during deployment, from your tools, your rules and your thresholds.
The architecture points named in these exchanges — location, local execution, isolation, encryption, role-based access, logging — are not a guarantee attached to the demonstration: they are those of the architecture set out in your quotation, and verified before commissioning.

The company in this demonstration

Fictional company

Ravel & Ostende — an insurance brokerage handling motor claims under delegated authority for three insurer principals

Sector
Brokerage and delegated motor claims handling (NAF 66.2, insurance auxiliary activities)
Headcount
48 staff, including 11 claims handlers and 2 settlement referents
Market served
26,000 motor policies held by private drivers and tradesmen's fleets, under delegation for three risk carriers
Order of magnitude
210 claim filings per working day; 6 in 10 arrive as a phone photograph of the claim form
Tools in place
A claims@ mailbox, a filing extranet, the firm's own claims management software
Who decides what
The settlement referent sets the liability share and signs the file open; the agent prepares and signs nothing
Room for improvement
38 minutes on average between a form being filed and the file being opened; 1 form in 4 goes back out as a request for a missing tick, signature or page

The firm receives its motor claim forms in the hardest shape there is: a carbon-copy pad filled in with a ballpoint pen at the roadside, two different handwritings, seventeen tick boxes, a freehand sketch — then photographed at an angle with a phone, sometimes on the bonnet. The agent reads the batch, returns every value with its position in the image, and leaves the settlement referent holding the decisions that set the liability share. It is connected to the claims@ mailbox and to the filing extranet.

This company, its figures and the exchanges that follow were invented for the demonstration. They illustrate a common situation; they describe no real client.

Document processing · hosted in France
210 claim forms filed since 6 pm yesterday. They are read. The batch is on your desk at 7:40, every value tied to the pixel it came from.
· The 17 circumstances boxes are captured across all 210, tick by tick, each with its image zone.
· The photographs taken at an angle at the roadside are straightened and read — 126 of the 210 came in by phone, 34 of them past twenty degrees.
· The carbon-copy pads are read through the print-through from the back, and the two handwritings on a single form are told apart.
· 14 duplicates are matched before you open anything at all.
· Setting doubtful documents aside came into play 11 times — 7 pads with no back page, 4 photographs too oblique. They are set aside with their reason rather than half-extracted: a form missing its back page, with its first eight fields pulled out, enters your system looking like a complete file, and nothing afterwards will say the signatures were missing. All 11 go back out as a request the same morning.
And here is what I found reading back over your last six months, which nobody had time to look for: 1 form in 4 goes back out as a request for a missing item — and 3 times in 4 it is the back that is missing, never the front. Your filing screen asks for « your claim form » in the singular and accepts a single file. Two lines to change on the extranet, and those requests stop.
What stays yours, because you wanted it that way: no file is opened without your move. 179 forms are ready to open in one grouped click; 31 wait for you one by one, because they touch the box that sets the scale. morning-batch_210-read.pdf210 forms read by 7:40 · 34 photographs straightened · the filing screen to fix
⛓ Sourced · fictional batch of 210 forms · 126 phone photographs, 34 past 20°
What recognition rate do you claim?
I give you better: the exact count on YOUR forms, redone every morning — and the protocol for checking it yourself.
Why better, and not less: a recognition rate describes a set of documents, not a piece of software. The same engine, on your forms photographed at the roadside and on forms run through a flatbed scanner, does not return the same result — and that gap is wider than the gap between two engines. A brochure figure would tell you something about somebody else's documents.
What I give you instead, this morning, on yours: 210 received, 179 read in full and ready to open, 31 placed in front of your referent with the zone enlarged. That count is redone for every batch and compares with yesterday's: yesterday 168 of 204, the day before 171 of 209. The curve has been rising since Monday filings started going through the extranet rather than the mailbox.
And the figure that actually decides, because it is the one your board will look at: 38 minutes on average between filing and opening, today. On this batch: 179 files openable at 7:40the night's batch ready before the first coffee — with the other 31 settled during the morning.
The protocol I propose, and it takes a day: give me 500 forms you have already handled by hand. I read them and publish the gap field by field against your own keying — registration, date, time, the 17 boxes. You get a figure that is about your documents and that nobody can argue with, and I get the per-field threshold setting that goes with it.
The next step I would recommend straight after: start with the 500 forms from your two largest introducers. They account for 61 % of your filings, and that is where the tuning pays back fastest. the-count-on-your-forms_and-how-to-check-it.pdfThe batch count, day by day · the measurement protocol on 500 of your forms
⛓ Sourced · 179/210 this morning, 168/204 yesterday, 171/209 the day before · protocol on 500 forms
Local inference · no data outside the EU

Your case is not here? That is exactly what a 15-minute conversation is for. Book the free audit

Use cases

What does the agent actually do?

One agent, several steps in the document chain. All these uses work in support, subject to your approval.

Included in your agent The 4 capabilities essential to this promise are included, at no extra cost.
From 700 € excl. VAT / month

Reading and type recognition

Identifies what the document is and applies the matching extraction configuration.

Field extraction

Picks up the fields you have defined, each linked to its position in the document.

Setting doubtful documents aside

Puts aside anything illegible or truncated, with the reason, without extracting anything.

Indexing for search

Makes the collection searchable, without moving or transforming the originals.

Controls and safeguards These 5 controls are built into the agent: they frame what it does, whatever plan you pick. They are not chosen and are not added to your order.
Human validation, exceptions and escalation Status, safe closure and audit trail Sources, access rights and handling of questions with no answer Inherited access rights, versions and freshness of the corpus Explicit handling of questions with no answer and of conflicting sources
What the agent must be connected to This connection is required for the agent to work. It concerns your information system and is scoped during the audit.
Measurement of accuracy, coverage and time saved

Need to go further?

These agents handle a different business process, with their own owner and their own price. They are added to this one.

Does your need fall outside this?

In 15 minutes we identify the most relevant agent — without oversizing the project.

Book the free audit Build your agent
The gain

How much checking can a team focus where it matters?

By extracting and ranking by confidence level, reviewing a whole batch is replaced by checking the flagged cases. The scale of the gain depends on your volume and remains to be confirmed by a pilot.

Extracting the data from a batch
Today · done by hand
Fields extracted, to approve
Checking every value
Today · done by hand
Targeted checking on the flagged cases
Searching the document collection
Today · done by hand
Collection indexed
Illustrative, non-contractual reference points, to be confirmed by a pilot on your volume and the quality of your scans. Extractions are approved by the officer, starting with those the agent flags as uncertain. An illegible document is set aside, never interpreted.
How it works

The stages of your AI agent project

1

Audit & scoping

15 minutes to target the use case with the best return.

2

Quote or direct sign-up

A catalogue offer is bought online; a specific need gets a costed quote.

3

Design

We design the agent and its guardrails.

4

Integration & testing

We connect your tools to the agent, which is itself hosted in France.

5

Rollout

Going live and training your team.

6

Operation

Continuous supervision and improvement.

Pricing

One package, one agent

A document processing agent (reading, extraction, indexing), installed and operated for you. Prices exclude VAT — annual subscription, the time it takes for the gains to settle in.

Agility

Setup + controlled subscription

7,400 € excl. VAT setup
then 700 € excl. VAT/month — you invest at installation and pay a reduced subscription. Ideal for keeping the cost under control over time.
  • Installation, configuration and training for your teams
  • Operation, human oversight, updates and support
  • Sovereign hosting in France, a dedicated and isolated resource
Order →
The simplest Serenity

All inclusive, no setup fee

1,110 € excl. VAT /month
all inclusive, immediate start. No upfront investment: a single subscription. Ideal for starting quickly and simply.
  • Setup included (installation, configuration, training)
  • Operation, human oversight, updates and support
  • Sovereign hosting in France, managed end to end
Order →
100% Sovereign

On site, you own it

10,895 € excl. VAT setup
then 896 € excl. VAT/month · + hardware from 2,491 € (one-off purchase, in addition) — a sovereign computer installed on your premises, maintained remotely. Models run locally, your data returned at the end of the contract. 36-month commitment.
  • Hardware installed on your premises (you own it)
  • French / European AI models run locally
  • Secure remote maintenance (Pro support included)
Order →
Not included in the packages: AI consumption (model tokens), re-invoiced at real cost with no margin, and tracked in real time in your client area. Maintenance and supervision subscription for an initial term of 12 months for the Agility package, 24 months for the Serenity package and 36 months for the 100% Sovereign package, renewable; support levels (SLA 72 h / 24 h / 4 h) optional. Bespoke development, additional integrations or exceptional volumes are quoted separately. Support Monday to Friday, 9am to 6pm. Prices exclude VAT.
AI model: none of the AI models offered currently carries a fixed surcharge. When the selected model carries a cost, that cost is shown when you choose it, before you order, and re-invoiced at the cost incurred, with no mark-up; usage is billed at the publisher's price. Publishers' prices are published in US dollars: the amount re-invoiced is the amount in euros actually borne by Blue Lemon Agent on the publisher's invoice, at that invoice's exchange rate, with no commission or mark-up.
Included components and additional components Components included in the base offer: the Blue Lemon Agent software foundation, the AI models listed in the order journey, the standard channels (Microsoft Teams, Slack, WhatsApp Business, email, website chat, calendars, Microsoft 365 / Google Workspace, file storage, market VoIP telephony, professional social-media pages and accounts, Google Business Profile), hosting in France for the package chosen, backups, supervision, updates and support. If adapting the AI agent to your constraints, your needs or your requests requires other paid components — a third-party publisher's software licence, paid API access to one of your applications, hosting of health data, for which French law requires an HDS-certified host (art. L. 1111-8 of the French Public Health Code), SecNumCloud-qualified hosting, a speech synthesis service, particular hardware —, they are offered to you as an option or on quotation and re-invoiced at the cost incurred; nothing is committed without your written agreement. Where the artificial intelligence model you choose entails an additional cost, that cost is shown to you before you order and re-invoiced to you at the cost incurred, with no margin.
What to expect
Go-live 2 to 3 weeks
Agent designed, channels connected, team trained.
Steady state 4 to 7 weeks
After a few weeks of real use, once the agent's behaviour matches what you expect. Indicative estimate, adjusted to the options you keep. It is not a delivery commitment.
Our commitment

Four guarantees that matter to your documents

Your documents are uploaded nowhereLocal inference or an isolated resource hosted in France; no scan or supporting document entrusted to a third-party service, and no document used to train a model.
Data in France, under French lawThe content of your scanned documents: minimisation and location in France, architecture designed to reduce exposure to extraterritorial legislation, location alone not being enough to guarantee immunity.
The officer keeps the decisionThe agent produces extracted, verifiable data, editable and checkable; no approval is automated.
Human oversight & traceabilityOn your volume and the quality of your scans: systematic logging and monitoring, compliant with the AI Act.
Frequently asked questions

Your questions, our answers

Which documents are supported?
The ones we have seen. The scope is fixed at the audit, document type by document type: we read a real sample of your archive, we show you what comes out of it, and the agreed scope goes into the contract. We do not publish a list of « supported » formats or languages: how hard a document is to read depends on how it was produced and scanned — a carbon-copy pad photographed at an angle and the same form run through a flatbed scanner do not read alike. A list would be reassuring and would say nothing about your documents.
How do you measure confidence, and what is that figure worth?
The confidence level measures the sharpness of the character — contrast, complete shape, absence of noise. It does not measure whether the character recognised is the one that was written: a perfectly crisp « 0 » in a typeface where O and zero look alike scores high confidence and can be wrong. That is why the fields exposed to substitution — references, identifiers, amounts — are checked against their expected structure and against your own data, whatever the confidence. We claim no recognition rate: a rate describes a set of documents, not a piece of software. What we provide is the exact count on your batches, redone every time, and a measurement protocol on a sample of documents you have already handled by hand.
What happens to a critical field?
It carries its own threshold, stricter than the document's, and you set it. A critical field — the one that triggers a payment, a decision or a write — is never set alone until you decide otherwise: the agent shows the value read, the enlarged image zone it came from, and, where two readings are possible, both side by side. You settle it with one click. The dial moves in one line of configuration, and the agent gives you the count for both settings before you choose.
Is handwriting understood?
It is read, and what is promised is not a coverage of handwritings — there is no such thing: two people in the same batch write, one neatly and the other not, and no catalogue says which. What is promised is the marked blank: where the reading does not settle, the agent leaves a visible blank and shows you the enlarged zone, rather than filling the gap from context or from history. A guessed value is indistinguishable from a read one, and it is the one move in this trade that produces a false and perfectly credible piece of data.
How do you avoid duplicates?
Through two distinct routes, which are not equivalent. Two filings of the same file carry the same fingerprint: the second is held, unambiguously. Two different scans of the same document have different fingerprints: the agent matches them on what is written on them and presents you with a match to confirm, never a merge made on its own authority. Merging wrongly would make a document disappear, and a disappearance shows up in no count.
Does the agent export automatically into our systems?
It prepares the export and waits for your move; automation is a mandate you grant — capped, traced and revocable with a word. On a direct connection to third-party software we promise nothing that is not proved: a version of the connector, a test suite and an acknowledgement that has been reviewed. Until those three exist for your tool, it is a separately priced project, not an included feature — and we would rather say so before signature.
Does a document scanned by the agent have the value of the original?
No, and that answer does not depend on us. A digital copy is presumed reliable — and therefore usable as the original would be — only where it results from a process meeting the conditions of décret n° 2016-1673 of 5 December 2016, made for the application of article 1379 of the French civil code (published in the Journal officiel on 6 December 2016). What the agent produces is a data extraction, not a reliable copy within the meaning of that text: it confers no evidential value on your scanning, and keeping your originals remains your decision. An evidential arrangement is a separate project, contracted separately.
How does this agent differ from the knowledge base and the supplier invoice agent?
By what it owns. This agent owns the image: it reads a document, extracts fields from it and ties every value to the pixel it came from; it indexes what it has read so the document can be found again. The knowledge base owns the corpus: it answers a plain-language question from documents that are already usable, and this agent can feed it after approval without replacing it. The supplier invoice agent owns the accounting entry: it matches document, order and receipt, handles the discrepancies and prepares payment. An invoice read here stops at the approved extraction; three-way matching and the payment circuit belong to that agent, and are ordered under its own identifier.
Are our documents protected?
Yes. The agent is hosted in France, running locally or on an isolated resource, with processing and access operated within the European Union as the deployment objective and an architecture designed to reduce exposure to extraterritorial legislation, location alone not guaranteeing immunity. Every read is bounded by the scope you have defined: data outside that scope enters no index, no log and no error message. Your documents are not used to train a third-party model.
How long does it take to deploy this agent?
A few weeks as a rule, depending on the document types to be covered and the systems to be fed, after a free audit and then a design, integration and testing phase.
Let's talk

Let us estimate the potential across your documents

15 minutes to assess your document types — hosted in France, supervised, with no commitment.