Data matching: the probable duplicates, the criteria visible
Two business databases often describe the same members of the public with different spellings. Your agent matches those records, presents the probable correspondences with the criteria that ground them and leaves the merge to your decision. Hosted in France: the public's data stays within your administration. No merge is applied automatically — a confused identity has lasting consequences.
Updated on
The correspondences are ranked by how far the criteria agree.
No merge is applied: every pair is presented for approval.
🔗 Sourced · records from both databases, criteria shown
The merge itself remains your act: confusing two members of the public's identities has lasting effects on their rights.
✎ Support · batch prepared, merge decided by you
A Blue Lemon Agent matching agent cross-checks your business databases and presents the probable correspondences with the criteria that ground them — name, date of birth, address, reference — ranked by how far they agree. No merge is applied automatically. It runs on local inference or is hosted in France: the public's data stays with you, architecture designed to reduce exposure to extraterritorial legislation, location alone not being enough to guarantee immunity.
These figures describe our offer, not results measured at a client. How large the gain is on your number of databases matched and their size is confirmed by a pilot.
What does an AI agent bring to the quality of your data?
A duplicate spotted and documented is dealt with in moments; unspotted, it spreads through every process.
! The issue
Matching two databases calls for comparing records written differently and explaining why they correspond. The comparison is systematic; the decision to merge engages a member of the public's rights. The agent takes on the first and documents every correspondence by its agreeing criteria.
✓ Our answer
Your data department approves correspondences that are already documented, ranked by how far they agree, and concentrates its attention on the ambiguous cases. No merge is applied without a decision: confusing two identities has lasting effects on people's rights. Local inference or an isolated resource hosted in France: members of the public's personal data does not leave the administration.
Members of the public's personal data: sovereignty & compliance
Matching databases by its nature handles personal data drawn from several processing operations. Here is how it is framed.
Local inference
The agent can run on a machine belonging to your organisation: no data about a member of the public and no record leaves the network.
Hosting in France
Otherwise, a dedicated and isolated resource hosted in France, under French law — your business databases and your reference data: processing and access within the European Union targeted by the architecture.
Reduced extraterritorial exposure
For members of the public's personal data, the architecture aims to reduce exposure to the Cloud Act and FISA 702; being located in France or in the European Union does not, on its own, guarantee immunity.
Isolated resource
No pooling: an environment strictly dedicated to your administration and its reference data.
Matching criteria shown
Every correspondence states the agreeing criteria and how far they agree; encryption, role-based access and logging of the matches proposed.
AI Act: governed deployment
The agent is strictly in support; no merge is applied and no identity is changed automatically; traceability and human oversight from end to end.
What depends on the architecture chosen These points are not general guarantees: they are settled deployment by deployment, in the quotation.
- The applicable location is that of the architecture set out in the quotation and verified before commissioning.
- Local execution is announced only for the configuration explicitly described and accepted in the quotation.
- The applicable isolation depends on the deployment mode set out in the quotation; no dedicated isolation is presumed.
- Roles and permissions are configured and accepted for the identities and systems actually connected.
- The events logged, their content, their retention period and who may access them are defined for the deployment chosen.
See the agent at work
4 real situations, taken from those that come up most often. Pick one: the exchange unfolds as it would in your organisation.
A scripted demonstration. These exchanges show how the agent behaves — its sources, its refusals, what it leaves to your teams. Nothing is sent from this page, no model is queried here, and the matters named are fictional. That is precisely what we promise your data.
The behaviours shown here — monitoring, automation rules, routing and reminders — are configured with you during deployment, from your tools, your rules and your thresholds.
The architecture points named in these exchanges — location, local execution, isolation, encryption, role-based access, logging — are not a guarantee attached to the demonstration: they are those of the architecture set out in your quotation, and verified before commissioning.
· One item of equipment is listed in two files with two purchase values. The gap is €12,400, and the accounts carry the lower one.
· Two records share an address, a date of birth and two spellings of the same surname. They will not be merged — I explain why below.
· A grants file and a contracts file disagree on 9 amounts. Both were updated, on two different dates.
· 412 rows in one file have no counterpart in the other. That is not an anomaly: the two files do not cover the same thing, and nobody had written that down. morning-watch_4-flags.pdf4 flags · 1 scope to write down
⛓ Source · asset register, accounts file, grants file, contracts
What agrees: same address, same date of birth, two close spellings of the same surname. Three strong signals.
What does not agree: two different first names, and two files opened four years apart on two unrelated schemes.
The most likely explanation, and it is ordinary: two people from the same family at the same address, born on the same day of different years — or the same day, which happens. A merge here would mix two lives into one record.
What that would cost: entitlements computed on two people's income, a medical or social history attributed to the wrong one, and a record nobody can separate again — because nothing says any more which line belonged to whom.
What I do instead: I show them side by side, with what agrees and what diverges, and an officer decides in thirty seconds with an identity document from the file. 2-records_what-agrees-and-diverges.pdf3 agreements · 2 divergences · no merge
⛓ Source · 2 records, 3 agreeing signals, 2 diverging signals
Routing follows the nature of the gap: a value gap goes to the department owning the reference file; two close records to an officer who can open an identity document, never to a dashboard; amount divergences to both departments, with both update dates; the 412 rows to the head of service, once.
With a chase: 7 days on everything, except what feeds a financial statement under way — 24 h in that case. Then a monthly summary: by type of gap and by pair of files, never by resident.
What this comparison has already produced: a €12,400 gap on one and the same asset raised with the department that owns the reference file before the accounts are closed, nine divergent amounts between grants and agreements presented with both their update dates, 412 rows explained — the two files do not cover the same scope, and that is now written down —, and two close records set side by side with what matches and what does not.
What that gives the service from tomorrow: financial statements that are true, and an officer who decides in thirty seconds on a pair of records instead of reworking two databases by hand. Merging stays their decision, and that is what protects the resident: two lives mixed into one file cannot be pulled apart again, because nothing says any longer which row belonged to whom. I hand that decision back fully worked — matching criteria, diverging criteria, level of match, and the identity document open at the right place. The day you want me to apply the merges myself above a level you set, you give me the mandate: written, bounded to that level, dated, withdrawn on a word. Residents' data never leaves the administration: I compare what you open to me, database by database, and every comparison is logged.
The next step is ready: the pairs are grouped by level of match, the safest first, so validation can move in batches. Tell me which level you want to handle first.
✎ Proposal · watch and chases to be configured — you set the thresholds
What carries weight: a shared identifier — the safest signal, and the only one sufficient on its own; a combination of fields improbable by chance; a cross-reference, where one row cites the other.
What carries almost none: a name alone, an address alone, a date alone. Two people can share any one of those fields without being the same person.
What I produce for each match: the fields that agree, the fields that diverge — given with equal care — and what it would take to decide.
The level of agreement I produce for every pair, and it is named rather than scored: "shared identifier", "improbable combination", "cross-reference", "weak signal alone" — and the matches arrive grouped by level, ready to approve in batches. Why a name and not a percentage: "87%" does not say what is missing, and an officer who sees 87% approves; an officer who reads "two different first names" opens the file. The "shared identifier" batch is approved in one action; the pairs with a divergence are looked at one by one.
Across 18,400 rows compared: 1,240 matches proposed, 96 of them with at least one divergence — those 96 are what matters. matching_18400-rows.pdf1,240 matches · 96 with a divergence
⛓ Source · 18,400 rows, 2 files, the department's matching rules
What they represent: 96 cases across 18,400 rows, half of one per cent. At two minutes each, that is three hours — once, and the backlog is cleared.
What I did for those three hours: each case arrives with the document from the file that would settle it, where one exists. On 74 of the 96 it already exists: an identity document, a tax notice, a deed. The officer opens it, decides, moves on.
On the other 22 the document is not in the file, and I say so rather than proposing anyway. Those 22 need a call or a letter.
What matters in that figure: it is not that the tool finds 1,240 matches — it is that it brings the human decision down from 18,400 rows to 96 cases, 74 of which settle with a document already there. 96-cases_74-with-the-document.pdf96 cases · 74 settleable · 22 to document
✎ Support · 96 cases prepared — 74 with the document that settles it
The two values: the asset register shows €34,200, the accounts file €21,800. A gap of €12,400.
What I looked for: a purchase document attached to either. The invoice exists, it is attached to the accounting line, and it shows €34,200.
What that means: the correct value is in the asset register, and the accounting line carries the wrong one despite holding the invoice. That is the opposite of what one would expect, and precisely why the document had to be opened rather than assumed.
What gets signed, and why: correcting the accounting line. A corrected asset value changes a depreciation and a balance sheet — it belongs to the accountant, with the invoice in front of them, and it carries their name. The work itself is done: the gap is established on documentary evidence, and I hand them the depreciation recalculated over the remaining life, next to the old one.
What I prepared: both lines, the invoice, and the 6 other items in the same state that I found while checking this one. 7-items_value-gap.pdf7 items · document opened for each
⛓ Source · asset register, accounts file, attached purchase invoice
What I looked at: the last update date of each row. The 9 divergences correspond to 9 variations signed between March and June. The contracts file took them in; the grants file stayed on the original amounts.
So it is not a keying error: it is one file that follows variations and another that never receives them. The problem is the circuit, not the data.
What that produces today: statements of grants paid are drawn up on pre-variation amounts. The cumulative gap is €47,600, running both ways.
What I propose: flagging each signed variation to the department holding the second file — and not copying the amounts across myself. A file fed by a tool without anyone knowing where the value came from is a file nobody checks any more. 9-variations_47600-eur.pdf9 variations · 2 files · cumulative gap
✎ Proposal · flagged when a variation is signed — no automatic copying
What I found on looking: the 412 are not scattered. 390 share one characteristic: they concern beneficiaries attached to a municipality that joined the joint authority in 2024. The second file was never extended to that perimeter.
What that means: the two files do not describe the same territory, and no document says so. Every cross-file statement produced for two years therefore understates part of the territory without anyone knowing.
The remaining 22 are isolated cases: 14 recent creations not yet carried over, 8 rows closed on one side and not the other.
What I propose: no data migration — first write down what each file covers. A one-page note, two perimeters, two dates. Migrating the data before writing that would amount to mixing two territories without knowing it. 412-rows_2-perimeters.pdf390 one cause · 22 isolated cases
⛓ Source · 412 rows, municipal attachments, file creation dates
What is kept: the match proposed, the agreeing and diverging fields, the decision taken, and its date.
What is not kept: no match between files you have not explicitly paired, no profile built by crossing several sources, no data copied from one file into another.
Why this is written here, and it is the most important point on this page: a tool that matches files is, technically, one step from a tool that cross-references all of them. That step is not taken, and it is not taken by design: every pair of files is paired explicitly by you, and two unpaired files never see each other.
What that gives you concretely: you can answer a resident asking what is done with their data — the answer is a list of file pairs, written down, short, and true. what-is-kept.pdf4 items kept · 3 prevented by design
✎ Framework · file pairings to be declared — nothing is matched by default
Your case is not here? That is exactly what a 15-minute conversation is for. Book the free audit →
What does the agent actually do?
One agent, several pieces of data quality work. All these uses work in support, subject to your approval.
Matching across databases
Cross-checks the records in your business databases and reference data.
Documented criteria
Shows the agreeing criteria and how far each pair agrees.
Batches prepared to approve
Groups the correspondences by level to speed up approval.
In 15 minutes we identify the agent that will give your staff the most time back — without oversizing the project.
How many records can a department match?
By taking on the comparison and the documenting, the effort shifts towards deciding the ambiguous cases. How large the gain is depends on your volume and remains to be confirmed by a pilot.
The stages of your AI agent project
Audit & scoping
15 minutes to target the use case with the best return.
Quote or direct sign-up
A catalogue offer is bought online; a specific need gets a costed quote.
Design
We design the agent and its guardrails.
Integration & testing
We connect your tools to the agent, which is itself hosted in France.
Rollout
Going live and training your team.
Operation
Continuous supervision and improvement.
One package, one agent
A data matching agent (cross-checking, criteria, batches to approve), installed and operated for you.
Setup + controlled subscription
- Installation, configuration and training for your teams
- Operation, human oversight, updates and support
- Sovereign hosting in France, a dedicated and isolated resource
All inclusive, no setup fee
- Setup included (installation, configuration, training)
- Operation, human oversight, updates and support
- Sovereign hosting in France, managed end to end
On site, you own it
- Hardware installed on your premises (you own it)
- French / European AI models run locally
- Secure remote maintenance (Pro support included)
Four guarantees that matter to your reference data
Related resources
Your questions, our answers
Does the agent merge the duplicates?
What criteria does a match rest on?
What does it do with ambiguous correspondences?
How is the GDPR respected?
Is the public's data protected?
How long does it take to deploy this agent?
Other agents for your data
Let's size up the potential in your reference data
15 minutes to frame your databases and your criteria — hosted in France, supervised, with no commitment.