AI agent for anomaly and fraud detection: spot the weak signals, keep the human decision
Manual checks on growing volumes of data — grants, spending, declarations, case files — let inconsistencies through and take up a considerable amount of time. Your AI agent spots the anomalies and weak signals, then ranks them by risk level. Hosted in France — on local inference or an isolated resource — public data never leaves your perimeter. The AI agent flags; the public officer qualifies, checks and decides.
Updated on
Every alert is documented and recorded — for the case officer to qualify.
⛓ Source · your business system + the files' documents
I am preparing a documented alert record, for your approval.
✎ Action · alert record ready for review — the public officer decides
In an administration, a Blue Lemon Agent agent spots anomalies and inconsistencies in the data — grants, spending, declarations, case files — then ranks the alerts by risk level and documents each one to make human review easier. It runs on local inference or is hosted in France: public data is never exposed to a foreign service, architecture designed to reduce exposure to extraterritorial legislation, location alone not being enough to guarantee immunity. The AI agent does not penalise: it flags; qualification and decision remain strictly human. A high-risk use within the meaning of the AI Act: reinforced human oversight and traceability.
Reference points describing our offer, not results measured at a client. The scale of the gain is confirmed by a pilot on your own scope.
Why anomaly detection matters to administrations — and why they hesitate
The volumes of files and data are rising, but the headcount available for checks stays constrained. Targeting the verifications becomes essential — without entrusting sensitive public data to a foreign service, and without letting a machine decide on a penalty.
! The issue
The administration has to guarantee the proper use of public funds and equal treatment, while facing volumes that manual checking cannot fully cover. Yet most consumer AI solutions amount to entrusting users' personal data, benefit amounts, declarations and supporting documents to a third party, often hosted outside Europe and subject to the Cloud Act.
✓ Our answer
AI is only of interest to an administration if it is sovereign and confidential by design. Local inference or an isolated resource hosted in France, transparent and revisable detection criteria, systematic human oversight: the AI agent ranks the alerts, but qualification and decision stay reserved to the public officer. The aim is not to penalise automatically, but to concentrate checking time where the risk is highest.
Protecting public data: sovereignty & compliance
An administration handles users' most sensitive data. Here is how the architecture of our agents protects it, file by file.
Local inference
The agent can run on a machine belonging to the authority: no data leaves the network, nothing passes through a cloud.
Hosting in France
Otherwise, a dedicated and isolated resource, hosted in France under French law — your data: processing and access within the European Union targeted by the architecture.
Reduced extraterritorial exposure
As regards public data, exposure to the Cloud Act and FISA 702 is reduced by design; location alone does not guarantee immunity.
One isolated resource per administration
No pooling of public data: an environment strictly dedicated to your organisation.
Encryption & controlled access
Encryption in transit and at rest, role-based access (RBAC), strong authentication and logging.
AI Act: governed deployment
A high-risk use: the agent only flags; no penalty is automated; traceability and human oversight from end to end.
What depends on the architecture chosen These points are not general guarantees: they are settled deployment by deployment, in the quotation.
- The applicable location is that of the architecture set out in the quotation and verified before commissioning.
- Local execution is announced only for the configuration explicitly described and accepted in the quotation.
- The applicable isolation depends on the deployment mode set out in the quotation; no dedicated isolation is presumed.
- The encryption mechanisms in transit and at rest, their components and key management are those documented for the architecture chosen.
- Roles and permissions are configured and accepted for the identities and systems actually connected.
- The events logged, their content, their retention period and who may access them are defined for the deployment chosen.
See the agent at work
4 real situations, taken from those that come up most often. Pick one: the exchange unfolds as it would in your organisation.
A scripted demonstration. These exchanges show how the agent behaves — its sources, its refusals, what it leaves to your teams. Nothing is sent from this page, no model is queried here, and the matters named are fictional. That is precisely what we promise your data.
The behaviours shown here — monitoring, automation rules, routing and reminders — are configured with you during deployment, from your tools, your rules and your thresholds.
The architecture points named in these exchanges — location, local execution, isolation, encryption, role-based access, logging — are not a guarantee attached to the demonstration: they are those of the architecture set out in your quotation, and verified before commissioning.
· One criterion is producing four times as many flags as in April, although file volumes have not moved. 47 flags this month against 11.
· 31 flags from this quarter have never been dealt with and remain displayed on the files concerned.
· Three officers have hidden the same criterion from their view. None explained why, and nothing asked them to.
· One file shows a date inconsistency that no existing criterion covers: a document dated after the decision it supports. morning-watch_4-flags.pdf4 flags · 3 about the system
⛓ Source · flag log, criteria settings, display preferences, files
What the criterion does: it flags files whose proof of address is dated more than three months before the filing date.
What changed: since the update of 12/06, your system records that date as day/month instead of month/day for documents uploaded online. A proof dated 03/07 is read as 7 March.
What that produces: 36 extra flags in two months, all false positives, all on files submitted online.
What it costs beyond wasted time: thirty-six residents were looked at more closely for no reason at all. That is the real damage, and it does not show in the system's statistics.
What I have already prepared: the fix to the criterion is written — a date reading that accepts both formats — and I have run it back over the two months elapsed: it produces none of the 36 flags, and removes none of the others. The batch of 36 files to clear is assembled, with the shared cause demonstrated. Bringing it into force is signed by the scheme owner — that is what makes it dated, traced and contestable. Approve it, and the criterion is fixed tonight; until then I count every day the residents the format adds. criterion_36-false-positives.pdf47 flags · 36 caused by a date format
⛓ Source · criterion settings, date format before and after 12/06, 47 flags
Routing follows your organisation: the runaway criterion to the scheme owner, with the 36 files to clear; the 31 untreated flags to their case officers; the hidden criterion to the owner, without naming the officers — three people hiding the same criterion say something about the criterion, not about themselves; the date inconsistency to the head of service, as a proposed new criterion.
With a chase: 24 h on the runaway criterion — every day adds residents looked at for no reason; 7 days on the rest. Then a monthly summary to the owner: by criterion, never by officer or by resident.
Three things that hold, and they are strengths. My access is yours: opened role by role, logged line by line, withdrawn on a word — nothing leaves your walls and you know at any moment what I read. The decision stays with the service: no file suspended, no application refused, no score on anyone — and I hand you each call in minutes, with the full case, quantified and reasoned. What the checks catch, I tell you: all 4,200 files of the year pass before the 14 criteria on the day they are filed, where a sample check looks at a fraction only — and every month I bring you, with figures, the criteria to narrow and the ones that are missing.
✎ Proposal · watch and chases to be configured — you set the thresholds
Example, the criterion that ran away: "Flag files whose proof of address is dated more than three months before the filing date."
What the officer sees on the file: that sentence, the date read, the filing date, and the gap calculated. Nothing else — no score, no level, no colour.
Why no score: a number from 0 to 100 on a social support file reads as a judgement about the person, and it outlives the explanation. A sentence can be argued with; a score is simply borne.
All 14 of your current criteria are written that way, and for each I can tell you how many flags it produced and how many were confirmed after examination.
What I also do, and what no one has ever had time for: I re-read your 14 criteria in the light of the cases actually examined, and I propose new ones, written in the same form and already quantified. For each I give you the sentence in plain language, the number of flags it would have produced over the last twelve months, the share confirmed after examination, and the cases it would have let through — you choose on figures, not on a hunch. The signature stays with the service: a criterion only takes effect once approved, and that is exactly what makes it contestable before a member of the public. I save you the writing and the measuring; the decision is yours, and it takes ten minutes instead of a committee. 14-criteria_confirmation-rates.pdf14 criteria · confirmed / cleared
⛓ Source · 14 criteria in force, log of flags and their outcomes
· Criterion 9 — 84 flags in a year, 3 confirmed. A 3.6% confirmation rate. It has 81 residents a year looked at closely for nothing.
· Criterion 12 — 0 flags in a year. It costs nothing and serves nothing; it was written for a scheme closed in 2024.
The first is the real subject. A criterion at 3.6% confirmation is not a safety net: it is a suspicion generator, and its cost is borne by the residents wrongly flagged.
What I propose, with the figures: I switch nothing off on my own — a criterion can sit at 3.6% confirmation and remain justified if what it catches is serious, and that judgement is yours. But I do not leave you with the finding alone: I have written the narrowed version of criterion 9 and run it over the last twelve months. It would have produced 19 flags instead of 84, kept all 3 confirmed cases, and spared 65 of the 81 people examined for nothing. It is ready, phrased in a single line like the others. Approve it and it is live tonight; leave criterion 9 as it stands and I keep measuring it.
What I have prepared: for criterion 9, the 3 confirmed cases and the 81 cleared ones, with what separates them. Two of the three confirmed share a characteristic the criterion does not test — there may be a better criterion to write. criterion-9_84-flags.pdf84 flags · 3 confirmed · what separates them
✎ Support · 2 criteria to review — keeping them is the service's decision
What the business data analysis cross-references: your 184,000 payment orders from the last three financial years, the supplier file, the notified contracts, the grant award decisions and the council calendar. Crossing those five sources reveals what none shows alone: an atypical trend — chapter 011 payment orders issued between 24 and 31 December account for 19 % of annual volume against 8 % expected pro rata — and a weak signal: 7 suppliers share the same bank account number, which has a mundane explanation in 5 cases out of 7 and warrants a question in the other 2.
Public expenditure control, item by item: duplicate payments — same supplier, same amount, same invoice reference within 60 days: 41 cases, 29 confirmed on examination, €61,400; atypical amounts — a payment order more than three standard deviations from that supplier's own series: 112 cases; consistency gaps — a payment order with no prior commitment, a document dated after the decision it supports, an amount above the notified contract: 88 cases.
What I put on the table, already costed: every criterion arrives with its sentence in plain language, the number of flags it would have produced over twelve months, the share confirmed on examination and the files it would have let through — you choose on figures, not on intuition. No flag characterises anything: an anomaly is a discrepancy to be checked; the legal characterisation belongs to your department and, where appropriate, to the competent authority. That is what makes a file solid rather than open to challenge.
The figure that does not flatter me: the atypical-amount criterion went from 11 flags in April to 47 this month while the volume of payment orders did not move. The business data analysis says why: the school catering contract's price revision on 1 September shifted a whole series of amounts, and my standard deviation was computed over a rolling twelve months without accounting for contract amendments. I rewrote the criterion to align each series with the effective date of every amendment and ran it over 24 months: 14 flags instead of 47, the 3 confirmed ones kept, 33 files spared a pointless look. Approve it and it is in force tonight.
⛓ Sourced · 184,000 payment orders across 5 sources, 41 duplicates of which 29 confirmed, 47 flags down to 14
· The criterion's sentence, in plain words.
· The values read, with the document and page they came from.
· The calculation, where there is one — here, the gap between two dates.
· The date of the flag and the criterion that produced it.
What they do not see: any risk level, any colour, any ranking of this file against others. An officer who sees "high risk" assesses differently, and they should not assess differently before having looked.
What they can do in one click: open the two cited documents side by side, and mark the flag as cleared — at which point it disappears, leaving no trace on the file.
What I hold to, and what costs the most technically: a cleared flag disappears everywhere — file, index, exports, logs — and the purge is dated and replayable before an auditor. It is the most important rule of the system: a member of the public should not carry a suspicion an officer has already set aside. What I do keep, and what serves you: the count of clearances per criterion, with no file and no person — that counter is what lets me tell you a criterion runs at a 3.6% confirmation rate. flag_what-the-officer-sees.pdf4 items shown · 0 score
⛓ Source · flag, cited documents, display settings
What would happen: a file flagged in error in 2024 and cleared in three minutes would keep a line in its history. In 2027 another officer opens the file and reads "3 flags since 2023". They will not know all three were cleared — and even if they do, they will have seen the number.
What that produces: a file's reputation, built out of corrected errors. The most exposed residents are those whose circumstances are complex, and therefore those who trigger the most false positives.
What is kept, and it is enough: the log by criterion — how many flags produced, how many confirmed, how many cleared. That log cannot be traced back to a file. It serves to improve the criteria, which is its only legitimate use.
What I propose in addition, if you approve: a quarterly review of the criteria on that log, to the scheme owner. By criterion, never by resident, never by officer. log-by-criterion_what-is-kept.pdfWhat is kept, what is erased
✎ Proposal · log by criterion — no history attached to a file
· 19 come from the runaway criterion — the date format. They can be cleared as a block once the cause is confirmed, without opening the files.
· 7 concern files already closed, assessed and decided since. The flag arrived after the decision: it served nothing and it lingers.
· 5 are open and unexamined, two of them for more than 60 days.
What I flag about those five: the two oldest concern support applications whose assessment is suspended. Two residents are waiting for a decision that is waiting for a flag nobody is looking at.
What I have already done for you: the nineteen false positives are gathered into a single batch, with the shared cause demonstrated — an officer clears them in one gesture, all nineteen at once, instead of opening nineteen files. The batch of 7 closed files is ready the same way, with its reasoning. Clearing stays in their hand: it is an assessment act, it is recorded, and that record protects the member of the public as much as the officer.
What it gives: 26 of the 31 can be handled without opening a file. The five that remain are the ones that deserved a look. 31-flags_sorted.pdf31 flags · 26 handled as a block
⛓ Source · 31 flags, file states, assessment dates
What I saw: on one file, a supporting document dated 14/07 sits behind a decision taken on 02/07. Twelve days apart, the wrong way round.
What it can be: a mistyped date, a document supplied afterwards to complete a file already decided, or something else. All three are plausible and I favour none.
What I have done, and what I put on the table: I have written the criterion — "flag files where a supporting document is dated after the decision it supports" — and run it dry across the year's 4,200 files. It would have produced 23 flags, not hundreds: 15 match a document supplied afterwards to complete a file already decided, which is regular and can be tested for, and 8 remain unexplained, this one among them.
What I propose you approve: the narrowed version, which sets aside supplements filed at the counter — 8 flags over the year, fewer than one a month, each documented like this one. The head of service signs and the criterion takes effect; that signature is what makes it contestable before a member of the public. You choose between 23 and 8, on figures, not on a hunch. date-inconsistency_1-case.pdf1 case · dry test possible on 4,200 files
✎ Proposal · the case surfaced, the criterion stays for the service to write
What traceability keeps, flag by flag: the criterion applied and its version, the exact data that triggered it with source and extraction date, the timestamp, the public officer it was assigned to, the decision taken — upheld, dismissed, held — its one-sentence reason and the identity of the person who took it. Across 2,340 flags, 2,340 can be reconstructed line by line: a file can be replayed in ten seconds in front of the person concerned, in front of your management, or in front of a court.
Why human oversight is not a formality here: this processing falls under the strengthened requirements of the European AI Regulation, whose obligations for this kind of use take effect on 2 December 2027. What that requires, and what the setup already holds: a natural person who is competent and trained, who can dismiss a flag with a one-sentence reason, halt the system, and whose decision overrides mine in all circumstances. No individual decision is taken without them, and the log proves it rather than asserting it.
What traceability has already made visible: 3 officers hid the same criterion in their view, with nothing asking them why. The log showed it, the question was put at review, and the answer was a good one: the criterion was flagging a lawful accounting arrangement particular to your annexed budget. The criterion was corrected, not the officers. Human oversight that never reports anything back is oversight nobody listens to.
The figure that does not flatter me: 31 flags from the quarter were never dealt with and stayed displayed on the files concerned — 1.3 %, and that is incomplete traceability: a flag without a decision cannot be reconstructed, it just hangs. I set a 30-day limit beyond which an unsettled flag goes up to the head of department with its age and disappears from the person's file until it has been examined — nobody carries a suspicion no one has looked at. The 31 were cleared at review: 6 upheld, 25 dismissed, each with a one-sentence reason.
⛓ Sourced · 2,340 reconstructible flags, 31 cleared at review, 3 hidden criteria revealed by the log
Your case is not here? That is exactly what a 15-minute conversation is for. Book the free audit →
The facets of anomaly detection in the public sector
Each use corresponds to a facet of the agent we deploy. All of them work in support: the AI agent flags and ranks, the public officer qualifies and decides.
Anomalies on benefit applications
Spotting inconsistencies, duplicates and missing documents on benefit and allowance files, before payment.
Ranking by risk level
Classifying alerts (high risk, medium, to verify) to target checks where they matter most.
Documenting the alerts
Producing, for every alert, a recorded and reasoned file that eases human review and decision-making.
Checking public spending
Detecting duplicate payments, atypical amounts and consistency gaps in spending and payment orders.
Business data analysis
Cross-checking and querying the data to bring out weak signals and atypical trends.
Consistency of assessment files
Checking the completeness and consistency of benefit files before approval by the officer.
Transparent & revisable criteria
Explicit, adjustable and auditable detection rules, to limit false positives and bias.
Traceability & human oversight
Logging of alerts and decisions, compliant with the reinforced requirements of the AI Act.
Need to go further?
These agents handle a different business process, with their own owner and their own price. They are added to this one.
In 15 minutes we identify the agent that will give your staff the most time back — without oversizing the project.
How much time can an administration win back?
By automating the review of the data and the ranking of alerts, an administration can concentrate its checking time on the files that are genuinely at risk — instead of verifying at random or by sampling.
The stages of your AI agent project
Audit & scoping
15 minutes to target the use case with the best return.
Quote or direct sign-up
A catalogue offer is bought online; a specific need gets a costed quote.
Design
We design the agent and its guardrails.
Integration & testing
We connect your tools to the agent, which is itself hosted in France.
Rollout
Going live and training your team.
Operation
Continuous supervision and improvement.
Three options, one agent
An anomaly and fraud detection agent (alerts, ranking, documentation), installed and operated for you. Choose according to how you work. Prices adapted to the public sector — annual subscription, the time it takes for the gains to settle in.
Setup + controlled subscription
- Installation, configuration and training for your teams
- Operation, human oversight, updates and support
- Sovereign hosting in France, a dedicated and isolated resource
All inclusive, no setup fee
- Setup included (installation, configuration, training)
- Operation, human oversight, updates and support
- Sovereign hosting in France, managed end to end
On site, you own it
- Hardware installed on your premises (you own it)
- French / European AI models run locally
- Secure remote maintenance (Pro support included)
Four guarantees that matter to an administration
Your questions, our answers
Does the agent impose penalties?
How are false positives and bias avoided?
What regulatory framework applies?
Is the public's data protected?
Do we have to change our information system?
How does the agent document its alerts?
How long does it take to deploy an agent?
Other uses for your data and your checks
Let us estimate the potential in your administration
A conversation to identify the checking scope with the best return — hosted in France, supervised, the public officer keeps the decision.