The AI agent for code generation & review: write faster, review better
Writing repetitive code, understanding an inherited module, reviewing a pull request, documenting a function: this work ties up your developers without always producing the value expected. Your AI agent absorbs that labour — it proposes code, refactors it and reviews it — while your teams concentrate on architecture and product. Hosted in France, on local inference or an isolated resource: your source code and your intellectual property stay with you. The developer keeps the lead.
Updated on
I can propose the fixes.
⛓ Source · your Git repository + your internal review rules
Patch awaiting your review — nothing is pushed without your approval.
✎ Action · patch proposed locally — the developer approves and merges
For a technical team, a Blue Lemon Agent agent writes, completes, refactors and reviews code — generating functions, reviewing pull requests, detecting vulnerabilities and regressions, documenting. It runs on local inference or is hosted in France: your source code and your intellectual property are never exposed to a foreign service, nor used to train a third-party model, architecture designed to reduce exposure to extraterritorial legislation, location alone not being enough to guarantee immunity. The developer keeps the decision and the merge. Live within a few weeks.
Reference points describing our offer, not results measured at a client. The scale of the gain is confirmed by a pilot on your own scope.
Why AI appeals to technical teams — and why they hesitate
Code assistants save real time, but they often mean sending the repository to a foreign service. Yet source code is the most strategic asset of a tech company: exposing it means exposing its intellectual property.
! The issue
Teams are caught between pressure to ship fast, accumulating technical debt, and code reviews that keep getting longer. Yet most consumer assistants amount to entrusting source code, secrets, business logic and architecture to a third party, often hosted outside Europe, subject to the Cloud Act, and liable to use your repositories to train its models.
✓ Our answer
AI is only of interest to a technical team if it is sovereign and confidential by design. Local inference or an isolated resource hosted in France, code never reused to train a third-party model, systematic human oversight, merging reserved to the developer: the time saved is never paid for in lost intellectual property. The aim is not to replace your developers, but to give them back thinking time for architecture and product.
Confidentiality of source code: sovereignty & compliance
A repository holds all the value of a tech company: algorithms, secrets, business logic. Here is how the architecture of our agents protects it, line by line.
Local inference
The agent can run on your own machines or infrastructure: no line of code leaves the network, nothing passes through a foreign cloud.
Hosting in France
Otherwise, a dedicated and isolated resource, hosted in France under French law — your code: processing and access within the European Union targeted by the architecture.
Reduced extraterritorial exposure
Architecture designed to reduce exposure to extraterritorial legislation, location alone not being enough to guarantee immunity.
Code never reused
Your repository is never used to train a third-party model: your intellectual property stays strictly yours.
One isolated resource per client
No pooling: an environment strictly dedicated to your company and your repositories.
AI Act: governed deployment
An agent strictly in support; no automatic commit or merge; traceability and human oversight from end to end.
What depends on the architecture chosen These points are not general guarantees: they are settled deployment by deployment, in the quotation.
- The applicable location is that of the architecture set out in the quotation and verified before commissioning.
- Local execution is announced only for the configuration explicitly described and accepted in the quotation.
- The applicable isolation depends on the deployment mode set out in the quotation; no dedicated isolation is presumed.
- The events logged, their content, their retention period and who may access them are defined for the deployment chosen.
See the agent at work
4 real situations, taken from those that come up most often. Pick one: the exchange unfolds as it would in your organisation.
A scripted demonstration. These exchanges show how the agent behaves — its sources, its refusals, what it leaves to your teams. Nothing is sent from this page, no model is queried here, and the matters named are fictional. That is precisely what we promise your data.
The behaviours shown here — monitoring, automation rules, routing and reminders — are configured with you during deployment, from your tools, your rules and your thresholds.
The architecture points named in these exchanges — location, local execution, isolation, encryption, role-based access, logging — are not a guarantee attached to the demonstration: they are those of the architecture set out in your quotation, and verified before commissioning.
· An API key was pushed to a branch 40 minutes ago. It is the only case where I block, and I did.
· Seven production incidents in the last six months concern code I had reviewed without saying anything.
· Eighty-nine per cent of my remarks were about style before an automatic formatter was connected. They were masking the rest.
· Forty-one of my remarks were rejected out of 312. I publish that figure. morning-watch_4-flags.pdf7 incidents on code reviewed without a remark
⛓ Source · 312 reviews, incident log, remark history
What I record: over the last six months, seven production incidents trace back to code I had reviewed and flagged nothing on.
What I did before telling you: I went back over all seven, one by one, and sorted them by what I had been missing. Two genuinely were subtle — a race condition, a time-zone edge case. Five were not: an untested value, two query-scope errors, two regressions on existing behaviour. And for those five I wrote the review rule that would have caught them: it has been running for six weeks and has already flagged eleven cases of the same kind, three of them fixed before merge. Explaining them by their difficulty would have been comfortable, and wrong.
Why I publish this: because the usual measure of a review tool is the number of remarks produced, and it improves as the tool grows chattier. What matters is what it lets through, and that figure only shows up months later, in incidents.
What it should change for you: my review does not replace yours. Of the five non-subtle incidents, three merges had been approved in under two minutes — the time to read my opinion, not the code.
That is this trade's real risk: I do not let bugs through, I lower attention. 7-incidents_3-merges-in-2-minutes.pdfI lower attention
⛓ Source · 7 incidents, causes analysed, approval times
Routing follows what can be recovered: a secret in the code blocks the branch and goes immediately to whoever pushed it and to security — the only block; a substantive remark stays on the line concerned, never elsewhere; an incident linked to a silent review to nobody — I record it in my own measure; a recurring rejection reason to nobody either, I fix it.
With a chase: none on a remark. An unaddressed remark stays a remark, not a reminder. Then a monthly summary: by type of remark and rejection reason, never by developer.
What it has already given you: an API key stopped 40 minutes after it was pushed, before it reached production; 89% of style noise moved out to an automatic formatter, so reviews finally bear on substance; and 41 remarks rejected out of 312, published by me, because a reviewer that measures what it lets through is a reviewer you can believe.
What it changes tomorrow: the first pass of a review arrives done — bugs, vulnerabilities, duplication, departures from your conventions, on the line concerned — and your developers open the pull request with time to read the code, not just my opinion. On the five non-subtle incidents, that is exactly what was missing.
Your code does not leave your walls: repository by repository, access by role, logged, withdrawn on a word, local inference or an isolated resource hosted in France, and nothing trains a third-party model. The merge stays with the named developer, and I hand it over in minutes: the diff reviewed, remarks ranked by severity, the missing tests, and the proposed fix ready to apply in one click. Open me a repository, and the first review lands on the next pull request.
✎ Framework · no modification, no approval, no per-developer statistics
What happened this morning: an API key was pushed to a branch. I blocked the branch, told the person and security, and touched nothing.
Why it differs from everything else: a bug is fixed by a patch. A pushed key stays in the history, even after the file is deleted — and it is on every machine that pulled the branch. Time matters: forty minutes is not nothing, and it beats three days.
What I prepared during those forty minutes: the purge command written for this repository specifically, with the list of the 9 known clones to resynchronise and their last access; the revocation request to the vendor, drafted and ready to send; and, on a branch, the key replaced by a reference to your vault, tested.
What is left to decide — and these are decisions: rewriting a shared history breaks everybody else's copies, and revoking a key can stop production. Each carries a name and a time, and that is what means somebody warns the teams first. Tell me the order, and the three gestures follow on.
What I supply so it can be settled: the file, the line, the time of the push, and the list of what must be decided — revoke, reissue, purge the history, inform the vendor.
Over twelve months: 4 secrets detected, 4 branches blocked. Three were test keys — and I do not tell the difference, because a test key in a repository looks exactly like a production key. 4-secrets_4-blocks.pdfA pushed key stays in the history
⛓ Source · 4 secrets over 12 months, block log
What I do: I block and supply everything needed to decide. I never lift my own block, even when told it is a false alarm.
Why that rigidity specifically: an agent that can lift its blocks can lift any of them, and it then suffices to be in a hurry to convince it. Urgency is exactly when people most want to override, and the worst moment to do it.
What lifts the block: a person, recording it. The note says who, when, and why — and it stays in the repository.
What that gave: of the 4 blocks, two were lifted within the hour after checking the key really was a test key and revoked. The notes exist, and that is all I was asking for.
And on test keys, a figure rather than a principle: three of your four detections were test keys. A repository containing test keys teaches its contributors that the repository contains keys — and it is that proportion, three in four, that installs the habit. That mechanism is what produces the real leak, not the key itself. 4-blocks_2-lifted-by-a-person.pdfUrgency is the worst moment to override
✎ Framework · the agent never lifts its own block
What I record: before an automatic formatter was connected, 89% of my remarks concerned spacing, line breaks, import order. All correct, all of no interest.
What that produced: a review with 40 remarks, 36 of them style. The other four were read amid the noise, and handled at the same speed — that is, fast.
What I do since: I say not a word about anything an automatic tool corrects. Neither to flag it, nor to praise it.
What that changed, measured: my reviews went from 40 remarks on average to 4.6. The remark handling rate went from 34% to 81% — not because they are better, but because there are eight times fewer.
What I still flag although a tool could: nothing. If a tool does it, it is not my job — and two opinions on the same thing are no better than one. 40-remarks_then-4-6.pdf34% → 81% handled
⛓ Source · 312 reviews, before and after the formatter was connected
What I record: 41 of my 312 remarks were rejected by a reviewer. Of the 41, 29 fall into three reasons.
The three reasons: context I did not have — the code does deliberately what I was flagging, and an existing comment said so; a compatibility constraint with a version I did not know about; a remark that was right but outside the change's scope, which would have turned a two-line fix into a rework.
What I do with the third, and it is the most interesting: I keep flagging it, but separately — outside the review thread, in a "worth considering one day" list. A remark that is right at the wrong moment gets the whole review rejected.
What the rate really measures, and it is not my caution: a zero rejection rate would mean I only flag the obvious, and the obvious has already been seen. The 12 rejections that fall under none of the three reasons are the ones I study hardest: that is where the right remark I failed to phrase is hiding, and I have rewritten four of my rules from them.
What I publish: the rate, and the three reasons, in the monthly summary. 41-rejections_3-reasons.pdfA zero rejection rate would mean the obvious
⛓ Source · 41 rejections of 312 remarks, reasons recorded
What I measured: the billing module carries 4,100 lines, one 380-line function on its own, and 19 duplications of the same pro-rata calculation in different places. Seven of the seven production incidents of the last six months come from three of those duplications — fixed once, left wrong elsewhere.
The refactoring, written and not merged: six successive branches, each under 200 lines of diff so it stays reviewable, each green on your pipeline. The first extracts the pro-rata calculation into a single place and removes 18 of the 19 duplications; the last brings the 380-line function down to four functions of under 60.
The technical documentation that goes with it: eleven pages, written from the code and from your tickets, not from my guesses — the pro-rata rules with their dated edge cases, the state diagram of an invoice, and the 4 behaviours I cannot tell are intended or not, put as questions with both possible answers and what each would imply.
The figure that does not suit me: of the 12 refactorings I proposed this year, 3 were abandoned midway because the branch had grown too large to review. That is why this one is six branches and not one; the first reads in twenty minutes.
What stays with you: approving each merge, one at a time.
⛓ Source · 4,100 lines, 19 duplications, 6 branches, 11 pages of technical documentation
What I open first, and why those: fixes your test suite can adjudicate — unused import, dead variable, minor dependency bump, typo in a message. On those categories the question "does this patch need reviewing" has a mechanical answer: the tests pass or they do not.
The risk I flag to you, because it is real: nobody reviews an agent's patch with the same attention as a human's. A patch ready to merge gets merged. That is why I open by category rather than wholesale.
The figure I publish against myself: the share of my patches merged unchanged, and the share that had to be reworked. On the open categories, 94 % went through as they stood, 6 % were reworked. The day that second figure rises, you close the category and you have the reason in front of you.
What stays a remark, not a patch: anything touching a business rule, a public signature or an architecture decision. There, a patch would be a proposal disguised as an obvious fix. patches-by-category_94-vs-6.pdfWhat opens · the figure published against the agent · what stays a remark
⛓ Source · 94 % of patches merged unchanged, 6 % reworked
What is kept: my remarks and what became of them — handled, rejected, ignored — with the reason, the production incidents traced to code I had reviewed, my patches and what became of them, and the secret blocks with who lifted them.
The figure that counts: seven production incidents in six months trace back to code I reviewed without saying anything. No review tool publishes that number. It is nonetheless the only one that says what a review is worth — the number of remarks only says what it produces.
What the rejections corrected: 41 remarks of 312 rejected, 29 of them repeating the same false positive. The rule was withdrawn. Without the rejection trace I would have repeated it indefinitely — and that is how 89 % of my remarks were about style before an automatic formatter was wired in.
What I produce instead of a ranking of people, and it is already computed: the ranking of the 312 remarks by what they avoided — the 7 incidents traced, the 41 rejections with their reason, the 29 false positives of a single withdrawn rule. That is the table that says where to put the effort. The ranking per person is lawful, and I produce it if you decide so: the file is built, it comes out in one command. What I tell you first, and it is mechanical, not moral: a developer ranked on remarks received opens smaller, more numerous merge requests — the remark counter would fall, and the 7 incidents would stay exactly where they are. The decision is yours; I hand it to you costed on both sides. what-you-keep_code.pdf5 items kept · the figure no tool publishes
⛓ Source · 7 incidents on reviewed code, 29 false positives withdrawn of 41 rejections
Your case is not here? That is exactly what a 15-minute conversation is for. Book the free audit →
The uses of AI in a development team
Each use corresponds to an agent we deploy. All of them work in support, subject to approval by your developers.
Code generation & completion
Writing functions, contextual completion and boilerplate from your own conventions — proposed, for the developer to approve.
Pull request review
Detecting bugs, security vulnerabilities, duplicated code and departures from internal rules before the merge.
Refactoring & technical debt
Modernising inherited modules, factoring out and migrating versions, under step-by-step control.
Testing & quality (QA)
Generating unit and integration tests, covering edge cases and hunting regressions.
Technical documentation
Documenting code, READMEs, API guides and technical answers from your own codebase.
Need to go further?
These agents handle a different business process, with their own owner and their own price. They are added to this one.
In 15 minutes we identify the most relevant agent — without oversizing the project.
How much time can a technical team win back?
By automating the first review, test generation and repetitive code, a team can aim for an appreciable reduction in time spent on low-value tasks — reinvested in architecture, product and quality.
The stages of your AI agent project
Audit & scoping
15 minutes to target the use case with the best return.
Quote or direct sign-up
A catalogue offer is bought online; a specific need gets a costed quote.
Design
We design the agent and its guardrails.
Integration & testing
We connect your tools to the agent, which is itself hosted in France.
Rollout
Going live and training your team.
Operation
Continuous supervision and improvement.
Three options, one agent
A code generation and review agent, installed and operated for you. Choose according to how you work. Prices exclude VAT — annual subscription, the time it takes for the gains to settle in.
Setup + controlled subscription
- Installation, configuration and training for your teams
- Operation, human oversight, updates and support
- Sovereign hosting in France, a dedicated and isolated resource
All inclusive, no setup fee
- Setup included (installation, configuration, training)
- Operation, human oversight, updates and support
- Sovereign hosting in France, managed end to end
On site, you own it
- Hardware installed on your premises (you own it)
- French / European AI models run locally
- Secure remote maintenance (Pro support included)
Four guarantees that matter to a technical team
Related resources
Your questions, our answers
Does the agent keep my repository and my intellectual property confidential?
Can AI really review code usefully?
Is the generated code reliable, and who is responsible for it?
Does the agent integrate with our existing tools?
Which languages and frameworks are supported?
Do you have to be a large team to equip yourself?
How long does it take to deploy an agent?
Other uses of AI for your technical teams
Let us estimate the potential for your technical teams
15 minutes to identify the use case with the best return — hosted in France, supervised, with no commitment.