Skip to content
Solution · AI in controlling

Not an AI pilot that stalls. A system that already runs.

For CFOs and finance directors of mid-size and large companies exploring AI in controlling: most pilots stall. GuardPilot doesn't — because it's built by an operator, has been running non-stop on our own organisation since autumn 2025, and enforces human-in-the-loop technically.

Last night 03:33 · examplesignal
S03 · Meta triage
four-eyes ✓ · both models agree
24/7
in production since autumn 2025

Why AI pilots in controlling stall.

Three causes recur: pilots built by people who never lived the domain, output without evidence that the organisation doesn't trust, and no integration with the existing decision practice — so AI hovers alongside the work instead of running through it. GuardPilot is explicitly designed the opposite way: lived knowledge, evidence before signal, human-in-the-loop.

Six reasons

Why this system does reach production.

01

Built by an operator, not a consultant

25 years leading a contract-heavy service organisation. 385,000 lines of personal knowledge on exceptions, edge cases and escalation paths — no consultancy deck, no theoretical framework.

02

Evidence before signal, always

Agents never flag blindly. Source evidence is built first, then two independent model checks follow. What reaches a human is pre-sorted and grounded — no alert fatigue.

03

Human-in-the-loop technically enforced

No mode lets the system autonomously execute bookings, credit notes or contract changes. The human decides; the agent investigates, informs and documents.

04

Own production as testbed

Running non-stop on our own contract and invoicing flows since autumn 2025. Every edge case improves the system — before it ever reaches an external customer.

05

Dutch and privacy-first

Personal data does not leave the network. Cloud analysis only on anonymised data. No US-only vendor lock-in, no mandatory US hosting.

06

Reproducible trail per decision

Every signal carries its source lines, model verdicts and decision path. Justifying afterwards or discussing internally uses the same data the agent used.

How it works in practice

AI investigates and informs. The human decides.

Every night 27 agents run through contract, execution and invoice. A triage agent decides what warrants attention; analysis agents build evidence; two independent models confirm. Only then does the signal reach a human — with evidence, a concrete proposal and a reproducible trail. No one has to trust AI without grounding; no one has to explain a decision after the fact.

GuardPilot data quality screen with 35 findings and concrete follow-up steps
Data quality as a work list: 35 findings, each stating what is wrong and what to do about it.
Proven result

Non-stop in production since autumn 2025.

27
Agents in production
20 core + 7 sub-agents
434.667
Documents processed
our own live operation
24/7
Continuous monitoring
since autumn 2025
Frequently asked questions

About AI in controlling.

Why do AI pilots in controlling stall?

Most AI pilots in controlling stall on three things: they are built by people who never lived the domain, they produce unreliable signals without evidence, and they don't change the existing decision practice — so the output hovers alongside the existing work. The result is a demo that impresses and a production environment that ignores it. GuardPilot is built by an operator who had the problem for 25 years, with an evidence-first architecture and an explicit human-in-the-loop model.

What does human-in-the-loop mean concretely?

The order is technically enforced: AI investigates and informs, the human decides. An agent never autonomously executes a booking, credit note or contract change. What the agent does do: collect evidence, have two independent models check it, and present the case to the right person with a concrete proposal and the underlying sources. Reviewing becomes a matter of minutes instead of hours — but the final responsibility stays with the human.

Why is your own organisation the testbed?

Since autumn 2025, GuardPilot has been running non-stop on our own contract and invoicing flows: 27 agents, 7 production collections, 3 compute machines. Every false positive, every edge case, every misinterpretation improves the system — before it ever reaches an external customer. Consultancy models can't replicate this; you have to have lived it to know where the edge of a clause sits.

How does this fit into existing controlling tooling?

GuardPilot replaces no ERP, no accounting, no CLM. It reads existing sources (documents, ERP, DMS, mail) and produces evidence-backed signals in the channels your team already uses. There is no migration, no new UI to adopt before value emerges. The first week covers source integrations and scope; after that the system runs in nightly cycles.

From pilot fatigue to a working system?

Request access to the waitlist for the first external pilots, or contact us for an intake call.

Related solutions

More of the GuardPilot system.