"Onboard EU clients three times faster without growing headcount at the same rate."
Shipped: 4.2 hours to 1.4 hours per audit.
When commercial insurance clients send ESG paperwork, someone still has to read it. At CIS that someone was an analyst, four hours a case. I was asked to make it faster. The real brief was to make it trustworthy enough to sign.
Analysts opened a PDF, found Scope 2 emissions on page 34, typed it into a 15-year-old system, then did it again. Transcription errors sat around 23%. CSRD was about to turn those errors from an ops headache into a legal one.
Leadership wanted 3x client onboarding without hiring in lockstep. Legal wanted zero hallucination, full stop. Engineering wanted an overlay, not a rewrite. If I missed any of those, the product would die in review.
"Onboard EU clients three times faster without growing headcount at the same rate."
Shipped: 4.2 hours to 1.4 hours per audit.
"If it cannot be traced to a cited internal document, it is a liability."
Shipped: closed-loop RAG, source-pinned, human approved before commit.
"Do not touch the legacy database. API overlay only."
Shipped: read/write layer. Engineers signed off in sprint 2.
The AI was not the hard part. Sarah needed to verify. Thomas needed throughput he could defend. Julian needed an immutable log. Miss one and the whole thing gets rejected.
"If I cannot verify where that number came from, I cannot sign. I do not care how fast it is."
Needs: inline citations, confidence next to every value, one click to the exact PDF page.
"Show me bottlenecks. Tell me if AI is speeding work, or just moving it around."
Needs: pipeline velocity, override rates, weekly quality delta versus manual audits.
"Every AI decision, every override, every source. Exportable for a regulator."
Needs: GDPR, CSRD, SFDR trail. Append-only. Named analyst. Timestamped.
Standard Double Diamond assumes you can predict the output. Aura is probabilistic. I treated model behaviour as a design material, the same way I treat a constraint.
Three weeks embedded with analysts. Shadow sessions, data audit, an AI readiness score before any UI love.
Intent maps across 240 audit sessions. Named the exact moments autonomy must yield to judgment.
Explainability seams: reasoning, confidence, and source as first-class UI, not a tooltip afterthought.
Wizard of Oz before the model. We planted a 5% calculation error to test automation bias. 100% of analysts caught it.
I pushed for closed-loop RAG early. Under GDPR, client emissions data hitting an external API is an egress problem. If the model can only reason about documents already inside the private cloud, hallucination becomes a structural issue, not a prompt issue.
If a figure is not in the indexed source, Aura returns "No verified source found." It never invents a number to be helpful.
Extraction, comparison, compliance check. Separate confidence rules, so one failure does not paint the whole screen green.
Nothing writes to the legacy system until a named analyst approves. That click is the audit trail Julian asked for.
Tooling stayed inside enterprise licences: Figma AI Enterprise, ProtoPie, Azure OpenAI in EU region, Dovetail, Maze. Legal blocked consumer AI on client data. That constraint made the design better.