It holds permissions
Credentials that reach real systems, which means the blast radius of a mistake is now a question someone has to be able to answer.
An agent that only talks is a demo. An agent that acts is a system — with permissions, failure modes and an audit requirement, and almost none of it decided by prompt quality.
Calling tools, writing to your ERP, sending messages, moving a ticket — the moment a model does any of that it acquires permissions, failure modes and an audit requirement. That transition is where most agent projects either become valuable or become a liability, and the difference has almost nothing to do with prompt quality.
Credentials that reach real systems, which means the blast radius of a mistake is now a question someone has to be able to answer.
Multi-step processes fail halfway. Loops retry. A wrong answer no longer stops at the screen — it lands in a system of record.
Somebody will eventually ask which tools were called, with what arguments, on whose authority, and what came back.
The most reliable way to waste six months is a use case workshop that produces a prioritisation matrix. The opportunities that pay are found by sitting with the people doing the work and watching where the time actually goes, which is usually somewhere nobody would have put on a slide. Our practical filter is four questions.
Does this task happen often enough that automating it matters?
Is the input available in a machine-readable form, or does somebody have to find it first?
Is a wrong answer recoverable, and how quickly would anyone notice?
Does a specific person own the outcome, or is it spread across three departments in a way that means nobody will adopt it?
None of this is exotic. It is the ordinary discipline of production software, applied to a component that happens to be non-deterministic.
An agent should hold the narrowest credentials that let it do its job, and the blast radius of a mistake should be something you can state in one sentence.
Reading is cheap to get wrong. Sending, paying, deleting and committing are not, and those steps belong behind a human confirmation until the error rate is measured and accepted.
Agent quality drifts with every model update, prompt change and tool change. Without a fixed set of cases and an automated way to run them, you cannot tell an improvement from a regression — and you will hear about the regression from a user.
Multi-step processes fail halfway. The system has to know what it already did, and be able to undo it.
Which tools were called, with what arguments, on whose authority, and what came back. This is what makes an incident investigable, and it is also most of what an auditor asks for.
Agent loops that retry are the standard way a pilot’s economics quietly stop working at production volume.
For European organisations the obligations follow from where the system sits in the risk classification, and a good deal of what the EU AI Act asks for is documentation, human oversight and traceability that a properly built agent produces as a by-product.
Our role is to make sure the engineering produces that evidence during the build, rather than having it reconstructed afterwards by people who were not there. We work alongside your legal and compliance function rather than replacing their judgement.
Build, Buy, or Both
The frameworks, model providers and agent platforms are moving quickly, and they are not where a manufacturer or an energy company builds an advantage.
The orchestration layer, the model providers and the agent platform. This is a fast-moving commodity layer, and running your own is rarely where the advantage sits.
Your data, your process knowledge and the tool integrations nobody else can replicate. This is the part that is actually yours.
Data residency requirements, how far your workflow deviates from what the platforms assume, and what you can realistically staff.
Against your actual requirements rather than a feature matrix. We hold no reseller agreements with any of them.
Where This Applies
Structured data pulled out of PDFs, scans and CAD exports — and then checked, which is the half that usually gets skipped.
Assembling a fact base from systems that do not talk to each other, ahead of a conversation where that base decides the outcome.
Workflows carrying a heavy documentation load, where the evidence trail is as much of the work as the decision.
Where the bottleneck is finding and reconciling information across systems rather than producing the final document.
A pilot needs a defined kill criterion, so that stopping is a normal outcome rather than an admission of failure.
From the real workflow to a scoped pilot
We spend time in the real workflow with the people who run it, then come back with what survives the filter.
Time with the people who run the process, watching where it actually goes rather than collecting nominations for it.
The tasks that survive the four questions — frequency, machine-readable input, recoverability, and a named owner.
A recommendation for the orchestration layer against your data residency, deviation and staffing constraints.
An architecture that includes the evaluation harness and the approval model from the start, and where it makes sense a scoped pilot with the kill criterion agreed before it begins.
The model layer underneath an agent — where it creates measurable value, and how to deploy it with governance and control.
Where an agent’s evaluation set comes from when you cannot collect or label enough real cases.
Vendor-neutral advisory on the decisions around an agent programme, from sourcing to delivery leadership.
Where the line between platform and domain logic falls in your case, decided against requirements rather than vendor slides.
What changes when the agent has to run inside your perimeter — device class, power budget, and an update path you control.
What Uber’s Agentic Pods get right: the most valuable workflows are discovered inside the work, not nominated in a workshop.
We will tell you whether it survives the filter, what it would take to run it safely, and where stopping would be the right call.
Book a Strategy & Architecture Review