Building the Agent Factory: A Blueprint for AI Product Management
In my last edition, I wrote about the gap between having {accurate, well formed, and timely} data and doing something with it and why closing that gap means moving beyond purely generative AI.
What I did not get into is how. And that is where, in my opinion, the fun lives.
The main question I get asked is some version of: how accurate can you really make an agent? It sounds like a technology question. It is really a process question. And the answer, in this context, is: as accurate as your understanding of your own process. This can be a humbling answer for most organizations, because it shifts accountability away from the technology and back to the workflow.
But once you can answer that question honestly, a different question opens up. Not how accurate can you make it but what is that accuracy worth when it runs continuously, across thousands of decisions, without fatigue or variance. That is where the value conversation actually starts.
The technology question that does matter — is where in the workflow you are using AI versus deterministic automation.
Here is what I have learned working through that question with real teams on real workflows.
Autonomous Action Is Not What You Think
The thing that makes people nervous about agentic AI is the word autonomous. It conjures a black box making decisions in your systems without supervision. That fear is understandable. It is also, in my experience, mostly misplaced — but understandable given the conflation with probabilistic generative AI.
In a well-designed agentic system, AI handles what AI is actually good at: reasoning through unstructured inputs, managing variance, interpreting ambiguous information, and making judgment calls within defined parameters. The action, the thing that actually happens in your systems, is executed by deterministic automation. High precision, rule-based, auditable, traceable.
That layering is not a technical detail. It is the architectural principle that makes the difference between a solution that is clever and one that is accurate, cost-effective, and safe enough to trust at scale. AI and automation are not interchangeable. They are complementary. Using them that way is what makes agentic systems… agentic.
Once teams understand that - that the AI reasons and automation acts, the conversation about oversight becomes much more productive. And when the oversight is productive, something else comes into focus: the relationship between agents and value. An agent that is not accurate enough is a liability. An agent that crosses the threshold of trust in a high-volume workflow is a different category of asset entirely - because accuracy at scale is not just reliability. It is economics.
Getting the Guardrails Right
I have sat in a lot of workshops where the guardrail conversation becomes a proxy for a deeper anxiety. How much latitude do we give the agent? How does this impact our compliance? The instinct is to say: as little as possible while still getting value.
That instinct makes sense. It also tends to produce systems that nobody trusts or adopts because nothing meaningful ever goes through them unsupervised.
The better starting point is a harder question: what is the actual consequence if this gets it wrong? Not in theory — in practice. What happens downstream? Who finds out, and how quickly? Can it be corrected?
In regulated industries, those questions have explicit answers. Penalties, policies, thresholds. The stakes are already legible. In less-structured environments, surfacing that answer is part of the design work. I have seen teams spend weeks debating oversight architecture before anyone had mapped the error scenarios. That is incomplete designing.
The factors I come back to are: how bad is wrong, how fast does it travel, and how clear are the rules. Errors that surface quickly and correct easily warrant different guardrails, and attention, than errors that move downstream before anyone catches them. Sensitive data and ambiguous edge cases need human judgment in the loop — not as a fallback, but as a deliberate part of the design.
What I have found is that the right oversight model is almost never the most conservative one. It is the one that matches the actual risk profile. Sometimes that means a human reviews exceptions before action is taken. Sometimes it means the human sets the thresholds and the system handles everything within them. In lower-stakes workflows, retrospective review is often enough. There is no default. There is only a decision made honestly.
Building the Team That Runs the Factory
Most organizations are training people to use AI. Few are training people to build and manage it. That distinction is going to matter more than most realize.
When an agent is running a workflow, the human's job is not to do the task - it is to understand the system well enough to know when something is off. That requires a different kind of knowledge than task execution does. It requires knowing where the process came from, why it was built the way it was, and what it was compensating for.
Most workflows were built around people, with assumptions and constraints baked in that were never fully documented — because someone had been handling the exceptions for years and everyone just knew that. Before you can hand that work to an agent, you need to decompose it: understand how often it changes, where the variance lives, how exception-heavy it actually is. That work surfaces the business rules the agent needs to operate within. It also reveals, almost every time, how much of the current process exists only because it was designed around human limitations.
A useful question to pressure-test any workflow: is this step here because it creates value, or because a person needed it? The audit that reviews a sample instead of everything — is that the right methodology, or a capacity constraint that became a standard? The approval that waits for a signature before the next stage can begin; is that sequencing logical, or does it trace back to a paper form that had to move between desks? Those constraints are worth naming before they get automated in.
On the operating side, the people running these systems need genuine fluency with the data, not just familiarity with outputs, but an understanding of where the data comes from and where it can mislead. They need to understand how the system fails, across both the AI reasoning layer and the automation execution layer. In my experience, people who come from support functions or whose work has centered on diagnosis tend to be naturals at this. They already know how to look for what is wrong rather than confirm what is right. Oversight that is not grounded in that understanding is not really oversight. It is comfort theater.
The New Org Chart
Something I have started noticing in more mature agentic deployments is a quiet shift in how work is organized — not announced, just emerging.
The people closest to the workflow stop being the ones doing it. They become the ones managing the system that does it. They define the rules, monitor the outcomes, handle the genuine exceptions, and gradually expand the scope of what the agent can take on. In effect, they have acquired staff. The staff just happen to be agents.
This is where enterprise AI is actually heading, and it is a bigger organizational shift than most leadership teams are accounting for. Not everyone will have an agent portfolio at the same time but the trajectory is toward a model where the question is not whether you have agents working for you, but who builds and maintains them, how many are there and what are they doing.
The model that works is not top-down deployment. It is enablement. It is literacy. Give the people who actually know the workflow, the ones who have been handling the exceptions for years, the tools and guardrails to build their own agents. Not IT. Not a Center of Excellence working from a requirements document three layers removed from the actual process. The person who knows where the edge cases live is the right person to build the first version.
What they need is a clear path: accessible tooling, guardrails that keep early experiments safe without strangling them, and a graduation model — from team-level agent to shared departmental capability to enterprise platform. That graduation is what keeps it from becoming shadow IT. The platform does not get handed down from a steering committee. It gets discovered through deployment and hardened through use.
The deeper implication is one I do not think we talk about honestly enough. When work shifts from execution to operation — when people's jobs become about managing systems rather than doing tasks — that is not a neutral change for the people in those roles. Getting this transition right means being direct with your workforce about what is shifting and why, not just optimizing the workflow and hoping people find their footing. The organizations that handle that well will build more durable capability than the ones that treat it as a change management footnote.
Where I Would Start — And Where I Was Wrong
If you had asked me a year ago which function was furthest along in agentic AI specifically — not traditional automation, but systems that reason, interpret, and act… I would have said supply chain.
I was wrong.
Finance keeps coming up instead, and for a while, that surprised me. What I have come to understand is that finance already has some of the properties that make agentic systems work: explicit rules, policy-driven processes, low tolerance for variance. The tribal knowledge that exists in every finance function — once surfaced and formalized — turns out to be more deterministic than I expected. Large swaths of finance operations already run on defined thresholds and policy rules. That part already thinks like an agent. It just has not had one yet.
Beyond finance, the functions showing the strongest results share a common pattern: reasoning through unstructured inputs before pushing a result into a system of record. The agent earns its keep at the interpretation layer. The automation earns its keep at the execution layer.
Insurance claims processing fits that pattern directly. Claims arrive as a mix of structured and unstructured data, need to be interpreted and routed, and have clear policy rules governing the outcome. The workflow has natural decision points and an unambiguous end state — which is exactly the architecture that works.
Quoting is another area gaining traction faster than I expected. We recently completed a project with a construction engineering firm where the workflow involves ingesting incoming requirements, gathering supporting materials, generating estimates, identifying opportunities to win, and generating a natural language response. The combination of structured business rules and unstructured inputs is where this architecture performs. And I think that pattern — interpret, apply rules, generate output, execute — is going to show up across more functions than most people are currently planning for.
Build the Foundation First
Most conversations about agentic AI start with the platform question. Which vendor. Which hyperscaler. Which architecture.
That question matters. Especially given the outlined importance of design and architecture in an agentic solution. It should not be the only one.
In concert, one should ask: what is the workflow? Understand it before any architecture decision is made — its rules, its exceptions, its data dependencies, its variance. Then look honestly at integration touch points and data readiness. An agent is only as good as what it can act on. A well-reasoned agent on a weak data foundation is still a weak agent.
There is no single right architecture at this stage of the market. However, there is a place for architectural archetypes based on categories or patterns of agent workflows. What makes one approach better than another is not the vendor name — it is how well the design fits the actual workflow, the true risk profile, and the gravity of the existing systems it has to connect to.
The factory gets built one workflow at a time. But something I have noticed is that the teams doing this well stop thinking about individual deployments after the first few. Each one teaches them something about their own processes, their own data, their own edge cases. That knowledge compounds. Good news is, the tenth agent is not as difficult as the first — it is significantly easier, because the organization has started to understand itself differently.
That is the outcome worth building toward. Not a portfolio of agents, but an organization that knows how or who to build them.
And that brings the accuracy question full circle. The teams that get the most out of agentic systems are not the ones that found a better model. They are the ones that developed a clearer understanding of their own processes — their rules, their exceptions, their data, their edge cases. The technology got them started. That understanding is what made them accurate.
Which means the real return on agentic AI is not just operational. It is organizational. You end up knowing things about how your business actually works that you did not know before. That knowledge does not disappear when the agent runs. It compounds every time you build the next one.
Connect with an Argano Expert!
Need specialized insights for your business challenges? Facing complex business technology questions? Don't navigate alone. Connect with an Argano subject matter expert who will personally respond within 24 hours.