Summary
An agentic AI harness connects a model to tools, data, memory and workflows. Its design shapes what an agent can access, decide and do. Businesses need to govern permissions, skills, handoffs and approval gates across the whole system. Choosing a capable model matters, but it does not replace process design, security controls or clear accountability.
In my previous article, I wrote that AI safety is a workflow design problem. The responses pushed the conversation somewhere worth following.
One commenter argued that workflow design gets harder when the model is a black box you’re renting by the token, and that running open-source models means you control the failure modes. Another said the best companies aren’t the ones with the most AI tools, but the ones that integrate them best.
Both are pointing at something real. Both still keep the model at the centre of the conversation.
I want to move it. The timing matters, because major platforms are connecting AI models to more of the tools, data and workflows organisations use every day.
The agentic wave has arrived at work
Agentic AI is moving from demos into workplace software. These systems can use tools and carry out sequences of tasks, rather than simply return an answer.
Microsoft’s Copilot Cowork lets users delegate multi-step work across Microsoft 365. Anthropic’s Claude Cowork uses plugins that bundle skills, connectors and sub-agents, with private marketplaces for organisations. Google’s Workspace Studio lets teams create, test and share skills based on their ways of working. Features and availability vary by product and subscription.
Look past the branding and a common architecture is emerging: a model, connected to data through connectors, given procedures through skills, coordinating work through workflows or sub-agents, and surrounded by admin controls.
That surrounding software is often called an agentic AI harness. It connects the model to tools, data, memory and workflows, and governs how proposed actions are executed. The model is one component, and it can be replaced while much of that surrounding system remains.
Even the labs’ boundaries leaked
The July 2026 Hugging Face incident showed why those boundaries matter. OpenAI’s investigation reported that models operating with reduced safeguards during cybersecurity evaluations circumvented isolation controls and accessed third-party systems.
These were unusual research conditions, not a typical enterprise deployment. The incident nevertheless exposed a combination of model behaviour, infrastructure weaknesses and oversight failures. A boundary in the design intent was not enough to contain the agents’ actions.
If well-resourced AI labs can leave gaps in their evaluation environments, organisations rolling out agents across email, files and business systems need to test their own boundaries too.
Australia’s cyber agency has drawn the same line
On 11 September 2026, the Australian Signals Directorate published Agentic AI Harnesses: The layer above the model. “Harness” covers everything in an agentic system other than the model: connectors, tools, memory, permissions and controls.
Its central point is that anything an agent reads, whether web pages, documents, emails or code comments, may be treated as an instruction. This is the weakness behind prompt injection, where instructions embedded in content try to redirect an AI’s behaviour, and ASD says no fully reliable technical mitigation currently exists. So the controls have to live in the harness, by limiting what an agent can access and what it’s allowed to do.
This adds a third failure mode to the two I wrote about last time. An agent can misjudge. An agent’s environment can leak. And an agent can be instructed by someone who isn’t you, through an email, a supplier document or a web page it reads while doing its job.
That “someone” may not even be a person. Increasingly, the content your agent reads will have been written by another AI: a supplier’s agent drafting emails, a customer’s assistant filling in your forms, AI-generated web pages. An attacker doesn’t need to handcraft the perfect malicious instruction. They can have a model generate thousands of variations and keep whichever ones work. That makes it risky to rely on detecting a fixed set of suspicious phrases.
Defences that rely on spotting suspicious wording are up against an adversary that never runs out of wording.
That’s why “never send customer data externally” in a system prompt is a policy, not a control. An attacker’s email may influence how the model follows a policy. Independently enforced permissions can limit what that influence allows it to do, provided the surrounding system has no bypass.
On owning the model
The open-source argument deserves a fair hearing, because part of it is right. Renting a model means accepting pricing changes, deprecations, behaviour shifts between versions and questions about where your data goes. For organisations with sovereignty obligations, running open-weight models on their own infrastructure can be the right call.
But owning the model doesn’t mean owning the failure modes that matter most. Prompt injection isn’t a vendor bug. It comes from how language models process context, open or closed. A self-hosted model with access to your inbox and an outbound email tool can face the same class of exposure as a rented one. The actual risk depends on the controls around it.
I know this first-hand, because I built my own. More on that below.
Open-source agents: you can see the harness, and you own it
The open-source argument goes well beyond models now. Hermes Agent, from Nous Research, is one example of an open-source harness built around these capabilities.
Hermes supports different models. Its proposition is the surrounding system: tools, skills, memory and ways to coordinate work.
It creates and updates skills, keeps persistent memory, connects to messaging platforms and external tools, and supports different execution environments. Nous reports using Hermes to refactor a million-line Python codebase with 1,393 sub-agents over nineteen hours. That is a vendor-reported example, but it illustrates the scale at which delegation can operate.
That’s a list of capabilities, and it doubles as a checklist of harness questions:
- A self-improving agent writes its own skills. Who reviews a procedure when the agent is both its author and its user?
- Persistent memory accumulates. What has it learned, and could any of it have come from content it shouldn’t have trusted?
- A messaging gateway is an input channel. Who can send it instructions?
- Scheduled jobs run unattended. What can happen overnight before anyone looks?
- Thousands of sub-agents means the chain question at industrial scale.
None of this is a criticism of Hermes. It is the review any agent needs: who can trigger a skill, which tools it gets, what credentials it can read and how to stop it. A limited toolset and an isolated execution environment can help reduce exposure.
OpenClaw provides a useful caution. In February 2026 guidance, Microsoft’s security team advised against running it on a standard personal or enterprise workstation. For evaluation, it recommended an isolated environment, dedicated credentials and non-sensitive data. Its analysis emphasised the combined risks of third-party skills, untrusted inputs and persistent access. Making a harness visible does not ensure its deployment is controlled.
Moving to a different harness also calls for a review of any skills, memories and permissions carried across. A migration should not automatically carry forward trust.
This is open source’s real advantage: you can read the harness. Nobody has to take a vendor’s word for how permissions, memory or tool access work. But an open-source agent gives you the ability to govern the harness. It doesn’t do the governing for you.
At enterprise scale, that governing starts with a posture decision: prohibit these agents, contain them in a sanctioned sandbox for pilots, or run them as a managed internal platform. From there it means one approved deployment pattern, central decisions about identities, tools and which channels can instruct an agent, and someone accountable for patching. Self-improving agents add something new: a promotion gate. Nothing an agent teaches itself should be relied on for consequential work until a person has reviewed it.
And because these agents install in minutes and connect to personal messaging and email accounts, the first question for many organisations is simpler: do you know whether your people are already running one?
On integration
The second comment is half right too. Integration beats accumulation, but every connection is a permission. Security researcher Simon Willison describes a “lethal trifecta”: access to private data, exposure to untrusted content, and the ability to communicate externally. Connect an agent to email, file storage and CRM, and you’ve very likely built all three. The best organisations won’t integrate the most. They’ll integrate deliberately.
What actually surrounds the model
In business terms, seven things around every AI worker decide whether it’s helpful or hazardous:
- What it reads: which content from outside the organisation could act as an instruction, including content written by other AI.
- What it can reach: which systems, data and identities it can access.
- What it can do: whether it informs, recommends, drafts or executes, and which tools and skills it can combine.
- Who checks: where human and rule-based checks sit, and whether they change the outcome.
- What it remembers: what persists between sessions and who curates it.
- What gets recorded: whether you can reconstruct what it saw, decided and did.
- How it stops: limits, kill switches and rollback.
These questions cannot be answered by choosing a model alone. Three newer ideas are putting pressure on all of them at once: multiple agents, skills and gates.
When one agent becomes many
The next wave isn’t a single assistant. It’s several agents working together, with one gathering, another analysing, a third drafting, and an orchestrator coordinating them. Different platforms implement this through sub-agents, orchestration or multi-step workflows.
It mirrors how teams work. It also brings a problem anyone who has mapped a cross-functional process will recognise: handoffs are where things go wrong.
A receiving agent may treat upstream work as trusted unless the workflow preserves its source and trust boundaries. The research agent read a supplier’s web page. The analysis agent reads the research agent’s summary. By the time the output reaches the agent that can act, content that started as untrusted has been laundered through layers of apparent legitimacy.
This is why ASD recommends treating a multi-agent system as a single agent for security purposes: a compromise in one component can spread through shared context and trust. No single agent may hold all three trifecta legs, but if three agents each hold one and pass work between them, the system does.
Assess the agents and the chain they form.
And the chain doesn’t stop at your organisation’s boundary. Agent2Agent (A2A) and similar protocols are designed to let agents from different vendors and organisations work together. When your procurement agent negotiates with a supplier’s sales agent, it is receiving messages and proposals from an AI you didn’t build and don’t control. Those messages must not become authority to act simply because another agent sent them. It may have been compromised upstream. It may simply be optimising for its owner at your expense.
Treat other organisations’ agents the way you’d treat a stranger’s email: useful, possibly legitimate, and never trusted to authorise anything on its own.
Skills are process design made executable
Skills are the most interesting development of the lot, and the most underrated.
A skill is, at its core, a reusable set of instructions an agent follows: how we do this task, in this order, to this standard. Skills may also include scripts, templates or supporting resources. Platforms increasingly support shared skills; Hermes can create and update them itself.
In other words, skills are process documentation that executes.
That’s a big opportunity. Organisations with clear, current process knowledge have a stronger starting point for using agents effectively. After mapping tens of thousands of processes, I’d argue it’s the strongest argument for process discipline I’ve seen in years.
It also comes with three risks. First, a skill codifies whatever process you have, including the workarounds, the outdated steps and the rules nobody remembers the reason for. Automating a poor process institutionalises it; turning it into a skill institutionalises it at machine speed, across everyone who installs it. Second, a skill is an instruction, not a control. “Always get approval above $10,000” in a skill is still a policy the model is asked to follow. If the threshold matters, enforce it in the system. Third, skills are a supply chain. A third-party skill is a set of instructions your agent will trust, written by someone else. Some vendors now scan skills for malicious content, but a scanner is a backstop, not an ownership model.
Treat skills with the discipline of business process improvement and controlled software changes: named owners, version control, review before release, and retirement when the process changes. That applies whether a person, a vendor or the agent itself wrote the skill.
Gates are where the design lives
Tools are where words become actions: sending the email, updating the record, raising the purchase order. In an agentic workflow the model often chooses which tools to call and in what order, so dangerous combinations may never have been designed by anyone. A file-reading tool and a messaging tool may each have a legitimate purpose. Combined in a manipulated agent, they can provide a route for sending private information outside the organisation.
One way to constrain that risk is a gate: a point where the workflow stops being probabilistic and becomes deterministic again. Something must be true before work can proceed.
- Approval gates, where a person decides.
- Validation gates, where rules check the output.
- Threshold gates, where value, volume or scope limits are enforced by the system.
- Trust gates, between agents (including agents from outside the organisation), where upstream content is checked before a downstream agent can act.
- Change gates, where a new model, tool or skill is tested before entering the workflow.
Platforms offer approval and administrative controls, but their scope varies. Microsoft’s Copilot Cowork documentation, for example, describes approval before actions are carried out. The practical task is to check which controls apply in your deployment, how they are configured and where they belong in each process.
Gate quality matters as much as placement. A bare “approve / reject” button, fifty times a day, can become a rubber stamp with a login. A good gate shows what the agent wants to do, why, and what prompted it, and appears only where human judgement changes the outcome.
None of this argues for locking AI down. Much of AI’s value is in finding, reading, drafting and recommending. Those activities may need less restrictive controls, but the sensitivity of the information and exposure to untrusted content still matter. Where autonomy makes sense, gates are what make it possible. Autonomy should be earned, not granted. An agent can start by recommending while humans act. As evidence builds, it executes low-value actions on its own, then sampled review replaces approval. If behaviour changes after a model update, a new tool or an incident, the gates tighten again.
What I learned building my own
I don’t just write about this. At home I run a small multi-machine AI setup: local open-weight models, a voice assistant, document search, and assistants connected to household tools for finances, property and records. I built it for the reason the first commenter gave: control over my own data.
It taught me more about harnesses than any vendor briefing.
The models were the easy part. I’ve swapped models several times, with routing that falls back from one to another when something is unavailable. Each swap took hours. The harness around them took months, and it’s the part that has lasted.
Self-hosting meant owning the operational failures too. A search component once broke silently, returning success codes and empty results with no errors logged. A background service was killed every three days by an operating system default I didn’t know existed, and nothing restarted it because I’d never built a watchdog. A managed provider may handle parts of that infrastructure. Self-hosting makes those responsibilities yours.
- The tool returns the figures; the model never supplies them. For anything financial, the model can explain a number but never produce one, regardless of model size.
- The model never supplies identity. Who is asking is established by authentication, never by what the model says.
- Exact match beats reasoning for deterministic tasks. If a rule can decide it, don’t ask a model.
- Inert by default. New capabilities start switched off.
I used separation to reduce exposure. My general assistant has tools but no command line. A separate coding assistant has a command line but no tools, and there’s no channel for one to delegate to the other. Confidential professional work runs on a separate machine, with data flowing to it in one direction only. Different family members get different access.
Write access was earned. Assistants that can change financial or property records stage their changes for my approval. The first capability allowed to write directly was a workout log, because it was the lowest-consequence data I had.
Tests had to prove they could fail. More than once, a check I trusted was passing because it couldn’t have failed. Now a test isn’t trusted until I’ve watched it catch a problem.
None of this required frontier AI. It required the same unglamorous discipline as any system design: identity, permissions, staging, logging, failure handling and documentation. An agent can expose gaps quickly, particularly when it has the authority to act without waiting for someone to notice.
Humans were the undocumented control
Most of what AI exposes existed long before AI: processes nobody documented, permissions nobody reviewed, data nobody owns, workarounds that became the process.
Those weaknesses were survivable because people absorbed them. The accounts clerk who knew that supplier always sends odd invoices. The warehouse manager who knew the stock was idle because of the renovation. The coordinator who noticed an email looked wrong and picked up the phone.
Humans were the undocumented control. Put an agent into that process without accounting for those checks, and the judgement quietly holding it together can disappear.
Everyone wants the genie’s wishes: faster cycles, lower costs, fewer manual steps. Some of those benefits are real and some are still only perceived. Either way, the demand is forcing a conversation many organisations have avoided for years. AI isn’t creating a new discipline. It’s collecting on an old debt.
Questions for vendors and boards
ASD’s guidance includes questions for board directors. The one I’d put in front of every executive is the last: what’s the worst outcome if the harness is compromised, misconfigured or manipulated, and which controls would contain it?
For any agentic platform, whether it’s Copilot, Claude, Gemini, ChatGPT, Hermes or something built in-house, ask:
- What identity does the agent act as, and can we restrict it below the user’s own access?
- Which tools, connectors and skills are enabled by default, and who approves new ones, including skills the agent writes itself?
- How does it handle instructions embedded in documents, emails, web pages or messages from other agents?
- Which actions need approval, and can the agent bypass that?
- What’s logged, and can we export it?
- Can we pause, cap or roll back what the agent does?
- What changes when the underlying model is updated or swapped?
If a vendor can’t answer these clearly, that’s your answer. With a self-managed open-source deployment, your organisation must be able to provide those answers, drawing on maintainer or commercial support where available.
The wish, not the genie
The model race will keep going. There will be new releases, benchmarks and arguments about open versus closed. It’s fascinating, and mostly beside the point for organisations trying to get real value safely.
The genie will keep getting more powerful. What organisations control is the wording of the wish, the reach of the genie, the gates it has to pass, and whether there’s a way to undo it.
That’s workflow design. It always was.
If you’re introducing AI agents into business workflows, SMC’s AI consulting services can help you consider the strategy, processes, data, systems and governance around them.
About the Author

Roger Molina
Roger Molina has more than 13 years of experience in enterprise system strategy, business process improvement and change management. His career spans change management, business analysis and leadership of SMC’s Strategy & Roadmap practice, with more than 400 projects delivered across mining, utilities, oil and gas, logistics, manufacturing and healthcare. He now serves as Director of AI & Innovations at SMC, helping organisations apply AI, automation and emerging technologies to improve the way they work. Roger’s research into leadership, motivation, structure and data quality has been published in the Reliability Engineering Journal.
Platforms series.
Get more insights from Solution Minds Consulting in Google Search.
Latest Insights
- Payroll Failures Don’t Start at Go-Live. They Start Months Earlier.
- What Is Oracle E-Business Suite and What Are Organisations Doing Next?
- MYOB Exo End of Life: What It Means for Mid-Market Organisations
- Microsoft Dynamics AX 2012 End of Support: What It Means for Enterprise ERP Strategy
- Microsoft Dynamics NAV End of Life: What Organisations Should Do Before 2028
