Part 2 of 6 in a series about Harness Engineering. Previous article:
Part 1: Harness Engineering
Tool orchestration and guardrails, or how a single GitHub issue steals your npm token
In February 2026, someone opened an issue on the Cline repository. Not a malicious commit, not a compromised dependency, not a phishing email to a maintainer. An issue. The kind of thing that arrives in every open-source project a hundred times a week.
The title of that issue contained instructions. Cline's automated triage workflow read it, and the agent behind that workflow did what the title told it to do. Four steps later — prompt injection through the title, extraction of the npm publish token from the workflow environment, poisoning of the CI artifact cache, publication of a malicious package to the registry — the attack was complete. Public disclosure came on February 9. It was exploited in the wild on February 17.
This was not an isolated cleverness. By April, researchers had demonstrated that a single injection pattern worked across multiple popular GitHub-integrated coding agents. By July, a technique nicknamed GitLost was tricking GitHub's own agentic workflows into leaking private repository data into public comments. And separately, an injection was shown to make Cursor's agent write a malicious MCP configuration file without user approval, which then produced remote code execution.
Different products, different vendors, same root cause. And it is not the models.
A model cannot tell instructions from data
This is the single most important thing to understand about agent security, and it is structural rather than incidental. A language model receives one undifferentiated stream of tokens. Your carefully written system prompt, the user's actual request, the contents of a file it just read, the text of a ticket, the body of an email, the comment someone left on a pull request — all of it arrives as the same kind of thing.
We have decades of intuition about this problem from a different domain. SQL injection exists because a database receives a string and cannot tell which parts were written by your developer and which parts were typed into a login box by an attacker. We solved it with parameterized queries: a hard structural boundary between code and data, enforced outside the database's interpretation of the string.
No equivalent structural fix exists for language models today. Prompt-level defenses help — telling a model to distrust content from certain sources measurably reduces successful attacks — but they are probabilistic, and a probabilistic defense against a deterministic attacker is a losing position. OWASP still ranks prompt injection as the leading driver of agentic AI security failures in production.
If you cannot make the model reliably refuse, make the refusal happen somewhere the model does not control.
That somewhere is the harness.
The rule: propose, then dispose
In a properly built harness, the model never executes anything. It emits a structured request — an intent to call a tool, with arguments. The harness receives that request and does four things before anything happens in the real world: it validates the shape and the arguments, it checks the request against a permission policy, it executes only if the policy allows, and it logs the whole transaction either way.
The result comes back to the model as structured data, not as free text spliced into a conversation. Success and failure both come back the same way, which matters more than it sounds: an agent that receives a clean, structured failure can reason about it, whereas an agent that receives a wall of stack trace tends to flail.
Run the Cline attack against that architecture and it stops at step two. The injected instruction still reaches the model. The model may well still decide to comply. But the request to read a credential from the environment, or to publish to a package registry, hits a policy layer that was written by a human, does not read the issue title, and cannot be talked out of its position.
Permission tiers, and where to draw them
The useful mental model is not a binary allowed/blocked list but a ladder, sorted by how hard the action is to undo.
-
Automatic — reads, and writes that are versioned and reversible. Reading a file, querying a database, writing to a branch, creating a draft. If the worst outcome is that someone reverts a commit, the friction of an approval gate costs more than it saves.
-
Confirm — anything irreversible or with external blast radius. Deleting, deploying, sending, paying, publishing, arbitrary shell execution. The agent drafts; a human presses the button. Note that arbitrary shell execution belongs here permanently, because it is not one capability but the union of all of them.
-
Never — credential access, permission modification, and anything that changes the harness's own configuration. An agent that can edit its own guardrails does not have guardrails. This is the tier the Cursor MCP-config attack exploited.
The tiers should be scoped as tightly as the work allows. An agent that only needs to read from one repository should not hold a token that reads from all of them, for the same reason we stopped handing out database admin credentials to application services fifteen years ago. None of this is novel security thinking. It is ordinary least privilege, applied to a new kind of caller.
Guardrails: four categories, not one
Permissions govern what an agent can do. Guardrails are the broader family of checks that sit between model output and anything downstream, and they fall into four distinct categories that teams routinely conflate.
- Behavioral — content safety, tone, and format — what the agent is allowed to say.
- Data — PII detection and classification, so sensitive values cannot leak into logs, traces, or output.
- Tool and action — the permission tiers above, plus draft-before-execute for anything irreversible.
- Operational — token budgets, wall-clock limits, and retry caps that bound the blast radius of a loop nobody anticipated.
The fourth is the one that gets skipped, and it is the one that shows up on an invoice. A verification loop with no retry ceiling is a perfectly ordinary piece of engineering until the day it fails to converge.
The compliance picture, accurately
There is a version of this argument that leans on regulatory deadlines, and it needs handling carefully, because the deadlines have been moving and several widely circulated summaries are now wrong.
Colorado's AI Act has been delayed twice. It was originally set for February 2026, pushed to June 30, 2026, and then — under SB 189, signed in May 2026 — moved to January 1, 2027 and substantially narrowed in scope. In the EU, the AI Act's high-risk obligations did become applicable on August 2, 2026, but the Digital Omnibus, in force since late July 2026, deferred the bulk of them: standalone Annex III high-risk systems to December 2027, and AI embedded in regulated products under Annex I to August 2028.
So the honest framing is not that a compliance cliff arrives next quarter. It is that the direction of travel is unambiguous and the timeline keeps getting handed back to you. Meanwhile the requirement that is already here, today, with no delay mechanism, is the one your auditors bring: SOC 2 engagements increasingly expect runtime AI controls as evidence, and evidence means logs. Which is to say the guardrail layer is not only a security control — it is the thing that produces the artifact an auditor will ask to see.
The failure mode nobody plans for
Guardrails rot. A permission set written when an agent did one narrow thing quietly becomes wrong as the agent's remit expands, and the danger is worse than having no guardrails at all, because a stale policy still produces the feeling of safety. Re-audit them on a schedule, and treat any expansion of an agent's scope as a trigger for review.
The opposite failure is real too. Every approval gate adds latency and asks for human attention, which is the scarcest resource in the system. A harness built for a regulated bank, dropped onto an internal prototyping team, will feel like filing taxes — and the predictable outcome is that someone builds a shadow workflow to get around it. The right amount of guardrail is a function of blast radius, not of institutional anxiety.
Where to start
If you have agents in production today and no permission layer, this is the first thing to build, ahead of every other layer in the series. Not because it is the most intellectually interesting — it is not — but because the failures it prevents are the ones that are hardest to recover from and most likely to end up in a disclosure.
The practical starting point is a single question, asked of every agent you are running: if a stranger could write text that this agent will read, what is the worst thing that text could cause to happen? Issue titles, ticket comments, email bodies, web pages, file contents, and code review comments all qualify. In most organizations, running that exercise honestly for the first time is uncomfortable — and considerably cheaper than the alternative.
Next in this series: why an agent that reviews its own work approves nearly all of it, and what to do about it.
This is the second article in a series on harness engineering. Part 1 introduced the concept and the seven layers. Part 3 covers verification loops.