TL;DR: agent harness engineering and its hidden attack surface
Agent harness engineering is the practice of building everything around an AI model (instructions, tools, permissions, memory and feedback loops) that turns it into a working agent. That harness, not the model, is where most of the real security risk now lives.
- The shift: teams stopped tuning prompts and started engineering harnesses. Agent = model + harness.
- The problem: the harness lives in plain-text files (rules, skills, MCP configs) that get reviewed like documentation, if at all.
- The evidence: the Rules File Backdoor and real MCP CVEs show attackers already target it.
- The fix: treat the harness as code. Inventory it, review it and scan it before it ships.
Every security conversation about AI eventually lands on the model. Can it be jailbroken? Will it hallucinate? Is the provider training on our data?
Fair questions. But they are increasingly the wrong ones. When an AI agent deletes a production table, pushes a backdoor or emails a customer list to a stranger, the model is rarely what failed. The model did exactly what its harness allowed it to do.
Agent = model + harness
Over the past year, the industry quietly changed what it means to build with AI. Prompt engineering gave way to context engineering, and context engineering gave way to something bigger: agent harness engineering.
The idea is simple. A model on its own is a reasoning engine with no hands. The harness is everything that gives it hands and a job: the instructions it follows, the tools it can call, the permissions it holds, the memory it keeps, and the checks that catch its mistakes. Martin Fowler’s site describes harness engineering as the work coding-agent users now do to make agents reliable, and an academic framework published this year breaks the harness into eleven responsibilities, from tool access and project memory to permissions and verification.
The results are real. Teams practicing agent harness engineering report that changing only the harness, with the same model underneath, can shift an agent’s results dramatically. This raises an uncomfortable point for anyone in security: if the harness decides how the agent behaves, the harness decides how it can be abused.
Take something as ordinary as installing a dependency. A developer pauses to check the package name. An agent with the right tool in its harness just runs the command, and if the package is malicious, nothing in the model will stop it. That’s exactly what we unpack in this episode of SafeDev Talks: what changes when AI agents install dependencies on their own, and why agent harness engineering has quietly become a supply chain problem.
What agent harness engineering looks like in a repository
Open a repository where developers use AI agents and the harness is right there, in files you have probably never scanned:
- Rules and instruction files (
AGENTS.md,.cursorrules, Copilot instructions) that tell the agent how to behave. - Skill files that package reusable capabilities the agent loads on demand.
- MCP server configurations that connect the agent to tools, databases and APIs.
- System prompts and prompt templates embedded in application code.
- Agent wiring: which tools an agent can reach, which data it can retrieve, and whether a guardrail sits in between.
All of it is plain text. All of it is committed alongside code. And all of it gets reviewed the way documentation gets reviewed: a quick glance, an approval, a merge. Yet each of these files can silently rewrite what an agent is instructed to do and what it is allowed to touch.
That is the real consequence of agent harness engineering: the most powerful configuration in your software is now the least reviewed.
Attackers already noticed
This is not a theoretical concern. The harness has a short but telling incident history.
- The Rules File Backdoor. In March 2025, researchers disclosed an attack where hidden Unicode characters inside a rules file instructed AI coding assistants to insert backdoored code, without mentioning it in their visible response. A reviewer reading the file saw nothing unusual. It is now catalogued as MITRE ATLAS case study AML.CS0041.
- Poisoned MCP tools. Through 2025, researchers repeatedly showed that agents trust tool descriptions implicitly, so a malicious MCP server can steer an agent simply by describing itself the right way. The same year, CVE-2025-6514 in a widely downloaded MCP client allowed remote command execution when connecting to an untrusted server, rated 9.6 on CVSS.
- Excessive agency. The OWASP Top 10 for LLM Applications lists excessive agency as a top risk: agents granted more tools, permissions, or autonomy than the task requires. That is not a model flaw. It is a harness design decision.
Notice the pattern. None of these attacks break the model. They break the harness, and the model faithfully does the rest.
Why your security stack doesn’t see it
Here is the awkward part. Most organisations already run good AppSec tooling, and almost none of it was built for this layer.
Static analysis understands code, but it does not know what a model is or why untrusted text flowing into a system prompt matters. Composition analysis inventories packages, but it does not enumerate MCP servers or skill files. Secret scanners may miss an API key sitting in an agent configuration file. None of these tools is defective. They were simply built for a world where configuration did not give orders.
So agent harness engineering moves fast, and security reviews the part it can see: the model and the code. The harness falls in between.
Five habits for securing the harness
Securing agent harness engineering does not need a new team. It needs a few habits that treat the harness as what it is: executable intent.
| Habit | Why it matters | What to do |
|---|---|---|
| 1. Inventory every harness | You cannot secure agents nobody declared. | Find every model, agent, MCP server, skill and prompt across your repositories, from the code and config files themselves rather than from a survey. |
| 2. Review harness files like code | Rules, skills and MCP configs can rewrite what an agent does, yet they get reviewed like documentation. | Give them the same pull request scrutiny as application logic, including checks for hidden characters. |
| 3. Pin what the agent loads | A prompt or model referenced by a mutable label can be repointed without a code change. | Pin prompts and models to immutable versions. |
| 4. Put a guardrail on every sink | Retrieved documents, tool output and user content can all carry injected instructions to the model. | Place a guardrail in the path wherever that content can reach the model. |
| 5. Keep credentials out of the harness | AI provider keys in prompt files and agent configs are an easy win for an attacker. | Remove them and keep scanning harness files for secrets. It is an easy fix. |
Securing the harness, not just the model
Xygeni AI Security starts where agent harness engineering leaves its footprint: in your repositories. It discovers the AI assets across your codebase (models, agents, MCP servers, skills, prompts, guardrails, and the AI coding tools in use) from code, dependencies, and the configuration files those tools leave behind.
It then analyses skill files, rules files, and MCP configurations as security artifacts, not documentation, and detects prompt-injection-class risks in agent wiring, such as untrusted content reaching a system prompt or a retrieval sink with no guardrail. Xygeni’s secrets scanning flags AI provider credentials in prompts and agent configuration files, and the dependencies of your AI stack get the same malware detection as everything else, before a signature exists.
FAQ
What is agent harness engineering?
Agent harness engineering is the discipline of designing the environment around an AI model (instructions, tools, permissions, memory, context, and verification loops) so that it behaves as a reliable agent. The model provides reasoning; the harness decides what the agent can see and do.
How is agent harness engineering different from prompt engineering?
Prompt engineering shapes a single instruction. Agent harness engineering shapes the whole system the agent operates in, across many steps, tools, and sessions. A prompt is one file in the harness.
Why is the harness a security risk?
Because it controls what an agent is allowed to do, and it lives in plain-text files that are rarely reviewed as security artifacts. Compromise a rules file or an MCP configuration and you control the agent, without touching the model.
Who should own agent harness security?
Application security, working with the engineers who build and configure agents. The harness is part of the software you ship, so it belongs in the same review, scanning and inventory processes as the rest of your code.







