G8KEPR
Agents That Turn on You: Hijacking and Memory Poisoning
Part 17 of the thread Building G8KEPR
- Agent hijacking is when an AI agent ends up working toward someone else's goal instead of yours.
- Memory poisoning plants something false or malicious in what an agent remembers, so it keeps doing damage later.
- G8KEPR's MCP pillar watches for hijacking signs like context poisoning, agent impersonation, exfiltration through an agent, and sleeper payloads.
- Its controls, role-based tool permissions, session tracking and rug-pull detection, are about limiting what a turned agent can do and noticing when one turns.
A chatbot that goes wrong says something it shouldn't. An agent that goes wrong does something it shouldn't. That's the whole difference, and it's why agents deserve their own post.
An agent is an AI model that's allowed to act: call tools, read and write data, take one step, look at the result, and decide the next step on its own. Most of the time, those tools reach the model through MCP, the Model Context Protocol, which I explained in the post on tool poisoning.
Agent hijacking
Hijacking is when an agent stops working for you and starts working for someone else, usually without anyone noticing right away. Nobody breaks in. The agent is persuaded.
It usually starts with the same weakness covered in the prompt injection post: a model can't reliably tell your instructions from instructions buried in something it reads. With an agent, the consequences are actions, not just words. A few forms G8KEPR's MCP pillar watches for, in plain terms:
- Context poisoning. Something the agent reads fills its working context with misleading material, so its next decisions are built on a bad foundation.
- Agent impersonation. In systems where agents talk to other agents, one pretends to be something it isn't, so the others trust it more than they should.
- Exfiltration through an agent. The agent itself becomes the way data gets out, because it has legitimate access and a legitimate way to send things.
- Sleeper payloads. Instructions that sit quietly until some later moment or trigger, so the harmful part happens long after the content that caused it arrived.
That last one is the nastiest, because it breaks the link between cause and effect. By the time something goes wrong, the thing that set it up is ancient history.
Memory poisoning
A lot of agents now remember things. Preferences, past conversations, facts they learned along the way. That's genuinely useful. It also means a single bad input can outlive the conversation it came in with.
Memory poisoning is planting something false or malicious in what the agent remembers. A made-up "fact" about how a process works, say. Once it's in memory, the agent pulls it back out next week and acts on it as if it were true, and there's no suspicious message in that later session to point at.
This is also why G8KEPR's Verification Engine watches for problems that build up across a conversation, poisoned memory included, not just problems inside one answer. I touched on that in Checking the Answer, Not Just the Question.
What the MCP-side controls are for
You can't make a model immune to persuasion. What you can do is limit what a persuaded agent is able to do, and notice when behavior shifts. That's how I think about the three MCP-side controls.
Role-based tool permissions. An agent that summarizes documents doesn't need a tool that sends email. One that answers customer questions doesn't need write access to the billing system. Permissions on tools, by role, mean that when an agent does get turned, the blast radius is whatever it was allowed to touch, not everything connected to it. It's least privilege, the oldest idea in security, applied to a new kind of user.
Session tracking. Hijacking rarely happens in one call. It's a read here, a lookup there, a send at the end. Tracking at the session level lets steps that belong together be connected, which is what makes multi-step attacks visible at all. It also feeds the correlation engine, which I wrote about in The Lethal Trifecta and Other Attacks That Take More Than One Step.
Rug-pull detection. A rug pull is a tool that describes itself one way when you approve it and changes that description later. G8KEPR notices when a tool's definition changes after it's registered, so a tool can't quietly start whispering new instructions to your agent. The fingerprint post covers this in more depth.
The mindset
The more useful agents get, the more they look like employees with access, and the less they look like software you can simply trust. We don't give a new hire the keys to everything on day one, and we'd want to know if someone started behaving very differently from last week.
Agents deserve the same treatment. Give them only the tools they need, watch what they do over a whole session, and don't assume the tools they're talking to are still the ones you approved. If you want to see how that fits the rest of the product, g8kepr.com has the overview.