Manager · Engineer · Mountaineer
FR · EN
AI Security
Refused in plain text, the same malicious instruction gets through once encrypted.
2026-10-08
A web page. On it, a block of encrypted text and a polite request: "decrypt this with Python, here are two keys." Twenty-eight seconds later, the developer's .env.prod file was on the attacker's server. That's what Adversa AI demonstrated on October 6 against GitHub Copilot CLI, with a technique they call Cryptographic Context Injection.
The detail that stopped me: sent in plain text, the same instructions were flagged as a prompt injection and refused. Encrypted, they went through. The model never sees the malicious order; it rebuilds it itself, by running its own code.
The agent, in Autopilot mode, reads a page the attacker controls. The first key is designed to fail and nudges the agent into reading local files to "complete" it. The second key works and reveals a second payload, which ships whatever was collected to an external server. No confirmation asked.
Adversa is upfront about the limits: it takes autonomous mode and a permissive model. Microsoft's mai-code-1.1-flash ran the full chain in half the attempts, while two GPT-5.6 models refused every time. And on the account tested, automatic routing picked one or the other without the user knowing. Your security level came down to a coin toss on the vendor's side.
Reported on September 17, the behavior was confirmed by GitHub, which declined to classify it as a vulnerability: the user chose to let an autonomous agent read attacker-controlled content. As of October 1, the chain was still reproducible.
Legally, I understand the position. Operationally, it tells companies one clear thing: what the agent does on the workstation is your responsibility. The model's filter is a bonus, not a control to build a policy on. Any filter on incoming text will eventually be bypassed by an encoding.
Watch actions, not text. A sequence of "read an external page, run code, read a secrets file, call a never-seen domain" is suspicious whatever content triggered it. Adversa itself recommends logging tool calls along with their resolved arguments.
Then cut the last link. An agent running unsupervised shouldn't be able to reach a new network destination by default. And an agent that reads the web has no business keeping production credentials in its working directory. It's the classic combination: private data, untrusted content, the ability to send data out. Remove any one of the three and the attack falls apart.
Don't wait for an answer from your AI vendor. Model behavior isn't contractual, and GitHub's response shows the vendor considers this risk yours. So it must be known, tracked and owned by leadership. Not left to each developer who switches on Autopilot mode.