This follows an earlier piece I wrote on how AI agents slip past existing controls[i]. That article asked whether your network and resources can even tell an agent apart from a human or a service, and what controls they need against an agent’s non-deterministic actions. Nearly every question I got back landed on two things: how does an agent make data exfiltration risk worse, and what does an optimal egress control look like. So, this piece focuses solely on egress.
An AI agent is an outbound machine. It takes a task, reasons about it, then reaches out: to a model API, a vector store, a SaaS tool, a Model Context Protocol server, another team’s agent, often an endpoint on the public internet. One request can fan out into dozens of calls the agent chose on its own, based on text it was reasoning over. Your network was not built for a caller that picks its own destinations at runtime.
Environments barely put sufficient checks on outbound traffic
In many environments there is little egress control, because outbound was never where the perceived risk lived. Ingress got the web application firewalls, the rate limits, the scrutiny. Outbound came from software the team wrote, so it was trusted by default and allowed to leave.
The controls that do exist are rarely thorough: a firewall sits on one path while a second gateway, a peered network, a managed service endpoint, or a developer’s convenience route stays uninspected. It looks governed on a diagram and is porous in practice. That is what an agent lands in: not a tight allowlist it must subvert, but a broad outbound path that is often patchily filtered.
Allowlists must now assume software improvises
For teams that have done the work, outbound policy rests on one assumption: the workload is predictable. A service talks to a known set of dependencies. You enumerate IP ranges, ports and domain names, deny the rest, and it holds, because the software does the same thing every run. That was sound until agents arrived.
Agents break that assumption. An agent picks which tool or model to call based on reasoning shaped by inputs you do not control, so the endpoints it might reach are not known when you write the policy. A static list ends up too tight, and the agent fails, or too loose, and it becomes a general-purpose outbound channel. Making the list longer does not resolve that.
Two problems compound this. Sensitive data usually rides inside ordinary application content, so exfiltration looks like a well-formed API call to a legitimate domain, the damage sitting in the words rather than the connection. And agent-to-agent and agent-to-tool calls create lateral traffic between workloads that previously had no reason to talk, extending a compromised agent’s reach.
The egress path leads to exfiltration
The most consequential agent attacks disclosed so far are, at their core, egress attacks. The attacker never breaks in. They get a trusted agent to send data out.
When Legit Security disclosed CamoLeak in GitHub Copilot Chat in October 2025, hidden instructions in a repository persuaded the assistant to push private source code and secrets out through a channel the network permitted[ii]. When Zenity Labs demonstrated AgentFlayer at Black Hat USA 2025, one poisoned document made enterprise assistants rummage through connected systems and hand credentials back out[iii]. In both cases the credential was valid and the destination reachable. What was missing was not authentication but any governance of where the agent’s traffic could go and what it could carry.
Start by finding and shrinking the exits
The first move is not policy authoring. It is inventory. You cannot inspect traffic leaving through a door you have not found, and there are usually more doors than anyone believes.
Enumerate every path traffic can leave by: internet and egress gateways, proxies, peered and transit connections into porous networks, VPN tunnels, managed-service endpoints, DNS paths and any developer-provisioned shortcut added for convenience and never removed. Then reduce that set deliberately onto a small number of controlled exits. Every path you remove is one less to monitor.
This matters more with agents than with conventional workloads. Paths proliferate as workloads multiply, and at scale inspecting every route individually stops being feasible. A compromised agent with excess privileges can influence routing, try alternate paths, and reach an exit nobody thought to cover. If one uninspected route exists, assume the agent finds it. The objective is structural: find every exit, minimize them, inspect what remains, so nothing leaves without passing a control.
Constrain destinations, per source
With the exits consolidated and inspected, the next control is what they permit. The baseline is default-deny: nothing leaves unless a rule explicitly allows it. That inverts the burden of proof. Instead of enumerating what is forbidden, an unbounded and losing exercise, you enumerate what is permitted; everything else fails by construction.
Default-deny only pays off when the allowlist is source-based. One global list applied uniformly at an exit means every workload behind it inherits the union of everyone’s permissions, and a single compromised workload’s blast radius becomes the entire allowlist. A source-based rule binds a specific source to specific destinations, so what one workload may reach does not open that destination for its neighbor.
Enforcing that list does not require reading traffic. The destination hostname is usually visible in the TLS handshake before the session is encrypted, and resolvers are a control point in their own right, since an agent cannot connect to a name they refuse to answer. Both have limits. Encrypted Client Hello conceals that hostname, so policy needs an IP and reputation fallback[iv]. And DNS controls nothing unless you also block outbound DoH and DoT, or the agent resolves elsewhere; that exact bypass has already defeated a CI egress allowlist[v]. Enforcement location matters more than any single observable.
For agents, source means identity
Destination lists alone leave the hardest question unanswered: which agent is this. The policy you want is not “can this host reach that address,” but “may this specific agent, owned by this team, acting for this user, reach this destination.”
That requires the agent’s identity and provenance to travel with the request to the egress boundary, the recognition problem the first article argued must be solved first. Provenance is the precondition for policy. Without it, every agent behind a shared exit is indistinguishable and source-based rules collapse into one network-wide list.
The model to aim for is an agent that starts with no outbound reach at all, with every destination it needs declared, reviewed and granted in its own definition alongside its identity and owner, not retrofitted into a firewall rule afterwards. A new endpoint later is a change to approve, not a runtime decision it makes for itself.
Denials must be explicit and legible: a refused call should return a structured rejection the agent can reason about and fall back from, because a denial that looks like a timeout invites retries and probing, exactly what you do not want from a system that improvises. And deny-by-default holds when the agent is hijacked, because the boundary never consults the agent’s judgment about where it may go. That is why this belongs at the network, not inside the agent.
Decrypt selectively, based on destination trust
Destination policy governs where traffic goes. It cannot tell you what that traffic carries, and since sensitive data rides inside ordinary application content, that gap is where exfiltration happens. The temptation is to intercept TLS everywhere. That is the wrong answer. Plenty of destinations will not accept an intermediate proxy’s certificate and simply break, decrypting at volume is expensive, and terminating all TLS gives the firewall plaintext access to everything, a fresh high-value target with new compliance obligations. Done indiscriminately, it solves a visibility problem by manufacturing a custody problem.
The opposite extreme is wrong too. Ideally you would permit only known, trusted destinations. But when enterprises open a path for general workforce or workload traffic, they usually deploy the inverse: allow everything except what is already known to be malicious. Traffic then splits into two populations treated identically: destinations you have vetted, and destinations whose posture is unknown. A steerable agent will happily use the second. Bounded agents avoid this through the source-based allowlist, but some legitimately need broad reach, browsing or calling third-party APIs chosen at runtime, so the unknown population is unavoidable.
So tier the control by destination trust. For trusted, vetted destinations, do not decrypt; you already know the endpoint. For unknown or uncategorized ones, decrypt selectively: outbound, run data-leak prevention so sensitive content cannot reach an endpoint nobody has evaluated; inbound, scan what returns, because content from an unvetted source is where malware and poisoned instructions enter, and an agent that ingests it will act on it. That confines cost and breakage to the traffic that warrants scrutiny and concentrates inspection where your knowledge is weakest.
Everything so far concerns traffic leaving your environment. The east-west direction, agents calling other agents and tools, needs the same discipline applied to internal reach, where least privilege does the work. That is what checks lateral movement by a compromised agent.
An agent should reach only the internal endpoints its task requires. A customer-support agent has no business reaching a code repository, and a build agent has no business reaching a payments API. Scope each agent to what its job demands, express that once in its definition, and enforce it at the boundary rather than trusting the agent to behave. On a flat internal network the opposite holds by default: any agent reaches almost anything, so one hijacked agent inherits the whole environment’s reach.
Treat these calls as trust-boundary crossings, not free movement: mutual authentication, inspection out and back, and segmentation so the internal endpoints an agent can reach are as deliberately chosen as the external ones. The return path matters as much as the request: a response from a compromised agent or tool is another way poisoned instructions arrive. The protocol layer is young too: the default MCP transport shipped at the end of 2024 with insufficient authentication, so do not assume these hops are safe[vi].
Attribute every outbound flow to an agent
Finally, make the traffic answerable. Every outbound flow should be attributable to a specific agent, its owner, and the human it acted for, and queryable afterward. “Which external destinations did any agent reach this month, and which agent reached them” should be one question with one answer at the boundary, not a forensics exercise after something has gone wrong. Where you do inspect payloads, pair the flow record with what was sent.
Egress is where recognition becomes control
Knowing an agent is an agent is necessary but inert on its own. That knowledge becomes an enforceable limit on the outbound path, the last point where a decision can be made that the agent cannot override. Corrupt the model, poison its inputs, talk it into a plan no designer imagined, and the data still has to leave through an exit you control, provided there is no other way out.
The agents are already calling out. The work is making sure something is deciding, on your behalf, which of those calls should ever leave the environment.
Sources
[i] HackerNoon, The Impostor in Your Environment Is the AI Agent Holding a Valid Credential.
[ii] Legit Security, CamoLeak: Critical GitHub Copilot Vulnerability Leaks Private Source Code.
[iii] Zenity Labs via PR Newswire, Zenity Labs Exposes Widespread “AgentFlayer” Vulnerabilities Allowing Silent Hijacking of Major Enterprise AI Agents Circumventing Human Oversight.
[iv] Cisco, Encrypted Client Hello (ECH) Defense Strategies: How Cisco Secure Firewall Tackles ECH.
[v] GitHub Advisory Database, Egress Policy Bypass via DNS over HTTPS (DoH) in Harden-Runner (Community Tier).
[vi] Vectara, MCP’s Rapid Journey: From Open Door to a Fortified Gateway