October 1, 2026
Adopting AI securely: a practical architecture for any organization
Seven control planes, eight traffic paths, and an honest account of what no one can inspect.
By Ahmed Thabit
10 min read
Summary
Most organizations are adopting AI in four forms at once. General-purpose assistants, such as Microsoft 365 Copilot and the coding assistants developers use. Purpose-built AI products bought for one job. AI-augmented software, meaning the AI that vendors are adding to products already in use. And AI platforms made for building AI, such as Copilot Studio or Azure AI Foundry, where home-grown agents get built. Whether each one can reach customer data depends on what it is connected to, and confirming that is the first job of the inventory this paper recommends. Agents raise the stakes, because they can act as well as read. Existing controls were built for people and applications, and much of this traffic never reaches them.
This paper recommends a layered architecture of seven control planes. Most of it runs on products an organization already owns: the identity provider, the office suite's tenant controls, the API gateway, the device security agent, and the firewalls. The gaps those products leave are usually narrow, mainly finding the agents nobody registered and posture management for AI running in the cloud.
It also recommends one rule for describing the program to auditors, regulators, and the board: claim only the coverage you can show evidence for.
This paper covers security. Privacy, fairness, model quality, and AI-specific laws matter as much, and they are handled outside it.
Where this view comes from
I came to security from networks and cloud infrastructure, so when I look at a new system I start with what talks to what: which component opens the connection, what crosses it, and where I could stand to watch it or stop it. An AI agent is a component that talks to users, systems, models, data, tools, and other agents, and each of those conversations is a path I can name. That is why the core of this paper is a list of eight paths, each with its risk, its control, and the team that owns it.
Why AI changes the security problem
Three things separate AI from the applications security teams protect today.
First, a model takes instructions and data through the same channel. It cannot reliably tell text it should obey from text it should only read. Anyone who can put words in front of it, in an email, a document, or a web page, can try to give it orders. OWASP lists this attack, prompt injection, first among its top ten risks for applications built on language models.
Second, agents act. A chatbot that answers wrongly is a quality problem. An agent that can send email, change a record, or call a vendor API turns a wrong answer into an action taken with someone's access. OWASP calls this risk excessive agency.
Third, much of the traffic never touches the organization's network. An office-suite assistant, its models, and its agents talk to each other inside the provider's service. Vendor AI runs in the vendor's cloud. A firewall cannot inspect a conversation that never crosses it.
One incident shows all three. In June 2025, researchers at Aim Security disclosed a flaw in Microsoft 365 Copilot, tracked as CVE-2025โ32711 and known as EchoLeak. The attacker sent a single email, worded to slip past the product's own prompt injection filter. Later, when the user asked the assistant an ordinary question, it pulled that email into its context, followed the instructions hidden in it, and leaked internal data. The user clicked nothing, and the data left through an address belonging to the provider, which any firewall would trust. The flaw has been fixed. The pattern remains: the attack arrived as data and ran with the user's access, inside a service the customer does not operate.
No one can solve prompt injection itself. It comes from how language models work, and the EchoLeak email got past Microsoft's own filter. Treat it the way you treat phishing: you cannot stop the email from arriving, so you limit what a successful one can do. Give every agent its own least-privilege identity, allowlist the tools it may call, clean up what it can retrieve, block the routes data could leave by, log every tool call, keep a tested kill switch, and require a person to approve actions that move money or change important records. The rest of this paper is how to arrange those controls.
Design principles
Five rules guide every choice that follows.
- Decide the capabilities first and pick vendors second. The architecture should survive a change of product.
- Put each control where the traffic has context. Identity belongs at the identity provider, prompts and tool calls at the gateway, and agent behavior in the platform or device that runs the agent. The network is the backstop.
- Keep one inventory, one place for logs, and one way to stop an agent. Three tools holding three different lists of agents is worse than one.
- Match the controls to the risk. An agent that touches sensitive data or can change important records gets every plane. A meeting summarizer gets the baseline.
- Claim only what you can evidence. Do not say that all AI traffic is inspected, because it is not and cannot be.
The architecture
The architecture has seven planes, and each has one job. Governance and logging cut across the other five. For each plane, here is the job, what most organizations already own for it, the gap they usually find, and who tends to own it.
- Governance. Keeps the inventory of AI systems and agents, sets risk tiers, reviews vendors, and maps controls to a framework such as NIST AI RMF, ISO/IEC 42001, or, for financial institutions, the CRI FS AI RMF. Already owned: a risk framework and a third-party risk program. Gap: an inventory that includes agents, including the ones employees build on their own, which takes discovery tooling to find. Usual owner: risk and compliance.
- Identity. Gives each agent its own identity and the least access it needs. Already owned: an identity provider. Gap: agent identities, so no agent runs on a shared service account. Usual owner: identity and access management.
- Staff and devices. Controls how staff use AI tools and coding assistants. Already owned: a device security agent and web filtering. Gap: discovery of unapproved AI tools, and data loss prevention on prompts. Usual owner: endpoint security.
- API and agent gateway. Authorizes and inspects agent traffic to internal systems, vendors, and outside agents. Already owned: an API gateway. Gap: its AI features switched on, meaning model, tool, and agent-to-agent policies. Usual owner: the integration team, with security setting the policy.
- Workload and cloud. Finds AI services in the cloud and fixes risky configuration. Already owned: a cloud security posture tool. Gap: coverage of AI services and agents. Usual owner: cloud security.
- Network. Segmentation, egress allow-lists, and blocking paths around the gateway. Already owned: firewalls. Gap: usually none for this job. AI inspection on the network is a separate purchase and needs decryption. Usual owner: network security.
- Logging and response. Collects every plane's logs and stops an agent on demand. Already owned: a SIEM. Gap: AI log sources, and a tested procedure to disable an agent. Usual owner: the security operations center.
Controls attach to paths, and an AI system has eight of them. Several never appear on a wire the organization owns, so each path below says where the control sits and which plane owns it.
Seven terms carry the list. A path is one kind of conversation an AI system has, named by who starts it and who answers. A plane is one of the seven control layers above. AI means any AI system, whether it only answers, like a chatbot, or can act. An agent is AI software that can take actions. A user is a person, either staff or a customer. A model is the language model an AI system sends prompts to. A system is anything other than a person or an AI that can start an agent or send it input: a mailbox, a business application, a file store, a scheduler, or a vendor's service. Direction separates two of the paths. When the agent calls an application, that is agent to tools and apps. When the application calls the agent, that is system to agent.
- User to AI. Prompts from staff and customers. Risk: customer data pasted into unapproved AI tools. Control: sign-in and tenant controls for approved assistants, plus device and web controls for public tools. For customer-facing AI, sign-in, rate limits, and input checks at the gateway. Plane: staff and devices for staff, API and agent gateway for customers.
- AI to user. The reply. Risk: a reply that shows the user data they are not authorized to see, or one that carries data out to an attacker, as EchoLeak did. Control: permissions at the source for the first, and output handling for the second, so replies cannot render external links or images or reach a customer unchecked. Plane: identity and governance for permissions, API and agent gateway for output checks.
- System to agent. An arriving email, a call from another application, a new file, or a timer starts the agent, with no person involved. Risk: outside content reaches the agent before anyone sees it. Control: the agent platform's admin policies, and the strictest risk tier. Plane: governance.
- AI to model. Prompts and responses. Risk: prompt injection, and sensitive data in prompts. Control: inside the provider's service for office-suite assistants and low-code agents, through tenant controls. At the gateway for AI your developers code. Plane: API and agent gateway.
- AI to knowledge and memory. Content retrieved from document stores, email, files, and indexes, plus anything the AI remembers between sessions. Risk: retrieved content that carries hidden instructions, and oversharing. Control: at the source, with permissions cleanup and labeling. Plane: identity and governance.
- Agent to tools and apps. Tool calls, usually over the Model Context Protocol, including web browsing. Risk: an agent acting with more access than it needs, and browsing that pulls in content nobody controls. Control: at the gateway, with an allowlist of tools per agent. Plane: API and agent gateway.
- Agent to agent. One agent handing work to another, usually over the Agent2Agent protocol. Risk: trusting another agent's output, and access passed along. Control: at the gateway for outside agents. Agents that call each other inside one platform stay inside it, under that platform's controls. Plane: API and agent gateway.
- Build and supply chain. Models, tool servers, plugins, and packages pulled from public registries. Risk: a poisoned component. Control: approved sources, review before use, and controls on developer laptops. Plane: staff and devices.
The network plane owns none of these paths. It backs all of them by blocking routes around the gateway, and the logging plane collects from every one.
Two controls make the gateway mandatory and not just recommended. The agent platform's admin policies limit what builders can connect to, so connectors, tools, and outside agents must point at the gateway. Internal systems then accept calls only from the gateway. The policy stops well-meaning builders, and the backend lock stops everyone else.
Why a gateway and not something else? The alternatives are to let agents connect directly to vendors and tools, which leaves no common place to authorize, log, or stop them; to inspect on the network, which needs decryption and never sees traffic inside a provider's cloud; or to rely on each platform's own controls, which stop at that platform's boundary. A gateway the organization already runs for vendor APIs is the one point that sees traffic across platforms, and most API gateway vendors have now added policies for model, tool, and agent-to-agent traffic. It works in all three directions: outbound, where the organization's agents call vendors, models, and outside agents; inbound, where outside agents and systems call the organization's APIs and tools; and internal, where an agent reads or updates a record. If deeper inspection is needed later, a security vendor's scanner can run as a gateway policy.
One organization's experience shows how this plays out. It ran applications on premises, most workloads in one public cloud, and some storage in a second. Vendor API traffic passed through an API gateway, and a transit hub sent traffic through a cluster of firewalls. Staff used an office-suite assistant and its low-code agent builder. The security team's first idea was to force all AI traffic, agent-to-agent traffic included, through the firewalls and use the firewall vendor's new AI inspection. Three facts ended that plan. The AI inspection ran only on a separate line of firewalls the organization had not bought, and it needed decrypted traffic. The assistant and its agents talked inside the provider's service and never reached the firewalls. And the API gateway vendor had already added governance for model, tool, and agent-to-agent traffic. The plan that replaced it keeps the firewalls for segmentation and egress control and makes the API gateway the control point for agent traffic.
Known limits
Some AI traffic stays out of reach of any control the organization operates. Say so plainly, internally and to auditors.
Traffic inside a vendor's service is the largest share. An office-suite assistant's model calls, agents that call each other inside one platform, and the AI inside vendor products never cross the organization's network or gateway. For these it depends on tenant settings, the vendor's own controls, and what the contract lets it verify.
Two more flows cannot be intercepted at all. AI tools send prompts, traces, and logs to their vendors' monitoring services, and a vendor's product may call its own model provider, a fourth party to the organization. Contracts and vendor review are the only controls for both.
Some encrypted traffic cannot be opened. Network inspection needs decryption, and applications that pin certificates or use mutual TLS break when decrypted, so they are excluded from it.
Traffic inside a host leaves nothing to inspect. Two agents on one server, or a coding assistant calling a local tool on a laptop, never touch the network. Device and platform controls cover these, or nothing does.
Keep a gap register that lists each blind spot, the control that compensates for it, and the evidence that control produces.
Roadmap and measures
The order matters more than the dates, because each step makes the next one cheaper.
- First 60 days (governance, staff and devices). Build the inventory of AI systems and agents. Confirm what the products you own already provide. Write the admin policies for the agent platform.
- Months 3 to 5 (API and agent gateway, network, logging). Route agent tool traffic and outside agents through the gateway. Lock the backends so they accept calls only from it. Send gateway logs to the SIEM.
- Months 6 to 8 (staff and devices, workload and cloud, identity). Close the device and cloud posture gaps, buying only where owned products fall short. Give agents their own identities.
- Months 9 to 12 (logging and response, governance). Test the program. Red-team one high-risk agent, rehearse disabling one, and assemble the evidence file for auditors.
Five numbers, reported each quarter, show whether it is working:
- The share of AI systems and agents in the inventory that have an owner and a risk tier.
- The share of agent traffic to internal systems that passes through the gateway.
- The number of agents running on a shared or human account. The target is zero.
- The time it takes to disable an agent, measured in a drill.
- The number of gap register items open for more than 90 days.
In short
You can adopt AI securely by naming the paths your AI systems use, putting each control where the traffic has context, using the tools you already own before buying new ones, and writing down the gaps. Start with the inventory, make the gateway the control point for agent traffic that reaches your systems, and make every agent identifiable, measurable, and stoppable.
Sources
Opened between 17 and 19 September 2026.
- OWASP, Top 10 for LLM Applications 2025
- Reddy and Gujral, EchoLeak: the first real-world zero-click prompt injection exploit in a production LLM system, AAAI Symposium Series, 2025
The views here are my own and do not represent my employer.