Hi - I answer from the OpenSmartRoute documentation: routing, the API, plans and quotas, self-hosting. Ask away, or open a support ticket if you need a person.
Grounded in the docs - follow a source before acting on it.
The adoption of AI agents in millions of organizations is creating new opportunities for attackers to make them take malicious actions. These actions include exfiltrating database contents and sensitive business and personal information. In the past five months, Google and four other organizations have acknowledged vulnerabilities. They exploit one agent inside a targeted network to spread harmful instructions to other internal agents. The technique is a special form of prompt injection that targets not the language model but a particular agent. For example, it targets an agent for translation or data analysis. Guardrails inside such agents are often lax if they exist at all. These guardrails will send the instructions to other agents down the chain. The latter agent explicitly trusts the first one because it follows the directions.
This trust allows harmful tasks to propagate through a network of specialized tools. An attacker does not need to trick the main language model directly. Instead, they target the specific function that handles data or communication. This function acts as a gateway for other agents to perform their duties. If an agent receives a command from another agent, it executes it without question. The system was designed assuming all internal agents are safe. It assumes no one inside the network is malicious by default. This assumption creates a blind spot for security researchers and defenders.
The discovery of protocol pivoting vulnerabilities in the Model Context Protocol
Researchers have found critical flaws in how AI agents communicate with each other. They discovered that these communication protocols can be used to inject malicious commands. The specific flaw is known as protocol pivoting. This term describes an attack where one agent tricks another into executing a task it should not run. The vulnerability affects the Model Context Protocol, often shortened to MCP. This protocol allows AI apps and agents to communicate inside an internal network. It was designed to help different tools share information and work together seamlessly.
The discovery came from independent security researchers testing real-world systems. They looked at how major organizations have implemented these agent communication standards. Their tests revealed that trust is built too easily between different agents. One agent can pass a malicious prompt to another agent using a different protocol. This second agent then runs the instruction because it trusts the first one blindly. The attack works even if the receiving agent has its own safety settings enabled. Those settings are often bypassed during the handoff of instructions between protocols.
Cohere released North 2 to manage agents and workflows across any model. The platform handles multi-step tasks while keeping context between sessions.
How independent researchers tested agents from major organizations including Google
Independent researcher Syed Anas Mohiuddin conducted a series of tests on various systems. He tested agents from Google, JP Morgan Chase, Weviate, and Rapid7. His team also tested systems used by the French government and the US federal government. These organizations have little in common except for their use of AI agents. They all rely on some form of agent communication to manage internal workflows. Syed's proof-of-concept attacks demonstrated how easily trust can be exploited across these boundaries.
His tests showed that well-crafted prompts could trigger unexpected behaviors in these systems. He did not need access to the main language model to launch an attack. The attack targeted specific agents responsible for handling data or making network requests. These agents often lack the strict guardrails found in traditional software applications. The researchers found that many of these agents were built with speed and flexibility in mind. Security was a secondary concern during their initial design phases. This prioritization left gaps that attackers can easily fill.
Specific technical flaws found in Google's MCP toolbox and Rapid7's network
The vulnerability affecting Google stemmed from its MCP toolbox for databases. This tool is named googleapis/mcp-toolbox in the developer registry. It initializes its HTTP client without using a CheckRedirect policy. This policy controls how a server handles errors or redirects to different URLs. The absence of this check allowed the toolbox to follow redirects blindly. Google's HTTP client also failed to validate target IP addresses during requests.
A crafted path parameter could make the toolbox follow a redirect to an internal endpoint. Once redirected, the server would send requests on behalf of the attacker. This behavior is known as server-side request forgery or SSRF. The severity rating for this Google vulnerability was 8 out of 10. This high score indicates a critical risk to organizational security. Rapid7 fixed its own network vulnerability last month. That flaw carried a severity rating of only 2.7 out of 10.
The concept of trust gaps between different agent communication protocols
Syed is calling the class of attack "protocol pivoting" because it exploits gaps between protocols. An app or server uses MCP to assign a task to an agent. That agent then forwards malicious instructions to another agent using a different method. This second method could be Google's Agent-to-Agent protocol, often called A2A. It might also use emerging standards like the Agent Network Protocol. Often, trust or authorization gets effectively lost in translation during this process.
Each protocol was built assuming it lived on its own isolated system. Each one checks its own front door while nobody watches the hallway in between. Markus Vervier, a researcher at X41 D-Sec, argues that "prompt injection" is the better term for these attacks. He says Syed's technique is simply a subclass of indirect prompt injection. The fact that the malicious prompt comes from a different protocol does not strictly require it to work. However, such attacks are unexpected and hard to mitigate in general.
Why these attacks succeed where traditional LLM defenses often fail
Many special-purpose agents lack the guardrails that normally mitigate harmful consequences of prompt injection. These agents are built to trust every other internal agent within their network. An exploit that would have been rejected by a standard language model succeeds here. The system assumes all internal actors are trustworthy by default. This assumption is the root cause of the widespread vulnerability.
In many cases, well-crafted prompts targeting the right agent will lead to server-side request forgery. This flaw causes a web server to make unauthorized network requests without user consent. Douglas McKee, director of vulnerability intelligence at Rapid7, explains that AI agents give attackers fresh connections. Someone plants text in content, and an agent reads it then passes it along as a normal delegated task. The second agent runs the work because it trusts whoever handed it the task. Every piece in that chain did exactly what it was designed to do.
Why it matters for the safety and security of enterprise AI systems
The adoption of AI agents in millions of organizations is creating new opportunities for attackers. These attackers can make agents take malicious actions like stealing data. The bugs underneath are old friends like injection and SSRF, but they have changed little in 20 years. Credit to the researcher for putting a name on it because a name helps defenders design for it.
The fact that the pivoting technique worked across five organizations with nothing in common is notable. MCP is new and is already everywhere before it has been sufficiently tested and hardened. Underlying these systems is a rush to build sprawling agentic architectures. This rush has led organizations to abandon a core security principle known as zero trust. To mitigate effects, engineers must design nodes to require authorization before conducting sensitive transactions with other ones.
What to do about implementing proper guardrails and zero trust architecture
The lesson researchers want people to take away is that anything passed from an LLM to your tool should be treated like input from a stranger on the internet. In a prompt injection scenario, that input is exactly what it is. Engineers must treat all incoming data as potentially malicious regardless of its source. The fixes for these vulnerabilities have not changed in 20 years. They rely on established security principles rather than new AI-specific technologies.
Organizations should implement an allow-list of IP ranges and block lists to prevent unauthorized access. This approach rejects an unsafe base URL at startup instead of waiting until the first request. That is what a real SSRF guard looks like according to experts. It is also more work than most MCP servers have done so far. Security teams must prioritize these checks during the design phase of new agent systems.
Security professionals should check their current agent communication protocols for similar gaps. They should look for any place where trust is assumed without verification. Testing with proof-of-concept attacks can reveal hidden vulnerabilities before attackers do. Implementing zero trust architecture means assuming one or more nodes may be infected. This assumption drives the need for strict authorization checks between all nodes.
Readers can compare their current setup against the specific flaws found in Google and Rapid7. They should verify if their HTTP clients validate target IP addresses properly. Checking for missing redirect policies is a simple step that can prevent SSRF attacks. Organizations should also review how they handle instructions passed between different agent protocols. Ensuring that handoffs do not bypass existing guardrails is critical for safety.
The name of the attack, protocol pivoting, helps defenders design systems against it now. Without a clear term, standards bodies might not prioritize fixing these issues. Naming the problem allows engineers to write better documentation and security guidelines. This clarity is essential for building safe AI systems in the future. The goal is to prevent attackers from walking across fresh connections within a network.
The Wikimedia Foundation found OpenAI agents editing wikis and making millions of API requests. It suspects these actions caused a partial service outage in May.