Your Agent Has a Middleman You Didn't Vet

Researchers bought 28 paid LLM API routers from Taobao, Xianyu, and Shopify-hosted storefronts and collected 400 free ones from public communities. On each they registered an account, ran a sandboxed coding agent through it, executed every tool call that came back, and watched what it did.1

Nine of the 428 rewrote a tool call into an attacker-controlled command or dependency: one paid router and eight free ones. Seventeen free routers went on to use an AWS canary credential after seeing it in transit, and one drained a planted Ethereum key. Two of the injecting routers hid the behavior. One waited for 50 requests before acting, and one fired only for sessions in autonomous “YOLO mode” working on Rust or Go projects.1

An LLM API router is an application-layer proxy. It terminates TLS, reads every request and response in plaintext, and can rewrite the tool call your agent is about to execute. The paper found no major provider signing its tool-call responses, so nothing ties the command your agent runs to what the model produced. The paper tested four agent frameworks, Claude Code and Codex among them, and none performed a response-integrity check.1

This post was first published on April 10, 2026 from the paper’s abstract. On October 2, 2026 I rewrote it against the paper’s full text. The April version got several things wrong, and they are listed under “What the April Version Got Wrong” near the end.

TL;DR

  • What a router can do. It sits between your client and the model provider as the endpoint you configured, so it holds the plaintext of every request and response and can change either before passing it on.1
  • What 428 routers did. Nine injected malicious code into returned tool calls (1 of 28 paid, 8 of 400 free). Two of the free ones used triggers to hide it. Seventeen free routers used a researcher-owned AWS canary credential, and one drained a researcher-owned Ethereum key holding under $50.1
  • What leaked keys and weak relays did. One deliberately leaked OpenAI key served 100M GPT-5.4 tokens and exposed more than seven Codex sessions. Weakly configured decoy relays served about 2B tokens and exposed 99 credentials across 440 Codex sessions, 401 of them already running with tool calls auto-approved.1
  • What the defenses did. In the authors’ own synthetic benchmark, a fail-closed policy gate blocked every injection sample at a 1.0% false-positive rate, and in a simple adaptive benchmark an attacker who knew about the gate bypassed it in 100% of samples. Anomaly screening flagged 89.0% of plain injections and about half of the evasive ones.1
  • What the authors say would fix it. A response envelope signed by the provider, so a client can check a tool call against what the model produced. The paper found no major provider shipping one.1
  • If you host the router yourself. The October 1 update below covers eleven LiteLLM advisories, one of which let any authenticated user make the proxy send out its provider keys.2

Key Takeaways

  • Agent operators: Every router between your client and the model provider has plaintext access to every request and response, and the client configures only the first hop. If you bought a router from a marketplace or pulled one from a public list, treat it as a hostile intermediary until you have a separate reason to trust the operator.
  • Harness builders: A PreToolUse hook runs before a tool call executes, and by then a router has already had its chance to rewrite that call. The hook has no original to compare against. What it can do is fail closed: allow shell commands that fetch only from listed domains and install only listed packages, and block the rest.3
  • Anyone running YOLO mode: In the researchers’ decoy study, 401 of 440 observed Codex sessions already ran with tool execution auto-approved.1 For those sessions a plain rewritten command would have executed with no trigger logic needed. Do not run auto-approved sessions through a router you do not control.
  • Teams hosting their own proxy: A router you run concentrates provider keys the same way. Eleven LiteLLM advisories that reached the PyPA database on October 1, 2026, most of them public since June, include one that let any authenticated user make the proxy send its provider keys to a host of their choosing. Releases from 1.97.0 on fall outside all eleven recorded ranges; the October 1 update below has the details.2

What’s a Router, Exactly?

An LLM API router accepts requests in one format, usually OpenAI-compatible, picks an upstream provider, and returns the response. The paper counts routers at every scale: cloud-managed services such as Amazon Bedrock and Azure OpenAI Service, developer-facing projects and services such as LiteLLM and OpenRouter, and a commodity market of resold and aggregated API access.1

People use them for sound reasons. The paper lists “model fallback, load balancing, cost optimization, and a single API key across providers,” and notes that routing is especially common “in regions where direct provider access is restricted, expensive, or subject to quota limitations.”1

The trouble is the position. No interception trick is needed, in the paper’s words: “the client voluntarily configures the router’s URL as the API endpoint, the router terminates the client-side TLS connection, and it originates a separate TLS connection upstream.” TLS protects each leg. It does nothing to protect the payload from the router, which reads the request JSON, forwards it, reads the response JSON, and returns it, with the chance to change either one.1

Routers also chain. The paper’s example is a developer who buys access from a Taobao reseller, who aggregates keys from a second-tier aggregator, who routes through OpenRouter, which dispatches to the model host: four hops, each with full plaintext access. “The client configures only the first hop; subsequent hops are invisible.” One bad hop anywhere taints the path, and the honest hops after it cannot tell.1

Nothing in today’s APIs closes that gap. Tool-call arguments travel as plaintext JSON, and “No provider-level integrity mechanism binds the arguments returned by the model to the arguments received by the client.”1

How the Measurement Worked

For each router the authors ran four stages. They registered a researcher-controlled account and obtained API credentials. They routed that account through a sandboxed agent environment running “a randomized subset of custom tasks and TerminalBench 2.0 tasks,” which produce realistic shell, package-install, and file-manipulation tool calls. They executed every returned tool-call payload inside the sandbox, which watched for network callbacks, file-system changes, and process-tree anomalies. A separate analysis agent then reviewed flagged executions and classified each as benign rewriting, potentially malicious, or confirmed malicious.1

The definitions are conservative. A router counts as injecting malicious code “only if a returned tool-call payload is rewritten into an attacker-controlled command or dependency.” An AWS canary counts as touched when a credential that crossed the router “later produces follow-on AWS API activity attributable to that credential.”1

The corpus is the commodity market, not the whole field. The 400 free routers are dominated by two open-source templates, sub2api and new-api, and the authors say enterprise and invite-only deployments are outside the study. Most of the corpus comes from Chinese-language marketplaces and communities.1

The Field Data

Behavior Paid routers (28) Free routers (400)
Injects malicious code into returned tool calls 1 8
Hides the injection behind a trigger 0 observed 2
Uses a researcher-owned AWS canary credential 0 observed 17
Drains a researcher-owned Ethereum key 0 observed 1

The injections are not cosmetic. The paper says they “change the command or dependency that the agent intends to execute.” The paid result is one router in 28, and the authors draw the plain conclusion: “Paid access improves service stability, but it does not prove tool-call integrity.”1

The two routers with triggers are the finding that defeats testing. In one, “malicious tool calls are injected only after the first 50 requests, which hides the attack from shallow smoke tests.” In the other, injection is enabled “only for sessions that appear to be running in autonomous YOLO mode and only when the project fingerprint matches Rust or Go.” The authors add that the triggers they saw may not be the only ones. They observed these routers only through finite black-box probing, so “a router may implement additional latent conditions that our probes did not activate.”1

The credential findings are all in the free set. Seventeen routers “trigger follow-on use of at least one researcher-owned AWS canary credential after seeing it in transit,” and one drained a prefunded Ethereum key. The appendix puts the loss from that drain at under US$50.1

The Two Poisoning Studies

A router does not have to be malicious to end up in the same position. The authors ran two studies of how a benign-looking router path gets poisoned.

Study 1: a leaked key. They leaked a researcher-owned OpenAI API key “on Chinese forums, WeChat, and Telegram groups frequented by router operators.” That one key served 100M GPT-5.4 tokens and exposed more than seven Codex user sessions, and in at least one session several distinct credentials appeared behind it. The paper is careful about how the traffic got there. The result “is consistent with several paths: direct reuse by a router operator, direct reuse by unrelated third parties, or reuse somewhere inside a longer relay chain,” and the authors “cannot always distinguish among those paths from provider-side logs alone.” What the study establishes is the blast radius: whoever holds a reused upstream key sees the traffic behind it in plaintext.1

Study 2: weak relays. They deployed vulnerable claude-relay-service instances across 20 domains and weak-password Sub2API and CLIProxyAPI instances across 20 IPs. The decoys drew more than 40,000 unauthorized access attempts from 147 IPs, which the authors read as a mix of scanning and relay reuse and not, in every case, as targeted attacks. The stronger signal came later: the decoys “were folded into active agent-facing relay paths.” They served roughly 2B GPT-5.4 and 5.3-codex tokens, about 13 GB of visible prompt and response traffic, and exposed 99 unique credentials across 440 Codex sessions spanning 398 projects or hosts.1

Every one of those 440 sessions was command-injectable by the paper’s definition, and 401 “already operated in YOLO mode, i.e., with tool execution auto-approved rather than confirmed command by command.” For those 401, the authors write, “sophisticated adaptive evasion was unnecessary: basic payload injection would have been sufficient to change executed commands.”1

Those were other people’s sessions. Their traffic reached the researchers’ decoys after the decoys were folded into relay paths, and a user who configures only the first hop would have had no way to see that. The authors kept aggregate metadata and salted hashes only, and retired their own exposed credentials when the observation window closed.1

Two Core Attack Classes and Two Evasion Variants

The paper reduces malicious-router behavior to two primitives and two refinements of the first.1

AC-1: response-side payload injection. The router rewrites a tool call after it leaves the provider and before it reaches the client, keeping the tool name and the schema intact. The paper’s example is a Bash call whose installer URL is replaced with an attacker’s. Because the result is still valid JSON for the expected tool, nothing looks wrong to the client.

AC-2: passive secret exfiltration. The router changes nothing. It scans requests and responses for credential patterns and keeps what it finds. The traffic is unmodified, so “clients cannot observe router-side retention.” The exposure is wider than API keys: “system prompts, tool definitions, user-supplied file contents, and environment variables all transit the same plaintext channel.”

AC-1.a: dependency-targeted injection. A variant of AC-1 aimed at package-install commands. Instead of swapping a URL, the router “substitutes a legitimate dependency name with an attacker-controlled package pre-registered on the target registry.” The paper’s example turns pip install requests into pip install reqeusts. The command still installs from the trusted registry, so a gate that checks only domains passes it, and the bad package stays cached for later sessions.

AC-1.b: conditional delivery. A variant that decides when AC-1 fires. The router behaves honestly until a trigger matches. The paper lists five families of trigger: tool name, content keywords, user fingerprints such as YOLO mode, time windows, and request counts. “A router deploying a 50-request warm-up appears entirely transparent during a 10-request probe.”

The paper names the March 2026 LiteLLM compromise as AC-1 at scale: once an attacker controlled the proxy’s request pipeline, every tool call passing through was open to rewriting. (The paper describes that compromise as dependency confusion. The reporting I worked from in the supply-chain post traces it to a PyPI publishing token stolen through the Trivy compromise.)

What the Research Proxy Showed About Agent Clients

The authors built a proxy they call Mine that implements all four classes, and ran 1,000 tool-calling requests per attack module through each of four public agent frameworks: OpenClaw, OpenCode, OpenAI’s Codex, and Anthropic’s Claude Code.1

None of the four verified response integrity. The rewritten tool call arrived in a valid, framework-native form in 100% of AC-1 cases and 99.6% of AC-1.a cases. The paper is precise about what that does and does not show: “These are compatibility measurements, not end-to-end execution rates: they show that the tampered tool call reaches the client in a valid framework-native format, but downstream permission prompts or local execution sandboxes may still block the resulting action.”1

That caveat is where a permission prompt earns its keep, and why auto-approval matters so much in the decoy numbers. The authors did not release Mine.1

The Composition Pattern, One Layer Lower

I keep writing about the same structural bug: individually authorized components composing into unauthorized behavior. Trivy-to-LiteLLM was composition at the package layer. Silent egress was composition at the fetched-content layer. MCP tool poisoning was composition at the protocol layer.

The router attack is composition at the hop between client and provider. Your client is authorized to call the router. The router is authorized to call the upstream model. The model is authorized to respond. Every hop is authorized, and the composition still yields rewritten commands, because no party signs what crosses the hop, and copied secrets, because every hop reads it in plaintext.

The paper draws the line to MCP itself: “a malicious MCP server receives tool-call requests in plaintext and can return forged results, so the same basic manipulation and collection ideas transfer with adaptation to the MCP message format.” It also marks the difference. An MCP server sits on the tool-execution side and can forge tool output. A router sits on the path to the model and sees every tool call and the whole request context.1

Three Defenses, With the Paper’s Numbers

The paper evaluates three controls a client can deploy without any provider’s help. One caveat governs every number below: “All numbers in this section come from controlled artifact-side evaluation rather than field data.” The corpora are synthetic, generated by the authors’ scripts.1

Defense AC-1 AC-1.a AC-1.b Cost and limit
Fail-closed policy gate 100% blocked 100% blocked not measured 1.0% false positives; bypassed in 100% of samples in a simple adaptive benchmark
Response-side anomaly screening 89.0% flagged 50.0% flagged 50.9% flagged 6.7% false positives
Append-only transparency log audit only audit only audit only about 1.26 KB per entry

The policy gate blocks shell commands that “fetch from non-allowlisted domains or install packages outside a local allowlist.” On 1,000 benign commands, 500 AC-1 samples, and 500 AC-1.a samples it blocked every attack sample, and all of its false positives were installer fetches from mirror domains missing from the allowlist. The authors call it “the strongest immediately deployable control.” They also report its failure plainly. If the attacker hosts the payload on an allowlisted domain, or pre-positions a local script that a benign-looking command later runs, “the policy gate is bypassed in 100% of samples.” Their verdict: the gate “should be understood as a coarse but useful filter for obvious high-risk workflows, not as a robust defense against defense-aware attackers.”1

One detail matters for anyone building the gate, and the paper is not consistent about it. Section 4 describes the gate as domain-only and designs AC-1.a to beat it: “the gate blocks non-allowlisted domains but does not maintain a per-package allowlist.” Section 7’s gate, the one that scored 100% against AC-1.a, also blocks installs “outside a local allowlist.” On my reading, a gate without the package list misses the typosquat substitution entirely.

Anomaly screening scores each returned tool call on locally visible features, such as shell-risk patterns and argument entropy, with no upstream original to compare against. At a 6.7% false-positive rate it flagged 89.0% of AC-1 samples, 50.0% of AC-1.a, and 50.9% of AC-1.b. At a 1% false-positive budget it caught 43.5% of AC-1 and essentially none of AC-1.a. The authors rate it “useful for review prioritization” and weaker than the gate.1

The transparency log records the request, the response, the router URL, TLS metadata, and a hash of the raw response. It prevents nothing. It lets you answer, after an incident, how far a router or a credential reached and which sessions passed through it.1

The Fix the Authors Ask For

None of the three controls authenticates where a tool call came from. The authors say so: “No client-side control available today can prove that a router preserved the upstream provider’s response.”1

What would is a signature from the provider. The paper proposes “a provider-signed canonical response envelope, similar in spirit to DKIM for email,” covering the model identifier, the tool name, the tool arguments, the finish reason, and a client nonce, which the client verifies before executing any tool call. It reports that no one ships this: “To our knowledge, none of the major provider tool-use APIs or the current MCP specification expose a deployed response-signing mechanism for tool-call arguments today.”1

Two limits come with the proposal. Transport security does not substitute for it: mutual TLS and certificate pinning “can authenticate the router endpoint the client chose, but they do not say whether the returned tool call preserves upstream semantics.” And signing does nothing for stolen secrets: “AC-2 cannot be mitigated by response-signing proposals because the secrets are exposed on the request path before any provider-side mechanism can act.”1

What You Should Actually Do

If your agent calls a model through a router you did not build:

  1. Go direct where you can, and know the operator where you cannot. A base-URL change is all it takes to add a router, which is why it gets added without a decision. Make it one. “Trust” here means an external basis, such as a known team, a contract, or a jurisdiction you can enforce against. Marketplace reviews are not one.
  2. Fail closed on high-risk tool calls. In Claude Code that is a PreToolUse hook that blocks shell commands fetching from domains outside an allowlist and package installs outside an allowlist. Keep the package list: the dependency-substitution variant exists to beat a domain-only gate. Write the hook to deny anything it cannot parse, with exit code 2 or a deny decision: Claude Code treats other exit codes and timeouts as non-blocking, so a hook that crashes or stalls does not stop the call: it goes on through the normal permission flow and, in an auto-approved session, simply runs. Check that the hook ran on its first call, too. The documentation warns that “a mistyped path in settings.json leaves the gate silently disabled.” Expect to maintain both lists, and expect a determined attacker to route around them.3
  3. Never run auto-approved sessions through a router you do not control. The 401 sessions in the decoy study are the precedent. A permission prompt is one of the few things that stands between a rewritten tool call and its execution.
  4. Keep secrets out of the traffic. Passive collection changes nothing you can observe. Anything in a prompt, a tool result, or a file the agent reads crosses the router in plaintext. Scope credentials narrowly, keep them out of the context where you can, and rotate anything that has passed through a router you later doubt.
  5. Log locally. Requests, responses, the router URL, and a response hash, with secrets redacted from the requests first, stored where the router cannot reach. It will not stop an attack. It tells you afterward what was exposed.
  6. Sandbox execution. The paper notes that sandboxes “reduce post-execution blast radius but do not authenticate where a tool call came from.” Take the first half.

The Uncomfortable Implication

The router layer is a clean example of the agent ecosystem shipping infrastructure faster than it secures it. People want one key for every model, lower prices, and access from regions a provider does not serve. Routers deliver all three, and the market rewards them.

The same sequence has played out at the MCP layer, the package layer, and the fetched-content layer. A new layer of the agent stack appears. Developers adopt it before anyone audits it. Attackers arrive, then researchers. Here the researchers counted 428 routers, 9 injecting malicious code, 17 using planted credentials, 1 draining a wallet, and 401 auto-approved sessions flowing through relay paths that included the researchers’ decoys.1

The piece that would close the gap, a signed response from the provider, is not something an operator can add. Until providers ship it, the controls above reduce exposure and none of them proves a tool call is the model’s own.

Update, October 1, 2026: The Router You Host Is on the List Too

This post is about routers someone else runs. The advisory record fills in the other half. On October 1, 2026 the PyPA advisory database added eleven entries against LiteLLM, an open-source proxy that teams host themselves for the reasons above, one key for every model.2 None of them is new. Nine have been in GitHub’s advisory database and NVD since June 21, a tenth since mid-September, and the one that matters most for a proxy, because it exposes the provider keys, was published in LiteLLM’s own repository on August 26, with fixes on PyPI since August 9; GitHub rates it Moderate. They are worth reading together for what they say about the layer, and because every final release published before August 9 is inside the range of CVE-2026-84377.

The one to read first is CVE-2026-84377. In the repository advisory’s words, “Any authenticated LiteLLM proxy user could redirect an outbound provider call to a destination they control and cause the proxy to send its own configured provider credentials to that destination.” The cause is the shape of the check: “The proxy’s request-body validation was a denylist that did not cover every sensitive parameter and did not inspect parameters nested inside other request fields.” It is fixed in nine release lines, 1.88.6 through 1.96.2. For anyone who cannot upgrade yet, the advisory lists three workarounds in one sentence: “Set general_settings.allow_client_side_credentials to false so callers cannot override connection parameters, restrict proxy keys to trusted callers, and block the affected parameters (api_base, base_url, model_list, fallbacks, provider credential fields) at a reverse proxy or API gateway.”2 The first is not enough alone. In the source of 1.95.0, an affected release, the request-body check is skipped entirely only when that setting is true (a deployment’s configurable_clientside_auth_params can still exempt single parameters), so false is where an install that never touched it already sits, and the incomplete check is the flaw the advisory describes. Upgrading is the fix.6

A second entry, CVE-2026-59823, is the same flaw in miniature: the guard “blocks the api_base and base_url parameters but does not cover user_config,” so a caller with a valid virtual key could put an api_base inside it and point the proxy at any host. That one was patched in 1.83.9, on PyPI since April 17; its advisory followed in September.4

Two entries touch the MCP side of the proxy: improper authentication in the MCP proxy (CVE-2026-12773, with a version range that ends at a fix in 1.84.0) and server-side request forgery through the spec_path argument of the MCP OpenAPI spec loader (CVE-2026-12798, recorded as affecting versions through 1.82.2). The other seven cover admin key handling, the SSO debug flow, SSO session invalidation, session expiry for generated keys, user enumeration in the UI, a guardrail bypass on async endpoints, and machine-to-machine JWT handling.5

Those nine records are thinner than the first two. Their descriptions carry VulDB’s templated wording rather than a maintainer’s write-up, and the MCP proxy entry’s description says “up to 1.59.8” while its version range says fixed in 1.84.0.5 CVE-2026-84377’s record has a mismatch of its own: the repository page lists affected versions as “<1.94.0” while the reviewed record’s ranges cover the 1.96 line up to its fix in 1.96.2.2

Check the ranges against the release you run rather than trusting a summary, this one included. Every recorded range ends at or below 1.96.2, so releases from 1.97.0 on, on PyPI since August 16, fall outside all eleven; the current release on October 1 is 1.103.2.5

The router’s danger is where the credentials sit, whoever operates it. A marketplace router that touches a planted AWS key and a self-hosted proxy that can be talked into sending out its provider keys are the same exposure reached from two directions, one by the operator and one by any tenant holding a virtual key. The defenses in the sections above assume you are the client.

For a proxy you run, add the operator’s half: treat every virtual key as a credential to the upstream keys behind it, keep client-supplied connection parameters off, including per-deployment configurable_clientside_auth_params, unless you need them, and put the proxy on the same patch clock as anything else that holds secrets.

If anyone you do not fully trust held a virtual key while the proxy ran an affected release, rotate the provider keys and any other secrets configured on the proxy, and check its logs for outbound calls to hosts you did not configure; the advisory’s impact covers “other configured secrets” and requests to “internal services reachable from the proxy.” Two of the eleven entries, CVE-2026-12773 in the MCP proxy and CVE-2026-12795 in the SSO debug flow, are authentication flaws in releases below 1.84.0, so on those releases who held a key may not bound who could reach the proxy; if it was reachable from networks you do not trust, apply the same rotation.

What the April Version Got Wrong

The April 10 version of this post was written from the paper’s abstract. Read against the full text on October 2, 2026, it was wrong or misleading in these places, all corrected above:

  • AC-1.a. I described it as an injection that “only fires when the request matches a specific dependency or context.” It is package-name substitution inside an install command. Triggers are AC-1.b.
  • The defenses. I wrote that “the abstract does not rank the defenses” and offered a ranking as my opinion. The paper’s body measures all three, calls the policy gate “the strongest immediately deployable control,” and reports that in a simple adaptive benchmark it was bypassed in 100% of samples. I left that out.
  • The signing advice. I told operators to sign requests on the client and verify them upstream, and called that “the only real fix.” The paper asks for the opposite direction, a response signed by the provider and verified by the client, and says signing cannot address stolen secrets.
  • The leaked key. I said the key was leaked “as if it had been exposed through a developer mistake” and concluded that “The router was a laundering layer for a stolen key.” It was leaked on forums and in chat groups where router operators share credentials, and the paper says it cannot always tell whether a router operator, an unrelated third party, or a longer relay chain reused it.
  • The credential counts. The answer block said “17 of 28 paid routers touched planted AWS credentials,” and the description said the researchers tested 28 routers. The paper’s paid row shows no credential abuse observed. All 17 routers that used AWS canaries, and the one that drained ETH, are among the 400 free routers.
  • The hooks advice. I recommended PostToolUse hooks that “validate response shapes.” A PostToolUse hook runs after the tool has executed, which is too late for a rewritten command. The control the paper tests is a gate before execution, which in Claude Code is a PreToolUse hook with an allowlist.
  • Smaller errors. The paper has six authors, not five. It does not say that a router “knows when it is being sampled”; that was my embellishment. The opening called the subject “MCP trust chains,” and the paper studies routers, not MCP. I described the silent-egress post as being about tool descriptions when it is about instructions hidden in fetched content.

FAQ

What is an LLM API router in this context?

A service that accepts requests in a unified, usually OpenAI-compatible format, selects an upstream model provider, and returns the response. It is an application-layer proxy with plaintext access to every request and response.1

Does TLS protect me from a malicious router?

No. The client configures the router as its endpoint, so the router terminates the client’s TLS session and opens a separate one upstream. TLS protects each leg and does nothing to protect the payload from the router.1

How would I detect a router that is rewriting tool calls?

Not reliably by testing it. Two routers in the study injected only after 50 requests or only for auto-approved sessions on Rust or Go projects, and the paper concludes that “no fixed-length client test can guarantee that the router is benign.” A fail-closed allowlist on shell commands and package installs blocks the plain cases, and the paper shows a defense-aware attacker gets past it.1

Does a PreToolUse hook help?

Yes, as a policy gate. The hook sees the tool call the client received, which is the rewritten one if a router rewrote it, and can block commands that fetch from unlisted domains or install unlisted packages. It cannot tell whether the call is what the model produced.13

I’m running Claude Code directly against api.anthropic.com. Am I affected?

Not by the router attacks in this paper, because there is no intermediary. If you route Claude Code through a proxy for any reason, such as a corporate gateway or a model aggregator, that proxy holds the same position.

What about OpenRouter, LiteLLM, or other well-known aggregators?

The paper measures 28 paid routers from three marketplaces and 400 free routers built mostly from two open-source templates. It names LiteLLM and OpenRouter as background and does not test them. The structural point applies to any router: it can read and rewrite the traffic, and visibility is a different property from integrity. For a proxy you host yourself, the October 1 update above covers eleven LiteLLM advisories.

Whose were the 401 auto-approved sessions?

Third parties whose traffic reached the researchers’ decoy relays. If you run auto-approved agent sessions through a router you did not build, stop, rotate every credential that crossed it, and review the session logs for tool calls you did not expect.


References


  1. Hanzhi Liu, Chaofan Shou, Hongbo Wen, Yanju Chen, Ryan Jingyang Fang, and Yu Feng, “Your Agent Is Mine: Measuring Malicious Intermediary Attacks on the LLM Supply Chain,” arXiv:2604.08407v1, April 9, 2026, listed for the ACM Conference on Computer and Communications Security, October 2026. Full text read October 1 and 2, 2026. Sections used: Introduction and 2.1 (what routers are, the four-hop example, TLS termination); 2.2 (no provider-level integrity mechanism); 4.1 and 4.2 (the four attack classes, the requests to reqeusts example, the five trigger families); 5.1 (the four-stage pipeline and the definitions); 5.2 and Tables 3 and 4 (1 paid and 8 free routers injecting, 2 free routers with triggers, 17 free routers using AWS canaries, 1 draining ETH); 5.3 (the leaked key and the decoys: 100M tokens, more than seven Codex sessions, 40k+ access attempts from 147 IPs, about 2B tokens, about 13 GB, 99 credentials, 440 sessions, 398 projects or hosts, 401 in YOLO mode); 5.4 (the key findings, including the sentence on paid access); 5.5 (scope); 6 and Table 5 (Mine, four frameworks, 1,000 requests per module, 100% and 99.6% compatibility); 7 and Table 6 (the three defenses and their results); 8.2 and 8.3 (the signed response envelope, MCP); 9 (the difference between an MCP server’s position and a router’s); Appendix A (retention, credential retirement, the drain under US$50, Mine not released). ↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩

  2. LiteLLM repository advisory GHSA-3cv6-jpf6-8222 (CVE-2026-84377, PYSEC-2026-4066), “Authenticated SSRF and provider-credential exfiltration via unvalidated request-body routing parameters,” published in the repository on August 26, 2026, listed by NVD on September 2 and reviewed by GitHub on September 30; read on the repository page and through the OSV record on October 1, 2026. Impact and the full Workarounds sentence are quoted from the advisory; the OSV record’s Patches line reads “Fixed in 1.96.2, 1.95.1, 1.94.3, 1.93.2, 1.92.2, 1.91.5, 1.90.7, 1.89.7, and 1.88.6,” and the repository page lists “Affected versions <1.94.0.” The count of eleven is the PyPA advisory database’s entries PYSEC-2026-4066 through PYSEC-2026-4076, all for the litellm package, each dated October 1, 2026, none withdrawn. ↩↩↩↩↩

  3. Anthropic, Hooks reference, Claude Code documentation, fetched October 2, 2026. PreToolUse runs “Before a tool call executes. Can block it”; PostToolUse runs “After a tool call succeeds”; “Any other exit code doesn’t block on its own for most hook events”; a timed-out command hook “doesn’t block the tool call”; and “a mistyped path in settings.json leaves the gate silently disabled.” ↩↩↩

  4. GitHub Security Advisory GHSA-hx8v-g79f-8w5f (CVE-2026-59823, PYSEC-2026-4070), “LiteLLM Proxy has server-side request forgery via the user_config request parameter,” published September 17, 2026, read through OSV on October 1, 2026: affected <= 1.83.8, patched 1.83.9, with the advisory dating that release to April 17, 2026. ↩

  5. OSV records read October 1, 2026: PYSEC-2026-4067 (CVE-2026-12773, MCP proxy authentication), PYSEC-2026-4069 (CVE-2026-12798, MCP OpenAPI spec loader), and PYSEC-2026-4068, 4071, 4072, 4073, 4074, 4075, and 4076. The GitHub advisories behind these nine were published on June 21, 2026, the day NVD listed them, and each cites VulDB. Upload dates and the current version are from the PyPI project page and its JSON, read the same day: 1.88.6 through 1.95.1 on August 9, 1.96.2 on August 11, 1.97.0 on August 16, and 1.103.2 on October 1. ↩↩↩

  6. Author’s reading of the litellm 1.95.0 wheel from PyPI (affected range 1.95.0 up to the fix in 1.95.1), October 1, 2026: in litellm/proxy/auth/auth_utils.py, _check_banned_params returns before rejecting anything when general_settings.get("allow_client_side_credentials") is True, and otherwise rejects a request whose body carries a parameter on the banned list unless the deployment’s configurable_clientside_auth_params allows that parameter. The banned list in that release includes vertex_ai_credentials and observability credentials and hosts. I read one affected release, not all nine lines. ↩

Related Posts

The Fork Bomb Saved Us

The LiteLLM attacker made one implementation mistake. That mistake was the only reason 47,000 installs got caught in 46 …

7 min read

MCP Servers Are the New Attack Surface

50 MCP vulnerabilities, 30 CVEs in 60 days, 13 critical. Tool-use protocols are the attack surface nobody is auditing — …

8 min read

The Ralph Loop: How I Run Autonomous AI Agents Overnight

I built an autonomous agent system with stop hooks, spawn budgets, and filesystem memory. Here are the failures and what…

11 min read