Observed Ray listening sockets missing from declared port metadata: 15 of 17 (Source: Sorami Consulting technical report (September 29, 2026))
A floor plan shows the doors the architect drew. It does not show which doors are unlocked tonight. A review that works only from declared ports and default scanner output is working from the floor plan. One new test on a single Amazon EKS cluster shows how far that can sit from the running system.
Sorami Consulting published the study on September 29, 2026. Its experiments ran on September 28. The lesson is this: a list built from what software declares records intent. It does not record exposure.
What Sorami tested
The team stood up a new test cluster on Amazon EKS, version 1.35. It hosted nothing else. The cluster ran the vendors' default charts and images for Ray, vLLM and KubeRay. Together, these tools spread AI model serving across several machines.
The team probed from one pod it had built in an unrelated namespace. It aimed only at pods it owned, using a fixed list of their addresses. It scanned no wider address ranges.
The port list and the running system disagree
In the two-node Ray setup, the neighbour pod reached Ray's control-plane listeners across namespaces. These were the GCS on port 6379 and the raylet RPC ports 10002 to 10006. They spoke gRPC without encryption.
The pods' declared port metadata told a different story. The researchers observed 17 Ray listening sockets. Fifteen were missing from the declared ports.
Sorami calls this visibility drift. It would affect any tool that treats declared ports as its exposure inventory, in a setup like this one.
Scanners did not close the gap. Four default configurations skipped the RayCluster resource. They were Trivy, Checkov, Kubescape and kube-linter. The Ray pods it creates were not part of their analysis.
Sorami notes that these are runtime and network properties. The team did not test custom scanner rules.
What held, and what did not
One control held. A network rule that denied all incoming traffic by default shut out the neighbour pod. Every Ray port it could reach before the rule became unreachable after it.
That result covers steady state only. Sorami did not test the window when pods start up.
The API key result was weaker. On two /v1 routes, adding a key turned unauthenticated requests from a 200 response into a 401. Sorami says this shows the normal key path works. It does not show that API keys protect vLLM.
The vLLM release in the test, v0.11.0, falls inside the range hit by CVE-2026-48746. That bug lets a caller get around the API key through the Host header. It was fixed in 0.22.0. Sorami did not test the bypass.
Separately, a single-GPU vLLM engine with no authentication set up exposed 26 routes. These included tokenizer and prefix-cache telemetry.
What the test does not show
Sorami measured this once, on one setup. The researchers say these are documented vendor defaults, not new vulnerabilities. The defaults are Ray's assumption of a trusted network, vLLM's optional authentication, and charts that ship no NetworkPolicy.
On the default CPU Ray image, the Job API on port 8265 accepted requests with no login. The GPU vLLM setup did not run that listener. So the team did not show that an attacker could run code on the GPU serving stack.
The two hardening layers tested here showed no measurable steady-state cost. The network rule and the API key made no measurable difference in steady-state speed, within variation between sessions. That rests on one run per level. Median latency stayed within 5% of baseline.
Questions to put to your platform team
Nothing here says your cluster looks like this one. It does suggest a few checks before an AI serving stack goes live.
First, ask how the exposure inventory is built. If it comes from declared ports, ask what compares it with the sockets that are actually listening.
Second, ask whether your scanners understand custom resources such as RayCluster. Ask what they report when they skip one.
Third, ask whether inference namespaces have a default-deny network rule. Ask who tested it, and whether startup was included.
Fourth, ask which vLLM version is running. Check it against CVE-2026-48746.
Sorami released its logs, manifests and raw results under a CC BY 4.0 licence. Your team can check each claim. Run the probes only in a cluster you own that runs nothing else. The floor plan is a starting point. Someone still has to walk the building and try the doors.
Produced by the WebPulse Newsroom with AI assistance from the original reporting credited below, and checked against that source by our editorial review. How we use AI.
Original reporting: Sorami Consulting.





