- The Kubernetes project reports that Local SSD swap raised sandbox density on one node from 80 to 240 Python sessions, and from 80 to 160 gVisor pods.
- Idle agents hold memory while waiting, which caps pods per node. Swap moves dormant memory to disk, but pushing active memory there slowed a build by over 40%.
- Ask your team how much agent memory sits idle, test your own sandbox runtime, and check how paged-out memory is protected.
Waiting is the expensive part of an AI agent
An AI agent spends little of its life working. It starts up, runs a burst of code, then waits for the next prompt. While it waits, its memory stays in the server's RAM. In a benchmark post dated October 5, 2026, the Kubernetes project argues that memory held by sleeping agents limits how many pods fit on one node. The lesson here: for agent workloads, cost is driven less by what agents do than by how long they sit still.
Administrators have long balanced two risks. Generous memory limits waste money on RAM that sits unused. Tight limits can get a process killed when memory runs out, which is called an out-of-memory kill. Agents sharpen the choice. Running untrusted code safely means a sandbox must reserve a big block of memory for its start-up and its burst of work. Afterward it often sits idle, still holding that block.
How node swap works
Swap lets the Linux kernel move memory that nobody is using onto disk. Kubernetes support for nodes with swap reached General Availability, meaning it is considered stable, in v1.34. The Kubernetes team backed that swap with fast NVMe local solid state drives (SSDs) and measured the effect.
Swap was discouraged for two reasons, the post says. Under cgroup v1, the older resource-control system, memory and swap shared one combined limit. That made a container's real memory use hard to predict. Kubernetes swap support relies on cgroup v2, which tracks disk swap separately. The second reason was slow spinning disks, which fast local SSDs largely remove.
To use it, operators set the kubelet, the agent that runs on each node, to LimitedSwap. They also give workloads Burstable QoS, which means setting memory limits higher than requests.
What the tests found
The Python test ran sandboxed sessions that each analysed 5 million rows of MovieLens data, using about 375 MiB each. Without swap, the node hit its RAM limit and failed at 80 sessions. With swap it reached 240, a 3× gain.
The gain varied by runtime, the post's other tests show. Plain runc containers, with no extra security layer, rose from 512 to 768 pods, a 1.5× gain. That test ran on a c4-standard-32 node with 32 vCPUs and 120 GB of RAM. The post states that node type only for the runc test.
Pods in the gVisor sandbox doubled, from 80 to 160. The post says swap "naturally absorbs" the memory overhead of security runtimes. Kata Containers microVMs rose from 40 to 50, a 1.25× gain. That test ended when the CPU was saturated, so memory was no longer the limit.
A traditional workload showed a similar benefit. A Linux kernel build needed a 600 MB memory limit without swap. With Local SSD swap, 300 MB worked, and the build ran in 374 seconds against 433.
Where the limits are
Swap does not make memory free. Cutting the kernel build's limit to 200 MB forced active memory onto disk, and run time rose by more than 40%. The authors call swap an insurance policy for bursts, not a replacement for active RAM.
Latency needs care too. At peak density, the researchers say, slower responses came mainly from pods competing for CPU, not from swap I/O. They add that an operator aiming for a latency target would run below the peak numbers.
These are the Kubernetes team's own benchmarks. Other clouds, disks and workloads may behave differently. The 3× figure is the post's "up to" number, not a typical result. The post reports density gains, not dollar savings.
Questions for your platform team
First, how much of your agent fleet's memory is idle at any moment? Without that figure, you cannot tell whether swap would help or whether you are simply short of RAM.
Second, what does your isolation choice cost? Gains varied by runtime in these tests, so test yours rather than rely on the 3× figure.
Third, what lands on disk? The post does not say how paged-out memory is protected. Agents running untrusted code may hold sensitive data, so ask your security team to answer this before enabling swap.
Finally, who is watching latency at peak density? Set a latency target first, then size the node to meet it.
Agents made idle memory a budget problem. Swap offers a way to fit more pods on the same hardware, but only for memory that is truly idle.
Produced by the WebPulse Newsroom with AI assistance from the original reporting credited below, and checked against that source by our editorial review. How we use AI.
Original reporting: Kubernetes.





