- Cloudflare released on-demand CPU and memory profiling for Workers and Durable Objects, so teams can inspect live production code as a flamegraph.
- In Cloudflare's own example, monitoring code believed to be disabled made up roughly 66.7% of a Worker's allocations and pushed it past its memory limit.
- Ask which services run close to their limits, and whether anyone has measured code that is supposedly switched off.
Switched off is not the same as not running
Cloudflare engineers had a Worker that kept getting evicted with an "Exceeded Memory" error. A Worker is a small program that runs on Cloudflare's network. Its P999 memory sat around 133 MB, against a 128 MB limit. P999 is the memory level that 99.9% of requests stay under.
Errors and metrics told the team that memory was the problem. They did not say which code was responsible.
So the team took a heap profile of a Worker running in production. A heap profile records which code allocates memory. They opened it with a tool called pprof. Their Prometheus code, a monitoring library, accounted for roughly 66.7% of the allocations.
That code was supposed to be disabled. The profile showed it was only partly disabled, and the memory cost was still being paid in full. The team removed it. Cloudflare reports that this left the P999 about 10 MB below the 128 MB limit.
The lesson here is simple. Dashboards report symptoms, and profiles name the cause. The costliest assumptions are the ones nobody thinks to check, such as "that feature is off."
What Cloudflare announced
On October 9, Cloudflare announced on-demand CPU and memory profiling for Workers and Durable Objects. Durable Objects are Cloudflare's stateful, named objects. From the Observability page in the dashboard, or from the command line, a team can request a profile of an active Worker. The team can inspect the result as an interactive flamegraph and download the file for further analysis.
Cloudflare notes that the request can also be sent directly, "or have your coding agent do it for you." The dashboard is simply one client of that request.
How to read a flamegraph
A flamegraph turns thousands of samples into one picture. Each rectangle is a function call. Its width shows how much CPU time or memory that function used. The widest rectangles are the biggest consumers.
Cloudflare also points to a table view that sorts functions by how often they appear in the samples. One practical note: TypeScript projects need source maps enabled. Without them, function names may be unreadable.
Cloudflare's own CPU example came from the Worker behind its R2 storage binding. It profiled the Worker for 50 seconds. One function, genericR2JsonReplacer, used over 5% of CPU time and was calling itself recursively. The function was walking the data tree even though JSON.stringify already visited every node. Data at the fifth level of nesting was therefore handled five times over. Fixing it made the function 2.7x faster.
A second find was a duplicate call to a metrics function. That one call used about 1% of CPU time. Saving the result of the first call removed it. For a Worker with heavy traffic, Cloudflare says, small amounts of wasted work matter.
How it profiles live traffic without stopping it
Profiling production is harder than profiling a laptop. Workers run in many data centers and on many physical servers. Before a session starts, the platform has to work out which version to profile and where it recently ran. It must also check that the program is still loaded there and that it belongs to the requesting account. For a Durable Object, it must find the live primary copy.
Cloudflare says it does not start a new isolate, the sandbox a Worker runs in, to produce a profile. The aim is to watch real production execution. A Worker with little traffic can therefore be hard to profile.
For CPU profiles, the runtime starts a V8 CPU profiler that samples every millisecond. To keep traffic flowing, the runtime takes the isolate lock only briefly, to start and stop the profiler. Cloudflare says keeping the lock for the whole session would freeze the code being measured, which would defeat the purpose.
Durable Objects have names. A team can therefore choose a specific object to profile, and the runtime routes the request to the server that owns it.
The limits
Cloudflare lists two. You must start a session yourself, so you can miss the moment a Worker misbehaves. And the memory profiler only shows allocations made during the profiling window, so start-up allocations do not appear. Cloudflare says it is working on continuous profiling to capture rarer events.
This is one vendor's feature, not evidence of an industry shift. The finding that matters is the monitoring code that was assumed to be off.
What to ask your teams
Ask which production services run close to a memory or CPU limit, and what their P999 figures are. Ask which monitoring, logging or feature-flagged code is described as disabled, and whether anyone has measured it. Ask whether performance work happens on live traffic or only in local tests. Cloudflare notes that a local run does not see the same kinds or volume of requests.
If your platform is not Workers, ask your vendor what it offers for profiling live code. A dashboard tells you something is wrong. A profile tells you which code is doing it.
Produced by the WebPulse Newsroom with AI assistance from the original reporting credited below, and checked against that source by our editorial review. How we use AI.
Original reporting: Cloudflare.





