- Synacktiv's build log of an on-premise LLM server shows that privacy comes from isolation choices, such as rootless containers and no network, not from the server's location.
- The team chose a stateless first version and accepted a container-versus-VM compromise, with a ~9,000€ GPU needed to fit a 120B-parameter model.
- Before funding a private AI server, ask for its isolation decisions and trade-offs in writing.
On-premise is a location, not a control
Many companies say they will run AI at home to keep trade secrets secret. That is a sound instinct. But a server in your own building is only a place. The protection comes from how the server is built and what it is allowed to touch.
Synacktiv has published a detailed account of building exactly such a server. The lesson here is that private AI is a series of design decisions. Each one buys some safety and costs something in speed, effort or hardware.
What the team set out to build
Synacktiv's stated goal was fully air-gapped LLM instances, which removes network data exfiltration as a risk. Several teams wanted the capability. Their candidate uses included document translation, proofreading, log crawling and codebase analysis.
The first version is deliberately small. It holds no input data and involves no custom training. It also has no document-search layer (known as RAG, a way of letting a model draw on company files), no agent and no connector. Synacktiv's reasoning is that a narrower scope gets the tool into use sooner.
That choice matters for risk. Every connector or document store added later widens what a compromised model server could reach. The team started with the narrowest version.
Why the hardware bill is driven by memory
The team chose the gpt-oss-120b model after testing several in real work conditions. At 4-bit precision, its weights take roughly 60 GiB. A model must also keep a running memory of the text it has processed, called the KV cache. Without it, the work to process a text grows with the square of its length. With it, the growth is linear.
That cache has to sit in the GPU's fast memory too. Synacktiv's sizing came to a minimum resident size of about 70 GiB. One card with 96GB could hold the whole model and four parallel cache slots, which keeps users from sharing cache data.
The figures are estimates. They leave out padding alignment and temporary buffers. The estimate left almost no margin once those were counted, so the team cut the context to 126,000 tokens to keep a 2 GiB safety margin. Memory sizing is a budget decision before it is a technical one.
How the isolation works, and what it costs
The model runs inside a container using Podman in rootless mode. A container is a sandbox that shares the host's operating system core. Rootless means the sandbox itself runs without administrator rights.
Synacktiv contrasts this with Docker's default setup. There, the administrator user inside the sandbox is the real administrator on the host machine. A flaw in the sandbox's isolation could then become a serious privilege escalation.
The team maps the container's user to a high, unprivileged number. The llama.cpp inference server runs as user 1000 inside the container. On the host, that is user 101000, which does not even hold the rights of the account that started the container.
The team also gave the container no network at all. The llama.cpp server can listen on a Unix socket, a local file-based channel, so it needs no TCP/IP stack. Synacktiv describes this as a gain in both performance and security.
The trade-off the team chose to accept
Containers share the host's NVIDIA kernel modules. Synacktiv says virtual machines with GPU passthrough would probably isolate better. They would also bring slower performance, harder orchestration and more tedious upgrades.
Synacktiv chose containers anyway. It points to the isolated network and to auditable open-source driver code as the reasons it can accept the weaker wall. That is a stated, reasoned compromise. Leaders should expect to see decisions like it, and should ask to see them written down.
One limit applies to this story. The text we reviewed stops partway through the hardening work, so it reports no final results or test findings. It is also one team's build, not a survey of how companies deploy private AI.
What to ask your team
Before funding a private AI server, ask for the isolation decisions in writing. Does the model process run as an administrator anywhere? Does it have a network connection it does not need? Which shared components, such as GPU drivers, sit between it and the host?
Ask what the first version leaves out on purpose. A model with no stored data and no connectors has less to lose. Ask also where each trade-off was accepted and who signed it off.
Private AI is not safe because it is private. It is safe to the degree that someone decided, and recorded, what it may reach.
Produced by the WebPulse Newsroom with AI assistance from the original reporting credited below, and checked against that source by our editorial review. How we use AI.
Original reporting: Synacktiv.





