# QA Wolf Sandboxes Give Each AI Agent Its Own Computer

> Published 2026-10-06 · https://www.promptzone.com/saoirse_quiroga/qa-wolf-sandboxes-give-each-ai-agent-its-own-computer-20n2

QA Wolf published details on running every AI agent inside its own dedicated sandbox computer. The post appeared on Hacker News where it received 12 points and one comment.

The approach assigns each agent a separate virtual machine or container with its own filesystem, network stack, and process space. This prevents one agent from reading another agent's state or files.

## How the Sandboxes Work

Each sandbox starts from a clean base image that includes only the tools the agent needs. QA Wolf provisions the environment on demand and destroys it after the task completes. Agents connect to the sandbox over a controlled SSH or API channel rather than running directly on the host.

The system logs every file change and network request inside the sandbox. These logs feed into replay tools that let engineers inspect exactly what an agent did during a run.

## Measured Performance Numbers

QA Wolf reported average sandbox startup time of 4.2 seconds on their current hardware. Full teardown and log upload takes 1.8 seconds. Memory overhead per idle sandbox sits at 180 MB when using lightweight container images.

In one production workload, 47 agents ran concurrently across 12 physical hosts without cross-agent interference. The team measured a 3.1× reduction in failed task retries compared with their previous shared-environment setup.

## How to Replicate the Setup

Start with a base Ubuntu image and install only the minimal packages required. Use Docker with user namespaces enabled and bind-mount only the directories the agent must access. Add a seccomp profile that blocks syscalls such as mount and ptrace.

For orchestration, QA Wolf uses a simple queue that spins up instances on Kubernetes pods with resource limits set per pod. After the agent finishes, the pod is deleted and its persistent volume is wiped.

## Tradeoffs Observed

Dedicated sandboxes add measurable startup latency versus running agents directly on a shared host. They also increase total infrastructure cost because idle capacity must be reserved for peak concurrency.

The isolation benefit is clearest when agents handle untrusted code or external data. In lower-risk internal tasks the extra overhead may not justify the gain.

## Alternatives and Direct Comparisons

| Approach              | Startup Time | Isolation Level | Cost per Agent | Typical Use Case          |
|-----------------------|--------------|-----------------|----------------|---------------------------|
| QA Wolf per-agent VM  | 4.2 s        | High            | Medium         | Production agent fleets   |
| Shared container pool | 0.8 s        | Low             | Low            | Internal testing          |
| Firecracker microVM   | 1.9 s        | High            | Medium         | High-security workloads   |
| gVisor sandbox        | 2.7 s        | Medium          | Low            | Quick prototyping         |

## Who Should Use Dedicated Sandboxes

Teams running agents that execute third-party code or access production data benefit most. Organizations with strict audit requirements also gain from the per-agent logs.

Teams running only internal prompt-chaining workflows on trusted data can usually stay with lighter shared containers and avoid the added latency.

## Bottom Line

QA Wolf's per-agent sandbox model trades modest startup cost for strong isolation and replayability, a practical choice when agent failures carry high downstream risk.

The pattern is already influencing other agent platforms that need reproducible execution environments.