Push it until something breaks: is any of this actually sturdy, and why I stopped before hard isolation
Final post in the agent sandbox series. What twenty concurrent long tasks prove and what they do not; why the node-killing chaos experiments cannot be run on one laptop; and the theory behind "container isolation ends here, virtual machines start" — the point where my hardware ran out.
Each of the previous five posts cleaned up the mess the one before it made: isolation → cold start → idle cost → lost state → peak queueing. Every step holds on its own.
But “holds on its own” and “holds together” are different claims. This last post does two things: push it until something breaks, and state plainly where I stopped and why.
The load test: twenty concurrent long tasks
The final lab assembles the earlier pieces, and the load script is plain:
async def main():
print("launching 20 concurrent 60-second tasks...")
tasks = [submit_task(client, i) for i in range(20)]
await asyncio.gather(*tasks)
Twenty requests at once, each about a minute of work, against a pool holding a handful of containers.
What it does prove — the promises from earlier posts, delivered:
- none of the twenty tasks is lost; the surplus waits in the queue;
- users get a position in line, not an error page;
- progress can be queried at any time, from any container, because it lives on the shared disk;
- a container that finishes picks up the next one without anybody intervening.
What it does not prove — the more important half:
- twenty concurrent is a laptop-sized number. Real systems break in the hundreds or thousands: the queue filling faster than it drains, connection limits, disk writes slowing down;
- nothing was broken during the run. No container killed, no node killed, no network cut. A load test measures busy, not broken;
- neither how long it takes to notice a failure nor how long recovery takes was measured — there is nothing in the system capable of measuring them.
The chaos experiments in the plan are ones I cannot run
The original design: hold 500 active sessions, kill 30% of the cluster’s nodes, kill the queue’s primary, then measure three numbers — time to detect, time to recover, how many users saw an error.
None of those three is possible in my environment, for concrete reasons:
- kill 30% of nodes — the local cluster has one node. Killing it is not a drill, it is a shutdown;
- kill the queue primary — Redis runs as a single instance with no replica, so there is no failover to observe;
- 500 active sessions — this laptop cannot hold 500 containers each consuming memory.
That is not a “later” excuse, it is where the entry price for this kind of experiment sits: chaos engineering needs a cluster that can absorb damage, not a machine that can run the code. The first five steps fit on a laptop because they test architectural defects. The seventh tests behaviour at scale, and scale cannot be faked.
While being honest, one thing the previous post already admitted: genuine automatic failover does not exist yet either. When a container dies, the task it held is not re-dispatched — someone has to notice, and someone has to resend it. Closing that needs a task-state table (who is working on what, since when) plus a timeout rule (silent for too long, take it back and re-issue). That is the line between “self-healing” and “resumable”.
One step further: container isolation ends here
The last topic of the series is a step I worked out on paper and did not build. Saying so explicitly matters, otherwise it reads as done.
Back to the question from post one: an agent executes code a model generated moments ago. One container per task stopped tasks from interfering with each other — but how hard is a container’s boundary, really?
An analogy: containers are partitioned rooms in one building. The walls are real, but the foundations and the plumbing are shared. That shared plumbing is the operating system kernel: every container uses the same one, and a flaw in it makes the partition walls decorative. For ordinary workloads the risk is acceptable, because the code running is code you reviewed. For untrusted generated code, the premise has changed.
The industry offers two roads, both heading toward “separate houses”:
| Approach | How it works | What it costs |
|---|---|---|
| gVisor | inserts a user-space kernel between the container and the real one, handling system calls on its behalf | still starts fast, but some system calls are unsupported and certain workloads slow noticeably |
| Kata Containers | puts each container inside a lightweight virtual machine with its own kernel | the hardest boundary, paid for in startup time and memory — and it requires hardware virtualisation |
Switching between them in Kubernetes is not hard: a RuntimeClass sends a class of workloads to a different runtime, with no application changes.
I did not implement it, for a specific reason: the local cluster runs inside a container on a laptop, and nesting a virtual machine inside that needs nested virtualisation — and more memory than the machine has. This step is not unclear, it is out of hardware. It is written up anyway because knowing where the boundary is and what it costs is itself the output of the step: untrusted code ends at a virtual machine, not a container, and that judgement does not require running it first.
The ledger for the whole series
Looking back, no step was simply “better”. Every one traded something for something:
| Bought | Paid |
|---|---|
| Isolation | 45-second cold start |
| Instant response | a row of containers idling permanently |
| Progress surviving a crash | a storage dependency, and a write-frequency trade-off |
| Peaks queueing instead of failing | one more component that must stay alive |
| (not done) a hard boundary for untrusted code | startup time, memory, hardware requirements |
Architecture is not designed, it is forced out by invoices. That is why this is a sequence of experiments rather than one final diagram: taken alone, the warm pool, the shared disk and the queue all look like over-engineering. Walked through in order of the bills, it is obvious what forced each component into existence.
What is still on the table, in my order of priority: make the pull model real (cheapest, and it removes scheduling contention outright) → a task-state table with timeout re-dispatch (that is what self-healing means) → storage that several machines can write to → and only then the isolation upgrade.
In one line: running is not the same as holding up; stating what went unverified is worth more than one more component.
All the code: agent-sandbox-oss. Thoughts welcome as issues.
- Kubernetes
- AI Agent
- Chaos Engineering
- Security Isolation
- Architecture