Containers die whenever; two hours of progress should not die with them
Fourth post in the agent sandbox series. Move progress out of container memory onto a shared disk, then delete the working container on purpose — another one picks up at page 17. Including the honest part: this "automatic recovery" still needs a human to press something.
- Kubernetes
- AI Agent
- Fault Tolerance