Rendered at 21:43:08 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
srini-docker 1 days ago [-]
I work at Docker. Lot of valid and useful feedback here that we're looking closely at.
One correction: this isn't containers. Each session is a microVM with its own kernel on the platform's native hypervisor: Hypervisor.framework, WHP, KVM. We wrote a new VMM (not Firecracker) to make it more effective across platforms.
I'd like to see real numbers that compare Docker Desktop for macOS before microVMs to post-microVMs.
I stopped using Docker on macOS because host file system performance was so slow, even with all of the caching hacks piled on top of it, that it made the whole thing effectively unusable for development.
Directionally the post shared sounds great, but it seems "too good to be true" that we'd have a performant microVM for macOS.
I'm very glad this now exists - fwiw almost a decade ago I worked on https://github.com/takeoff-env/takeoff as a solution for making it easier for hot reloading your stack which is thankfully redundant today, and funnily enough the list of problems you identify is also something I've been working on.
Recently I've also been working on a VM stack for an agentic platform using pre-build images with some cloud injection scripts that simplifies the deployment of a private agentic cluster - in the end I went with full VM with a 4vCpu/8gb for the main agent and 2vCPU/4Gb - only the main agent had docker-in-docker, the rest rootless docker but I agree it's still an elevated risk.
I'll definitely have to give this a spin and see if I can simplify it to one larger box with this solution.
Humphrey 22 hours ago [-]
My feedback on sbx:
Concept is great - works quite well - I often have multiple short lived sandboxes running at once.
Docs [1] on overriding auth are incorrect. Sbx ignores inject[].username for basic auth and instead the stored secret needs to be the complete Authorization header. This should be made clear, or fixed.
Having to log in every couple of days SUCKS!! Opening the browser so I can login (which we shouldn't have to do) interfers with my scripts that create and destroy sandboxes as I need them.
I miss the old worktree functionality - I dislike the new clone concept - So I've created my own scripts that create a worktree for a feature, and run sbx create/run from there.
Do you have a strategy for secrets? Such as storing them, or using MitM to inject them (e.g., HTTP API requests)? I've used squid cache in the past, and currently use iron-proxy for this sort of feature.
mikesir87 1 days ago [-]
The secrets are stored in the OS-specific keychain. When a sandbox starts, the network proxy injects the secret into the request (as auth headers) only when the hostname matches.
Kits provide the ability to also define new credentials and how to inject them into new services (connect to internal systems, etc.).
filearts 1 days ago [-]
Is there any line of sight to open sourcing the vmm?
Is it based on libkrun?
kwakubiney 1 days ago [-]
Why’s it not on Linux? What are the difficulties with that platform?
srini-docker 1 days ago [-]
Linux is available today (Ubuntu): github.com/docker/sbx-releases. Our webpage showing only brew and winget is on us.
For the people upthread who asked about on customization: templates (like snapshotting a running sandbox) and kits (YAML applied at creation like install steps, files, network and credential rules, or define a new agent outright) are the supported path now. It's early but take a look here: https://docs.docker.com/ai/sandboxes/customize/
On MCP, since credential handling was mentioned here: the sandbox sees one gateway endpoint, and OAuth tokens stay in the host credential store rather than in the VM. https://docs.docker.com/ai/sandboxes/mcp-gateway/
All this is early. We're looking at more based on feedback from users like running sandboxes in the background for long-horizon work and a lot more (including what you all raised in the thread here). Keep them coming.
srini-docker 1 days ago [-]
Quick update: webpage reflects Linux support as well, thanks for the flag.
The parent didn't go into any detail. I can. Homebrew has a history of ripping out your foundation underneath you. One day you are on Python 3.8, then next day you are on Python 3.10 and all your packages are broken. MacPorts doesn't do that.
Now, whether you should you be using the Homebrew Python is a completely different question. YMMV for other platforms managed via Homebrew.
I've traditionally used MacPorts for dev tooling and Homebrew for everything else, but with more aggressive adoption of tooling like uv an nvm I'm not sure the different really matters for me anymore.
tacker2000 1 days ago [-]
Exactly. Same with PHP, MySQL etc…
Also they just block old versions and dont let you install them, you have to jump through a lot of hoops to use an old PHP version for example, so in no way developer friendly.
In the end I realized that Brew is a package manager for consumers, and as a professional i should’nt keep fighting it.
theshrike79 12 hours ago [-]
I use mise for dev tooling, that way I can have the exact correct version for every project.
Python through brew is the one I expect to be the latest one I use for one-off scripts.
lucumo 6 hours ago [-]
I switched to mise too for all my dev tooling. It just works so nicely for all kinds of ecosystems. I can use the same tool for Python, Node, Java, whatever and it just works.
Groxx 1 days ago [-]
A lot of that is simply formula authors / application devs who don't know what they're doing (python@3.10 and other versions are a thing, and have been for quite a while now, but they're not always used and devs don't always keep track of the version they need) and people not updating their software for years (pythons are on a 5 year cycle everywhere, homebrew included: https://devguide.python.org/versions/ and https://formulae.brew.sh/formula/python@3.10 ).
Python in particular is well known to not be a stable target. For anyone. By design. If you expect long term use of a specific version of code, use a different language. It is not at all homebrew's fault that they're how many people discover that.
xrisk 1 days ago [-]
pyenv has been standard tooling for far longer than uv. depending on package manager supplied Python packages only makes sense if you’re running rhel or Debian or something and your application is packaged/deployed/the maintenance path uses dnf/apt. Otherwise you should always use a venv and use an out of package manager update mechanism. Like, in a broader sense, vendoring dependencies only makes sense if you’re shipping an application, not on a dev box.
tacker2000 1 days ago [-]
This is not about python packages, this is about python itself.
stavros 1 days ago [-]
That's what the GP means as well, you can use pyenv and uv to install multiple versions of Python and create envs with whichever version you want to use.
cogman10 1 days ago [-]
Looks like they do support Ubuntu.
Is this open source? Can I install this on a non Ubuntu system?
rocfan 1 days ago [-]
CLI works on Fedora. Been using it daily for ~ a week.
See repo `docker/sbx-releases`. The `.rpm` there has Rocky Linux in the name but works on Fedora.
aborsy 1 days ago [-]
A limited form of it with different syntax comes with Docker Desktop. The sbx tool is not available for non-Ubuntu distributions.
mooreds 1 days ago [-]
What about inbound credential checking, for when agent A calls agent B?
I looked at the docs last week and didn't see anything about that.
atechboy 1 days ago [-]
So the idea is to give each agent a VM to do several tool calls? or one VM for each tool call?
theplumber 1 days ago [-]
I think the idea is to have the agent run in the VM
android_reverse 1 days ago [-]
VM? own kernel? I wonder if this could allow to run Waydroid on Windows without the hassle of recompiling WSL kernel with Binder and Docker
Geezus_42 1 days ago [-]
Sounds like what I'm doing with Nix and MicroVM currently.
Fun fact: QEMU runs natively on windows and supports acceleration with WHP. It works surprisingly well.
codethief 2 hours ago [-]
> supports acceleration with WHP
On Windows 11, too? At least for hardware virtualization in VMWare one would have to disable Windows Device Guard & Credential Guard for that.
srini-docker 1 days ago [-]
Yes!
adityazero 23 hours ago [-]
[dead]
rusch 2 days ago [-]
The login is annoying but, lacking an open source alternative, this has been my daily driver for a while now because it works great out of the box with two key features: outbound firewall and secret injection with placeholders.
I run it with superset and then each git worktree is mounted in a sandbox that is configured for each repo i work in.
Outbound firewall is `--network-isolated`: egress is denied except the agent's own API endpoints plus domains you allow, enforced sandbox-side (working on host-side enforcement now). `--network-none` if you want nothing.
Credential brokering works the way you describe (currently Claude-only, I'll add more as time allows). The API key stays on the host, a local proxy injects it into the outbound request, and the sandbox never holds anything worth stealing. Other agents' credentials currently arrive as read-only file mounts instead (weaker, and something I'll fix soon). Generalising the injector is the obvious next thing.
One difference from your setup: yoloAI copies your worktree instead of mounting it. The agent works on the copy, you `yoloai diff`, and `yoloai apply` replays the commits into your real repo. That's deliberate. Docker's own security docs talk about the dangers of bombs being left behind in a live-mounted dir (git hooks, package.json scripts, Makefiles, IDE task config), which diff/apply avoids.
Isolation is per-sandbox rather than fixed: runc, gVisor, or Kata VMs (QEMU or Firecracker) on Linux; Seatbelt or full macOS VMs via Tart on a Mac.
westurner 1 days ago [-]
> yoloAI copies your worktree instead of mounting it
A few months ago now I started adding seccomp sandboxing to jinja2rs and then liboverlayfs support to ansiblers (which are early Rust ports).
Haven't finished that, but
I started working on a VM format that stores signed machine state into an OCI container repository, using the hypervisor migration support of KVM/QEMU.
Though this is not safe yet if ever, VM migrations are probably another way to sandbox and deploy en masse.
kstenerud 13 hours ago [-]
[dead]
nextblock 1 days ago [-]
[dead]
jachris 2 days ago [-]
Agreed. Network control and secret injection together with a microVM setup is as good as it gets right now, although I believe that we need more fine-grained tools down the road. It sounds like Microsandbox would be the perfect fit for what you are describing. I also built my own coding agent workbench on top of it (https://github.com/isolade/isolade). Microsandbox is quite cool, check it out: https://github.com/superradcompany/microsandbox
It has network filtering + placeholders for secrets.
OSS, no logins needed
reddec 2 days ago [-]
I've put some effort to integrate it to my agentic workflow. The problem, however, with docker in smolvm: it work-ish (there is example), but quite hacky.
Another problem which I wasnt able to solve - persistent image without Dockerfile. CloudInit will be ideal.
Documention at this moment in an early stage.
Overall, its a great project but for me was simpler just use Virtual Machine Manager (libvirt GUI).
I wish all luck to the maintainers, but probably DX-wise I will prefer to have more granular or predictable controls (eg micro cloud from Canonical).
(Not affiliated with them, just tried it out last week.)
binsquare 1 days ago [-]
I'm aware of them!
Yep - similar in some ways but headed towards different directions.
I am building a virtual machine to simplify/replace container infra. Ex. we run containers inside of linux VM's even in the `cloud`, resulting in managing both the vm, and the containers.
But smol machines is a lightweight, portable VM that you can package into a single portable .smolmachine file to be rehydrated on any platform, kind of like how containers are used for today.
Sandboxing happens to be a feature of virtual machines, so we are alike in being used for sandboxing.
Support for running agent harnesses in unprivileged podman containers is on my feature list. :-)
pojzon 1 days ago [-]
Docker containers are not enough isolation for anyone that cares about jailbreak scenarios.
Only real alternative is to use microvms. My goto solution for this are apple/containers.
Supermancho 1 days ago [-]
> Docker containers are not enough isolation for anyone that cares about jailbreak scenarios.
For the vast majority of developers, containers are enough, which is why they are ubiquitous while vms are less common. Ofc that ubiquity has led to lazy configuration, which is how the jailbreaking can occur. Knowing what you are doing with containers is a requirement to use containers as an AI sandbox.
zmmmmm 22 hours ago [-]
the big thing containers don't allow is for the agent to run and use docker itself without compromising the host
I'm not sure where "vast majority" cuts in but I would say a huge number of developers use docker and it is inconvenient at best if your AI harness can't actually run and test the infra it is building against
wbl 1 days ago [-]
The AI launches new kernel bugs as matter of course.
pjmlp 1 days ago [-]
There were not the solution for a while now, that is why Kata containers came to be in first place.
TacticalCoder 1 days ago [-]
> Only real alternative is to use microvms. My goto solution for this are apple/containers.
Why microVMs? I never ever run a container, AI harness or other, in something else than a full on VM. I could use a microVM but in any case I really don't see why I'd run a container on one of my bare metal OS: the place of a container is inside a VM (or microVM).
Especially for AI harnesses where the threat of an escape is very real: the more defense in depth, the better.
And If I can use rootless Podman instead of "rootfull" Docker, the better. Most of my containers are Podman btw.
> My goto solution for this are apple/containers.
To each his own: my goto solution is an actual server on my LAN with shitload of cores and memory and plenty of scripts to provision VMs etc.
I really don't understand why people are YOLO'ing containers on their bare metal OS.
graemep 1 days ago [-]
> I really don't understand why people are YOLO'ing containers on their bare metal OS.
You know about the people who do not bother with the container? Quite a few make “No problem so far” comments on HN discussions.
sgc 1 days ago [-]
I agree. I thought everybody knew to never use docker for high security, because it is "security lite". Might as well just use firejail. I presume that an agent knows more about networking and virtualization than I do. The only real solution is using multi-tenant level vm isolation, while presuming that the agent still might break out of their vm. So the vms need to be hosted on their own physical box that only runs the kvm provisioning host (or similar), and is firewalled on its own isolated network. It's a bit of a pain of course, but anything less feels almost like security theatre rather than meaningful to me. Otherwise you need to stick to the remote chatbots only.
cpburns2009 1 days ago [-]
The nice thing about running a microVM like a container is the interface is relatively easy, especially if you've used containers before.
1 days ago [-]
mirmor23 1 days ago [-]
[dead]
384028345 1 days ago [-]
Isn't Nvidia's openshell exactly what you're looking for?
I'm asking because I'm just learning about this stuff myself and tested openshell yesterday with pi for the first time.
I have a colleague who's using openshell. The advantage is it's independent of the containerization layer, yes?
SegmentTree 2 days ago [-]
Eclipse Enclave does exactly that: There is an outbound firewall and secret injections, so that the agent never sees a real key. And it's fully open source: https://github.com/eclipse-enclave/enclave
dvtkrlbs 2 days ago [-]
Looking at the Readme it seems like it only supports docker. Which is a dealbreaker for some
collabs 1 days ago [-]
What is missing in qemu + podman that we need rootful docker for this? Is there actual capability that is missing or is it more of a design choice by the eclipse enclave folks?
pbasista 1 days ago [-]
I use Linux Containers managed by Incus for working with Claude.
I have a dedicated container for that. It can run its own Docker daemon and other system services if needed.
Apart from the Claude login token, it has no SSH keys or other credentials. I push everything I need to it from the local machine. And I pull the Claude generated outputs from it.
Of course, this kind of setup requires a stack which can run or at least be tested without any credentials.
Tomte 1 days ago [-]
I do the same, but with pi.dev in Incus, mapping a project folder into the VM.
What I don‘t have compared to sbx is an outbound firewall, but my VM does not have any personal/interesting data, only a vanilla Fedora installation and the project dir with open source code, so I do not care much about exfiltration.
radlad 1 days ago [-]
Love Incus and I'm using throwaway restricted projects for testing. Highly recommend incus-windows if you need to do any Windows testing. Having agents validate Windows behavior has reduced so much toil for me.
We run https://github.com/NVIDIA/OpenShell on our workloads and we like it because it is backed by NVIDIA and fits natively within our k8s env.
It also provides a nice TUI for network policy management and prebuilt sandboxes which have claude code, codex, etc.
cpburns2009 1 days ago [-]
Gondolin looks interesting. It sounds like a TypeScript wrapper that achieves the same thing as my setup: Docker & Kata Containers 4 (KVM/QEMU backend) for microVMs, iron-proxy for egress and secrets, and dnsmasq for internal network name resolution (workaround for a Docker/Kata incompatibility).
Has egress and ingress filtering, egress can be bound to host/internet/subnet or even better to internal apps (which are each separate netns) meaning you can do your own firewall/vpn/whatever per sandbox. Plus you control what other components in the sandbox env the app can communicate with.
Really not built for day-to-day dev work though, more like automating your company/life / getting rid of SaaS (e.g. for technical Founders / Sales etc, not exactly useful for dev work)
chrisweekly 1 days ago [-]
https://smolmachines.com has "smolvm" microvms, for better security. The DX is whatever you decide to do with it.
notsirius 1 days ago [-]
Ooh - can you share more about your setup with superset? I tried getting it integrated with superset a while ago with no dice.
rusch 1 days ago [-]
Currently running codex. I run one sandbox per repo. So I create the sandbox in the repo root. Then it's a custom terminal preset:
sbx run --name "yoursandbox" -- --cd "$PWD"
This boots a sbx session in the worktree directory.
For Claude there is no --cd so it's more hacky, but I solved it by creating a sbx kit with entrypoint script that reads a flag (e.g --cwd) from the terminal preset command and then inside the sandbox cd's there and starts claude.
Def makes integrating easier but I try to avoid using direct mode for security (exposes .git folder, though there's probs a better way to protect it by disabling hooks or something)
rusch 6 hours ago [-]
Yes correct. We don't use git hooks so they are globally disabled. I also see they added host worktree mode which could work well with superset since it creates the worktrees.
It's an OS-level sandbox, though. It doesn't launch VMs or containers for sandboxing purposes; it uses whatever sandboxing features your host kernel offers.
grudg3 1 days ago [-]
Wouldn't say 'better' alternative, but I worked on making my own setup that I can trust by implementing a pi extension that leverages smolvm and agent-vault. The VM tooling is controlled by nix flakes. I can't share the source code (developed on company time), but I have a 'spec' of the whole thing, which you should be able to feed to your agent to replicate - https://gist.github.com/mahalel/c4e984292ff90bd4e11269555158...
We run cloud sandboxes, and have some experimental local sandbox support that is fully OSS.
Main thing for amika.dev is you can control the sandboxes and agents interchangeabley by SSH, web, or API, and can expose the services the agent is working on over signed URLs
Entire sandbox config is a TOML file
We're going to improve the local OSS sandbox mode and add better network controls over the next couple weeks.
Ultimately, what we're building kind of like if Tailscale and Firecracker had a baby, with a messaging protocol for remote controlling any sandboxed agent
It's free to try out. Still a lot to build, so we really appreciate any and all feedback about what we should focus on
Internally uses a single VM + Incus containers, supports docker/Kubernetes in each sandbox, has integrated worktree management.
nzjrs 1 days ago [-]
If you just need a python+venv sandbox with dev-first UX, no container build step needed, and no startup cost then I am using https://github.com/nzjrs/sandbubble in prod.
I havent used nor gondolin neither docker's solution, but curious to know what gondolin is missing (evaluating both for my personal use)? is it only the DX or something else, if DX, can you what exactly is missing?
thanks
codethief 1 days ago [-]
In my experience it's mostly the UX/DX where Gondolin is lacking. For instance, I don't want to set up a JavaScript project every single time I need a sandbox. Instead, I just want to place a config file somewhere in my repo or my home dir and be done with it.
You can run the agent in the gondolin sandbox if you wish.
Their example implementation with pi uses a pi extension so that pi runs on the host but the read/write/bash/etc tools run in the guest. Doesn’t have to be that way though.
olejorgenb 1 days ago [-]
Note that the pi-extension example in the gondolin repo is very outdated. Look in the pi repo instead.
Does secret injection really prevent that the agent send my GitHub key somewhere? If it has access to it via env var, can it not just paste it somewhere?
rusch 2 days ago [-]
The env var is just a placeholder in the VM, so no real secret is in there.
llimllib 1 days ago [-]
right, but say you give the agent access to github and it can push as you, or make a gist; now it can easily exfiltrate your secret.
And that's just an easy case - really if it has any network access at all it can come up with a clever way to route a request through the network such that the key comes back somewhere in the request. If you scan for it inbound too, the machine can obfuscate it.
Our agents are trained to be so intensely helpful and they have such intricate knowledge of how things work that they will do some incredibly clever tricks to do what you ask them to do.
skinfaxi 1 days ago [-]
The agent has no access to the secret. It has a placeholder that is replaced at a higher level. When it makes the network request the secret is substituted but that is outside of the caller's worldview.
ruszki 1 days ago [-]
But how? Normally the TLS handshake and encryption/decryption happen in user space. Even the kernel doesn’t know anything about it.
So, if the program or the proxy solution doesn’t support it, then it doesn’t work? Like with security solutions?
15 hours ago [-]
toksdotdev 15 hours ago [-]
microsandbox maintainer here. the custom certificate is installed in the guest's trusted root CA list, so it should work across any program, except where the program opts to explicitly pin certificates for a destination.
icedchai 19 hours ago [-]
The docs mention it can be bypassed for configured domains.
ruszki 17 hours ago [-]
Yes, I read it. That means that it doesn’t work in those cases. Btw, as a developer it’s very easy to have something like that. It’s not as trivial as it seems at all. I encountered with similar problems all the time, with similar solutions (mainly for security theater reasons) in the past. There are websites which simply doesn’t work if you replace certificates, regardless of browser or CA for example.
icedchai 9 hours ago [-]
Yes, I've worked with people who have run into issues with "security" solutions like ZScaler. I have tried it with some APIs (like GitHub) and it does work. Not to say it will work in your case.
ruszki 6 hours ago [-]
I was just interested how it works, because above it was sold as “it works”, when in reality, “it works*”.
icedchai 2 hours ago [-]
Isn't it that way with most software? "It works", except when it doesn't.
So what stops it from sending a network request to a git repo that pushes what that placeholder resolves to?
skinfaxi 1 days ago [-]
How would that work? You don't control github.com servers so your repo would never see the secret.
edit: You may want to look into tokenizing proxies as the general application of this concept.
llimllib 1 days ago [-]
Your agent writes secret.txt with the placeholder, and the tokenizing proxy replaces it with the token, then the agent reads secret.txt
skinfaxi 1 days ago [-]
It only replaces the token in the HTTP header that is sent to the server. Whatever you wrote in your files isn't touched by the proxy.
stavros 22 hours ago [-]
It sends a request to requestb.in and reads the public log of the headers. There are ways.
skinfaxi 8 hours ago [-]
But requestb.in is not api.github.com so the proxy wouldn't replace anything.
messh 1 days ago [-]
Couldn't it then just publish the mock in a public place... it would get replaced by the real secret.? How is this prevented
drdexebtjl 1 days ago [-]
Maybe the tokenizing proxy could work both ways? If the agent tries to read secret.txt, it gets back the placeholder.
1 days ago [-]
eli 2 days ago [-]
It’s injected into an outbound api call, not into an env var the agent can read.
kellpossible2 1 days ago [-]
what's to stop an agent creating an outbound call with the var to a malicious endpoint? (unless you whitelist what it has access to)
mpern 1 days ago [-]
At least for gondolin and microsandbox, you bind a specific secret placeholder to the target host. i.e. your GH token is only replaced/injected for calls to api.github.com, not other hosts. And you can set up both with deny-by-default
messh 1 days ago [-]
But then... why is replacent needed at all: just use sone permissions system.
messh 1 days ago [-]
Couldn't the agent post then the key in some public comment?
roywiggins 1 days ago [-]
Only if the replacement is global and not, say, only looking and inserting it into the actual (eg) Authorization header. If something is only transparently altering the Authorization header, then an agent inserting the dummy value somewhere else is totally safe.
mikesir87 1 days ago [-]
For clarity, there is no "search and replace" function going on. It's only setting the header.
The main reason a "proxy-managed" env var is set is because most CLI tools assume if the env var is set, auth is set. If the env var is unset, it will assume auth needs to occur. Fortunately, most don't do a pattern matching on what the value actually is.
eglintondust 1 days ago [-]
You just don't inject the real secret unless hostname/whatever rule matches the request, right? I don't know if that's how this works but it's my assumption.
mikesir87 1 days ago [-]
On the Docker DevRel team... yes! This is it. The secret is injected only into headers in which the hostname matches.
There's also an ability to create kits where you can setup credential injection into other services as well.
skinfaxi 1 days ago [-]
The replacement is on a url/host basis.
kellpossible2 1 days ago [-]
or an outbound call to a trusted endpoint with the env var in a way that can get exposed to the agent via a subsequent call?
skinfaxi 1 days ago [-]
How would that work, exactly? What are you envisioning?
radlad 1 days ago [-]
It's possible reflected instances are masked too, like GitHub Actions. But I don't know.
petesergeant 2 days ago [-]
What specifically do you want? I have:
https://github.com/pjlsergeant/byre -- slightly different security model, but lazer-focused on developer experience; my daily driver and I love it not just because I wrote it. The TUI is great for configuring and setting up instant boxes just how you want
https://pleasedonotescape.com/ -- a list of every other agent jail I could find, filterable by open-source and whatever else you want
What I run is one hardened QEMU/KVM VM per project holding the whole dev environment (editors, agents, containers), with nftables on the host allowing internet egress but dropping anything aimed at the host, the LAN, or any other private address, plus an allowlist for deliberate exceptions.
Basically, it's a plain QEMU/KVM VM on a stock Debian cloud image: device model stripped down to a virtio disk, a virtio NIC and a serial console, nested virt off, no passwordless sudo in the guest. It also ships a containment check that scans outward from inside the guest, so the network boundary is something you can verify.
Wrapping the whole environment rather than a single agent session puts supply chain attacks inside the boundary too. A poisoned npm or PyPI package, or a compromised editor extension, lands in the VM instead of on the host. That was the original reason I set this up; agents just made it more urgent.
There's no per-domain egress allowlist; the policy is "internet yes, private addresses no". Secret injection isn't built in either, though Infisical's agent-vault on the host as an egress proxy covers that part.
Wrote the whole setup up here, in case it's useful:
I currently do something similar, but this article was a nice read and gave me some new ideas.
TacticalCoder 1 days ago [-]
Yeah I discovered your blog a few days ago: I've got a setup not unlike yours.
> So rather than pick one, this post advocates layering both, in the spirit of defense in depth: a sandbox VM wraps your containers along with the whole toolchain, and that sandbox reaches the internet but has no route to anything private.
Yup it's the only proper way.
And that is true not just for AI harnesses/agents (that shall try to escape), but also for stuff like Plex/Jellyfin/Immich/private pastebin etc.
If you care about security, there really simply is zero reason to run containers on the bare metal.
1 days ago [-]
schmitthub 1 days ago [-]
[dead]
1 days ago [-]
1 days ago [-]
Grimburger 2 days ago [-]
> Each agent runs inside a dedicated microVM with your dev environment
What's a "microVM" and what's the security model here compared to using real virtual machines with actual constraints on breakouts?
“microvms” are real vms but the hypervisor and vm (guest kernel) shed most of the hardware / device emulation, support, and discovery which makes traditional VMs look / feel like real computers, as well as most guest interactions. This gives them extremely low overhead.
Firecracker is designed to start a VM in under 125ms and 5MB. Netbsd advertises that you can direct-boot a MICROVM kernel configuration in under 10ms.
jsiepkes 2 days ago [-]
If an agent fires up NPM, takes a boatload of memory, is that memory released back to the OS after NPM shuts down in the VM?
arendtio 2 hours ago [-]
> If an agent fires up NPM... :D :D
I think this is one of the big issues with today's MCP servers. Most are based on Node.js and take a lot more memory than they should, compared to the complexity that the job requires. Just run a few MCP servers locally, and all your RAM is gone...
dist-epoch 2 days ago [-]
In principle yes, in practice it's complicated, using something called "balloon drivers"
There’s also memory hotplugging via virtio-mem. But generally speaking downscaling live vm memory can’t be said to be a solved problem, it’s more of an active area of research.
GautamTalksDev 1 days ago [-]
[flagged]
TheqO 2 days ago [-]
There are many devs that have little to no experience of Linux, like the hundreds of thousands of .Net and Java CRUD devs in enterprise companies using Windows.
There is a need for a Docker desktop like GUI for this market.
newsoftheday 1 days ago [-]
Java CRUD devs have been getting into Linux a lot more in recent years since learning more about how DevOps does their work is important and integrated in with our projects, like the Dockerfile, Helm files, etc. I was already into Linux before Java was even released. I'd say even .Net devs are too since .Net Core is becoming more important in their world.
Izmaki 1 days ago [-]
When I was a teen I had no experience with software development. Then I went to school and learned about it. Now I'm making money doing this thing because I became quite good.
The world didn't fit around me, so I made "me" fit around "the world". I bet those stubborn dinosaurs can learn a new trick or two also, if management lets them do it during working ours...
Also, WSL (Windows Subsystem for Linux) has been baked into Windows for a long time and makes it very easy to play with Linux, as does using the Hyper-V VM system. Any developer unfamiliar with Linux because they use Windows, has little excuse.
pjmlp 1 days ago [-]
Agreed, even game devs that ignore Linux as target for their AAA games, actually tend to use Linux for game servers, the age of IIS with .NET/ISAPI is long gone, except for legacy code stuck in .NET Framework.
Modern .NET did not went cross platform by accident, and Java development has always been "develop on Windows deploy on UNIX", in corporations where Mac tends to have little presence.
1 days ago [-]
venatiodecorus 1 days ago [-]
docker's sandboxes are cli only atm
huflungdung 2 days ago [-]
[dead]
frio 2 days ago [-]
It’s real VMs, firecracker style.
Grimburger 1 days ago [-]
Haven't used docker sandbox but you can't just `apt install postgres` on firecracker, it needs to get baked into the image first.
That's my experience anyway, there's a lot of restrictions once you need to do some real basic things. For basic prompts maybe but interacting with a full stack ehh.
So bit hesitant to call firecracker a real VM myself.
vdfs 2 days ago [-]
Your example is not complete, you have to show how it will run claude/codex, you have to do extra things to install run and mount folders there, this one does that with less config, also with this agents can run docker, lxd doesn't allow you to do that
Grimburger 1 days ago [-]
You create your own image first rather than blank ubuntu.
packer init ubuntu-claude.hcl
packer build ubuntu-claude.hcl
incus launch ubuntu-claude my-claude-vm
incus exec my-claude-vm -- claude -p "solve the EC discrete logarithm problem, if it doesn't work keep going" --dangerously-skip-permissions
The reality of these things are that eventually you will want to do something useful or different with them and the flexibility simply isn't there compared to a real vm, if you desperately need boot times then there's plenty of other options here, especially with packer. Some of my prompts are often hitting 60+ minutes so it's not really something I think about.
dist-epoch 2 days ago [-]
An Ubuntu Server VM, like the ones started by Incus, use at least 512 MB of RAM per instance. If you spawn 10 sandbox VMs, you already pay 5 GB RAM just to sit there idle. You also pay a CPU cost, you have 10 kernels managing stuff, but arguably it doesn't matter that much given CPU core counts.
I use something in between - a single Ubuntu VM, into which I spawn multiple Incus LXC containers for the agents. The containers only use 50 MB or so per instance (separate systemd, ...). This way I pay the VM RAM tax only once, and the agents are still contained inside the VM if they manage to escape the LXC containers.
apitman 19 hours ago [-]
This is what I do. Nice benefit is it lets me passthrough my GPU and share it between multiple containers. Incus is awesome.
TacticalCoder 1 days ago [-]
Yeah this basically. I differentiate between "containers I wrote" (where I packaged the app/wrote the OCI "Dockerfile" / container file) and "containers from other people": all those I wrote (for our own use) go into one VM, while all the other containers go into another VM. Then I've got a third VM for containers for the AI agents.
This way I don't pay tens of VMs "tax" but basically only three (plus one or two VMs I use for testing enhancements to my VMs provisioning / optimization / securing setup, when I work on that).
I don't use LXC (I could) but regular containers, inside VMs.
I'll look into the stripped down "micro" VMs but then I don't spend my days launching VMs/shutting them down so it's not a big deal.
dizhn 2 days ago [-]
That's a full VM. Microvms are much smaller and they start up very very fast. In miliseconds.
randomint64 2 days ago [-]
[dead]
Roark66 1 days ago [-]
How about implementing proper permissions on the tool use or if you need more flexibility a dedicated model to analyse potential impact? (Like Claude Code's Autopilot but more configurable)?
I find solutions like this to be a like trying to patch a leaking boat on a lake with duct tape. It will help, but it's not a proper solution.
Also, often the tasks you want the AI to perform are in the outside world. Like "connect to my servers, and figure out X and Y".
The proper way is permission isolation. I run a small k8 cluster in the homelab and I have 3 types of pod/agent combinations for my AI agents. Read only, one that can change my gitops but it needs to create PRs that admin approves, and admin.
Likewise with code. I have a forgejo git instance where agents have ability to create feature branches and so on, but merging is gated.
Those things require "GH enterprise features".
In fact more and more things we do at home will require "enterprise features". Why? Because a person with AI is basically a small team, but some of team members behave like Chimps on crack... So security must be top notch.
AlotOfReading 1 days ago [-]
What are the proper permissions for an agent? An agent shouldn't be able to read ~/.ssh, but that means a bash tool that spawns `cat` is different than one that spawns an ssh client. I don't allow my agents to use git commit, except sometimes I ask an agent to split up a complicated branch that I can't be bothered to split myself. rm'ing intermediary files is fine, but rm'ing committed files is bad, unless the agent has done *.bak renaming and is cleaning up itself, etc.
I don't think "proper" permissions are possible without dramatically limiting the way people use these tools.
iury-sza 1 days ago [-]
Implementing that is trivial in the harness side. You code vibe that in minutes.
victor_edka 1 days ago [-]
yea...I think k8s is de wae for running proper proper rbac sandboxes for agents.
kwakubiney 1 days ago [-]
Any idea of anyone exploring this space? Sounds interesting
So with your solution you get the additional security benefit of containerizing the hypervisor on the host.
figmert 18 hours ago [-]
Once you have a vm, the container provides next to no additional security benefits. It's just unnecessary overhead at that point.
codethief 14 hours ago [-]
That's not correct. Virtio devices have different security properties and many of them expose the host system to considerable risks. Using containerization on the host is one way to limit the latter. See e.g. https://github.com/libkrun/libkrun/#security-model for more details.
figmert 8 hours ago [-]
Well, I stand corrected! Thanks for the link
hokkos 2 days ago [-]
Wow, I hope one day Linux will be able to support the exclusive MacOs/Windows technology of Docker Sandboxes.
(it's in the doc, but kinda strange to not see some instructions on the main page, probably distro related)
Yes, bubblewrap is superior to Docker for this. I wrote a tool to use bubblewrap for the purpose. It needs a tool to start it, or is at least much more convenient with a tool, because you need to take your session/auth data into the container, and if you want the agent to be able to start containers (agents love containers) within the container, you need some config magic mounted inside. You could manually do all that, or do it with a shell script, as well. But, this is how I did it, and you're likely to run into all the same little quirks I ran into:
Interesting. I’ve been trying to use bwrap, slirp4netns, and mitmproxy to create a simple Python script to get save shell for development. But it’s a huge time sink (and I might resort to podman)
I’ve been using this pretty extensively for a few months on Mac and Linux and have been super happy with it.
ethagnawl 1 days ago [-]
Thanks. That omission didn't smell right based on everything I know about Docker. Curious choice, indeed, not to show Linux install instructions.
brettpro 15 hours ago [-]
When I tried this a month or so ago Linux support was markedly bad, and a quick look at the GitHub issues confirmed it wasn't just me and wasn't a priority for the company. I wouldn't advertise it either.
The nails in the coffin were 1) login was required 2) login was broken because they "didn't consider" it would be run in a headless environment [0] and 3) they shipped with hardcoded binary paths and root requirements [1].
I moved on and use Incus directly with small helper scripts, smolvm, or a full fat VM running desktop Claude or ChatGPT if I (or someone I mentor) really needs the full app. I'm not yoloing every new claw agent in --dangerously-destroy-my-things mode, so network restrictions are best effort, though filesystem access stays tight.
Either way, docker sandboxes really didn't seem to be it, and the company didn't seem interested in trying to be anything beyond an enterprise solution.
I got excited for this not because this didn't exist before, but because Docker putting their weight on this would imply a broader adoption and better integration in the industry. I am sad that they are asking for a login here though, which doesn't make any sense to me.
KolibriFly 2 days ago [-]
That's docker, man. Tomorrow they're gonna add limits on sandbox runs without a premium account too
rvz 2 days ago [-]
microVMs (firecracker) have existed for years. This is not new.
PufPufPuf 1 days ago [-]
I built "Locki": something similar but open-source! A bit different approach -- single VM with Incus containers -- focusing on speed of spinning up new sandboxes and integraton with git worktrees. The core grievance that motivated me was the lack of docker/kubernetes support in existing sandboxing tools, with Locki there's no chance of footguns like "two agents rebuild :latest tag at the same time". Give it a try: https://github.com/JanPokorny/locki
katspaugh 1 days ago [-]
Looks cool!
I took a similar approach with https://runmachine.dev/ but later switched to OrbStack for iOS development.
It blocks network access by default, mounts only what you specify, and you can add a customization layer. This is all done in the container itself (srt for network blocking). It doesn't implement a central point for secret sharing, MCP exposure, etc. So it might not have enough features for some but it works well for my needs.
I just found through this thread yoloai which has an apple container backend, so its quite similar using that. My main issue would be that network access is allowed by default. https://github.com/kstenerud/yoloai
Several other projects listed here use libkrun which is an alternate implementation that works with Mac's HVF. smolvm, microsandbox, podman (with likrun backend), gondolin.
navigate8310 2 days ago [-]
I just made my own devcontainer that I copy on any project and load whatever harness I want in that repo. Harnesss' config and auth are simply mounted from the host, so no setup required at all.
I don't get how it works, or works well -- presumably those domains are behind CDNs, and IP addresses are unpredictable. So the firewall allows traffic to specific IPs that are resolved at the time the script is run, but not after that? What if the same domain is resolved again without going through the cache, and it becomes a different IP?
And even if that works, this is a very short list. As soon as you reach for Go, Rust tooling etc nothing works. So you need to manually maintain this list which is nothing but painful trial and error.
ai_fry_ur_brain 2 days ago [-]
[dead]
iamspoilt 1 days ago [-]
I recently wrote a blog post on using Tart for Macs for something similar that docker is doing here but with better persistence and control. My take is that the Tart approach is superior to this as it gives you a full dev machine with a single command line that allows agents to access files on host, install packages, maintain the vm and do whatever they want to do without compromising the host OS.
Haven't tested it yet, but it seems to address the same issue as Docker Sandboxes, but in a different way.
speedgoose 2 days ago [-]
I have tested it and the big advantage is that is has access to the local development tools.
But it’s not as well sandboxed for sure.
lemontheme 1 days ago [-]
Nono has been my daily driver since the start of the year. It's not a perfect sandbox -- that's for sure. For example, the default network rules let you escape via a global TMUX server. But it is extremely practical. It gives me enough guarantees to feel confident about running in YOLO mode. So far nothing has gone awry.
fg137 1 days ago [-]
As you as your Go build fails because you haven't put the local cache dir in the "allowed directories", you'll understand how painful this is, as well as most tools based on bubblewrap/sandbox-exec. There is a difference between a clean environment with standard setup vs a layer on top of everyone's existing tools/setup, especially in a enterprise environment.
(I'm sure you can spend time to come up with a proper bubblewrap configuration that allows go build to succeed, but it's probably not worth the effort.)
LeBit 1 days ago [-]
Why do you say that?
Eg, if used with Colima in macOS, it means I can run a devcontainer in an isolated VM and Nono inside the devcontainer can restrict a lot what can and cannot be done.
You get credentials proxying and network outbound limits.
How is Docker Sandbox better sandboxed?
speedgoose 1 days ago [-]
Yeah but that’s Colima and Nono then. Not only Nono.
isityettime 1 days ago [-]
True. But it's also an illustration of how relying on an OS' native sandboxing capabilities is nicely composable with other isolation techniques.
dhchun1203 1 days ago [-]
Neither of the recent ones was actually a container escape though. The OpenAI
one in July found a misconfig in the sandbox network, and Kimi K3 last week just
walked out to grab answers off GitHub during an eval.
Both went through stuff the sandbox was set up to allow.
codethief 1 days ago [-]
To everyone sharing their favorite container-based sandboxing solution: Docker Sandbox does not use containers for isolation. It spawns the workload in a libkrun-based micro VM, which has vastly different security properties.
Do you mean like Docker has supported for years…? (Just configure krun as Docker's OCI runtime.)
Obviously, there's a reason why Docker released Docker Sandbox as a separate product:
- Barely anyone bothers to configure Docker/Podman with a different OCI runtime like krun. Heck, most people don't even know about OCI runtimes in the first place. Case in point: Most people here in this HN discussion are proposing using "standard" containers (with the default OCI runtime) for sandboxing. This is what I was trying to get at.
- A sandbox for agent needs tighter network control.
As for differences between the krun OCI runtime and Docker Sandbox (which also uses libkrun), let's please continue the discussion here: https://news.ycombinator.com/item?id=49240662 .
I have a solution based on Nix that can be used to generate reproducible container images: https://github.com/nothingnesses/agent-images . It lets you customise which agents, harnesses, or any other packages you want included in the VM and it uses `agent-box` for sandboxing.
binsquare 2 days ago [-]
wonderful, will try to test this in smol machines as well
dizhn 2 days ago [-]
This looks like gvisor but is a vm like firecracker right? Any reason you did not want to use firecracker?
(I am testing this now as a backend for my pet project which currently supports firecracker and gvisor. No network.)
binsquare 2 days ago [-]
It's a batteries included alternative to firecracker with a couple of new ideas tossed into the mix i.e. portable like a container (bake into a single file and rehydrate the vm anywhere), dynamic resource allocation, etc.
Alifatisk 2 days ago [-]
What? Does using sbx require login? Bummer.
dSebastien 2 days ago [-]
Yes and they have a specific subscription for managing sandbox policies across the enterprise: Docker AI Governance
dSebastien 2 days ago [-]
You can create those manually but if you want to enforce those then you need the subscription
reddozen 2 days ago [-]
If any AI company was doing serious engineering isolated containers would have been a prerequisite to using their tools.
globular-toast 2 days ago [-]
Anyone serious about security will want to bring their own sandbox anyway, not trust these, often proprietary, agents. I've never run an agent outside a sandbox. My first bubblewrap script for `claude` is now over a year old. The tools are available and if you learn to use them you can run any program in a sandbox.
But, in any case, why put in effort doing something people don't expect or ask for? We can assume everyone running agents is either a) using their own sandbox, or b) doesn't care. I think we can guess which category most people fall into. You could maybe argue about responsibility, but I don't think you can argue about "serious engineering".
topspin 1 days ago [-]
> Anyone serious about security will want to bring their own sandbox anyway
Exactly. It's not as though it's difficult. It never occurred to me to not do this from day one, and it astonishes me that anyone runs this stuff bare metal. Since then, I've brought several other people on board, and that's all they've ever seen: I don't think they'd know how to run outside a sandbox, and that's just fine.
So what "sandboxing" does this add that is not already present in Docker, and how can users be any more assured that software cannot break out (which has happened at times with Docker).
Can a user blindly trust this sandbox, because that's how people will treat it based on the marketing. Sounds like it could be useful for far more than just AI though.
codethief 1 hours ago [-]
> So what "sandboxing" does this add that is not already present in Docker
Docker Sandbox spawns a micro VM, not a standard container isolated by host kernel mechanisms (Linux namespaces etc.)
mikedelfino 1 days ago [-]
Since everyone is sharing their setup, here’s my approach, just to give people an idea of how others are doing it, however impractical it might look: I run a full Linux VM (with a GUI) on my Linux host. I connect via virt-viewer to run Claude Desktop, as I’m not a fan of using the terminal for this.
The VM sits on its own libvirt network in a dedicated firewall zone, and specific directories are shared via filesystem passthrough. To keep the agent from accessing anything related to Git, the actual gitdir is stored on a separate path outside the mount point.
I review the git diff manually and commit it from the host.
everforward 1 days ago [-]
Do you find the permanence of a full VM useful? I’ve wondered about something like this but always defaulted to Docker for much the same reasons people use stuff like Ansible. I’m afraid the LLM will heavily customize its environment and I’ll be unable to replicate it when my laptop dies or I can’t upgrade the OS or whatever.
Then again, I guess GUI is a pain in Docker. I tend to operate through Zed and an ACP harness though, so my GUIs are sort of “inside the container” anyways.
geoka9 1 days ago [-]
My understanding is that the OP is not running any heavy tools/chains inside the sandbox? I use a similar setup, but using Incus and cli agents which I drive via ssh. I make sure I can (and do) rebuild the VM from scratch after every session or so.
Currently considering using ACP for codex so that I can do more of the driving from my editor (emacs, over ssh) and something similar for claude code (it doesn't seem to be as good as codex at supporting re-attachable sessions).
One concern is making sure my editor's ACP client doesn't enable/support fancy terminal stuff, because that would basically void all the benefits of using a VM sandbox.
Grimburger 1 days ago [-]
> I’m afraid the LLM will heavily customize its environment and I’ll be unable to replicate it
Would suggest Hashicorp Packer or cloud-init for deterministic images, not hard to setup or use. LLM's have little problem with them either I find.
If you need quicker environment rebuilds consider using something smaller like alpine as the base, though once you setup a golden image even heavy things like debian are fine.
mikedelfino 1 days ago [-]
The VM is allocated 2 cores and 4GB of RAM. For my workload, it doesn't feel any slower than running it directly on the host. The few times I've checked memory usage, it wasn't anywhere near full as far as I remember.
pkhamre 2 days ago [-]
I started building my own isolated and security-hardened docker image for OpenCode about half a year ago. Been using it daily.
if I was paranoid about security I wouldn't use docker in the first place.
d2p 2 days ago [-]
Does this support Linux yet? When I previously looked it did not (the reason being that they were already using VMs on Windows/macOS but not on Linux). Every time I see an announcement I think "great, they must've added Linux now then", but the linked pages always have Windows + macOS instructions but not Linux.
All the open GH issues about supporting Linux that I subscribed to have gone unresponded to.
OpenShell looks like a good alternative, but it still has "Do not use in production" plastered all over the website, which doesn't fill me with confidence yet
narinciye 1 days ago [-]
It must be a joke that this tool is not supported on linux yet, although docker is built on top of linux containers. Shame on docker.
I wrote a tool to use `bubblewrap` to containerize any agent (at least all the agents I've used a couple of times), and bind mount the system stuff read-only, so the agent has your "usual" environment, but they can only see the project. Their history persists (either through a bind mount or a "shadow" copy of the history that only the wrapped agent sees), the agent can still create and manage containers of its own using podman's rootless mode, etc. It's nearly instant to start because it's just a namespace (plus a few copied files for the container support and session history); no container needs to be built/fetched/updated/whatever. bubblewrap is extremely well-tested as it is used by flatpak and several other large projects, so I trust it quite a bit (more than I trust Docker).
Bubblewrap is not nearly as secure as a proper VM.
fg137 1 days ago [-]
bubblewrap may work well for you and your specific workflows/projects but not in an enterprise setting where everyone already has a different setup on the host and needs something different inside the container. It's impossible to deploy a solution like that with bubblewrap -- configuration itself is going to be a nightmare. Which is why Docker Sandbox is aimed at teams/enterprises.
SwellJoe 1 days ago [-]
Yeah, Podman would be a better basis for that kind of use case. I'd built an early implementation of `flar` with Podman first, but it was more annoying than simply having my regular dev environment instantly available in the container. But if you need a bunch of different dev environments, instead of just your usual one, then sure, a bunch of different custom containers makes sense.
But, Docker is rarely the right way to manage containers on Linux, IMHO.
fg137 1 days ago [-]
"Docker Sandbox" is not docker. Completely different (and almost unrelated) products.
I also hit the same issue recently. No Linux and no Windows on arm. AI sandboxing has a lot of options but none feel complete just yet. It's hard to commit to something, especially if reviewing tools to aide in company policies.
Regardless, I'm hoping something that isn't behind a login screen is going to win out.
This is not really an alternative if you care about the security of your host system. Docker Sandbox uses micro VMs for a reasons.
biehl 2 days ago [-]
Looks really nice. Would it be easy to make a qwen-cli wrapper?
nezhar 1 days ago [-]
Sure, I added an issue for this, so it will follow in one of the next releases
llimllib 1 days ago [-]
Operating systems ought to be providing us the utilities we need to safely sandbox processes (agent or otherwise), but they appear to not be interested in the job
pjmlp 1 days ago [-]
Apple, Microsoft, IBM, Unisys, HP, Oracle/Sun have done that for a while now.
llimllib 1 days ago [-]
Poorly! Apple’s facilities for this are the ones I know best, and they are woefully insufficient
pjmlp 1 days ago [-]
Well, the features announced at WWDC 2026 naturally are yet to be made available in a mature form.
dannyw 2 days ago [-]
I’d rather use another open source solution that doesn’t require a signup, and less likely to get rugpulled.
There is no reason to require a login for creating local mini sandboxes.
If you’re on Apple, native solutions like “container-machine init” come built in and are pretty good, if you’ll only be on Apple hardware.
yellow_lead 2 days ago [-]
I know some people want to run their agents when their computer is off, but I imagine a solution like this will be much more common than paying for a remote sandbox (i.e on fly.io or exe.dev), especially because it'll be free.
Though, they need to remove the login requirement.
dbmikus 1 days ago [-]
Agreed! I think the best user experience is:
1. You can run sandboxes locally
2. You can control them securely over the internet, for when you're on the go
3. You can migrate them to cloud VMs if you want
If I can toot my own horn, I'm trying to build that :)
Still early and the local sandboxes are experimental right now
Can someone more versed in Docker explain to me how this is different than building my own docker container from a Dockerfile for using Pi agent harness? That's what I do currently. I use Docker Desktop in windows as the backend for that.
Hugsun 2 days ago [-]
Docker containers use Linux kernel features to create an isolated environment, running on the same machine as docker is.
This creates a virtual machine, with its own kernel, and runs the container in there.
This gives stronger isolation and security guarantees.
eloisius 2 days ago [-]
I have the same question as GP. Your answer helps a little but not really. I might be naive, but I was under the impression that malicious code escaping a docker image and running amok on my host system was not something I should be too worried about. Especially if I run docker in rootless mode. Is that wrong?
For clarity I’m actually using podman, not Docker.
angry_octet 2 days ago [-]
Oh no, you should definitely be worried about that. Podman might make it harder to escalate to host root, or manipulate other containers, but it is still vulnerable.
Now I'm curious to know how hardened the Docket Sandbox orchestration interface is. I guess we can assume they have run Mythos against it for a few weeks maybe? It's unclear.
eloisius 2 days ago [-]
Unsettling. I mean, is there any reasonable way to develop software in 2026? I've already sworn off ever installing npm directly on my host. Containerizing everything is laborious enough, but running a separate VM for everything?
mihaelm 2 days ago [-]
You just need to work out the threat model for what you're working on. For trusted containerized workloads, where the attack surface is minimal, just containerization is fine. However, agents can do just about anything on your computer if you allow it and people aren't really shying away from `--dangerously-skip-permissions`, so better hardening (VMs, microVMs) is desirable.
cognitiveinline 2 days ago [-]
BS.
Unless we're talking 0-day/CVE, running an unprivileged container is as trustable as a VM. The only difference is how strictly you want to hold the memory/CPU bar. Infact on linux, containers are more lightweight than VMs.
So yeah, not "vulnerable".
angry_octet 1 days ago [-]
LLMs are great at finding 0-day, and people are rubbish at updating their containers and hosts to patch b-day.
Containers have access to the kernel ABI, and as shown in the latest kernel exploits, all the memory handling surface that exposes. The virtualisation interface, offering fewer services, is significantly harder.
Containers are obviously lighter than VMs, both to start and to schedule, but firecracker is pretty fast. gVisor pays overhead per syscall vs at startup.
So yeah, more vulnerable.
teravor 1 days ago [-]
gvisor's overhead is mostly IO. especially if you use the KVM backend.
cognitiveinline 1 days ago [-]
Got it. 0 days are possible so throwaway containerization. You should blog about it, will help millions of developers and companies. Heck, even consult with the hyperscalers - they will be riddled with their workloads.
angry_octet 1 days ago [-]
Who do you think created firecracker? gVisor?
sureglymop 2 days ago [-]
That depends on the runtime though. For example, libkrun lets you do this:
docker run --runtime krun hello-world
That starts/runs the OCI in a qemu microvm.
LeBit 1 days ago [-]
When the host is a Mac or window , docker always run in a VM anyway.
On Linux, you can run docker directly on the host, but you can also very easily setup a vm with incus and run docker from there.
angry_octet 2 days ago [-]
It's a VM.
meffmadd 2 days ago [-]
I tried Docker Sandboxes but last time I checked you could not configure custom volume mounts, making more complex setups impossible. For work I need two directories for context for the agent to have access to…
When starting a sandbox, you can specify the mountpoints you want. It just defaults to the current directory. You can also specify some of those mounts as read-only as well.
Example: sbx run claude ./ ../another-project:ro
2 days ago [-]
cyberpunk 2 days ago [-]
put them both inside another directory and share that? what am i missing?
meffmadd 1 days ago [-]
Of course, but that was not part of my workflow and I found it quite strange that this was simply not possible especially when docker-compose can easily do this
espadrine 2 days ago [-]
Models start going to extreme, damaging lengths to achieve ambiguous prompts[0]. Having good sandboxes is now a must IMO.
But sbx is a bit annoying to use with OpenCode for instance (which has zero sandboxing by default, unlike codex CLI or Claude Code). You cannot easily change ~/.config/opencode/opencode.jsonc AFAIK.
We run OpenCode currently, but need to improve the setup for users, so would like understand more of the issues you have sandboxing it, if you can share
weinzierl 2 days ago [-]
Before you use no sandbox at all use this or one the many similar projects but it's alway worth remembering that Docker is not a security boundary. It never has been meant to be and never will
become one.
cgroups are a mechanism designed for hierarchical organization and resource distribution. Against a malicious and capable actor, and that is how we have to treat AI agents, cgroups will not withstand.
Also, the kernal is an interface too big for what an AI agent needs and is therefore offering a gigantic attack surface completely unnecessarily.
codethief 58 minutes ago [-]
As the sibling said, Docker Sandbox is not based on standard Docker containers. It spawns micro VMs.
TheRoque 2 days ago [-]
Would you say podman is better, or is it the same as docker ?
weinzierl 2 days ago [-]
In general running containers rootless is better from a security standpoint and podman makes this much easier. So, yes.
This is not my main point though. Both are based on cgroups and cgroups are the wrong tool for the job.
TheRoque 2 days ago [-]
Wouldn't it need a super critical exploit, I mean zero-day vulnerability, to escape from that kind of sandbox ? And if you think further, then isn't that risk also applicable to pretty much any kind of sandboxing ?
weinzierl 2 days ago [-]
Container escapes are more common than you think. Common enough for AWS not to rely on containers for their serverless functions, common enough for Google to say: "Untrusted code shouldn't rely on the container security boundary [..]" [1]
The same is not applicable for any kind of sandboxing for two reasons:
1. The boundary is in the kernal’s own code, enforced by the thing you are trying to be protected from. -> Use a VM
2. The kernal is a gigantic attack surface -> Use gVisor
What if you use tools like bubblewrap or nono inside the container?
Say I want to use pi inside a container. If I wrap pi within a bubblewrap or within nono, how is that less secure than using a vm?
Also, I think most people run containers inside VMs anyway and not directly on their hosts (on Mac and windows you have to use a vm anyway).
akdev1l 1 days ago [-]
bubble wrap is just doing the same cgroups work
LeBit 1 days ago [-]
Depending on the configuration, bubblewrap can substantially reduce the attack surface.
It doesn’t change the fact a malicious process is still attacking the same kernel , but it can reduce what it can do to that vm.
dboreham 1 days ago [-]
Which is fine but this thread is about a feature that provides hypervisor isolation, not cgroups.
realexweb 2 days ago [-]
[flagged]
bob1029 2 days ago [-]
The sandboxing problem is perhaps the greatest justification for doing agent integration via existing human interfaces rather than low level shell access. Granting access to shell is a super obvious path (it's easy) so I can understand us wanting to fight for it. But we should consider the other paths as well before we make our final stand.
Automating browsers with LLM agents properly requires a lot more work than Process.Start into powershell, but the advantages can be immense once you have achieved integration this way. Incrementally maintaining this integration is generally easy because human users cannot tolerate rapid changes either.
It's a hell of a lot easier to convince management to adopt a robot that looks and acts like a human employee than one that looks like a combine harvester. The combine is far more efficient, but it is also totally indiscriminate. Nothing constrains its appetite except for the invisible fence imposed by GPS. The amount of infrastructure required to keep farm equipment from running astray is incredible. In the context of agriculture, the added complexity is definitely worth it. We don't want to have to recreate the same thing with our technology if it can be avoided. Sandboxes and security isolation boundaries are not things to aspire to. These are costs to be paid for admission to something more valuable.
CraftianAI 8 hours ago [-]
I would have put 'microVM' in the title. While reading it, I wasn't sure if it is based on microvms or just rebranded/hardened containers.
Also, what took them so long?
Anyway, I decided to try it in a VM. Got:
"You are not authenticated to Docker. Starting the sign-in flow..."
(Just to try it.) Joke's on me.
myshapeprotocol 1 days ago [-]
Disposable and isolated sandboxes are the right way to handle untrusted agent execution. Great architectural pattern.
aki237 1 days ago [-]
I'm currently facing this issue. I've resorted to implementing my own execution environment albeit limited.
It goes like this:
- bash script parser + interpreter (with hooks for things like file open, execute etc.,)
- wasm executor for execution.
- wasm implementations of common tools like coreutils, grep, sed etc., from the uutils project.
- wasm implementation of python by a VMware backed project.
- entirely virtualized filesystem using Go's io/fs.FS. (tmp dirs can be implemented using any backend)
Works like a charm for the limited usecase I have. There are definitely some drawbacks with threading and especially with preopens in wasm. But a cheap sandbox for simple file explorations and minimal computations.
garganzol 1 days ago [-]
I do not see any value proposition in this - if I need a sandbox, I make one with Dockerfile, Bubblewrap or virtualization. What I am missing? An enforced required login is a net negative value - it means rug pulls in the future.
On Linux I am using https://github.com/wrr/drop which configures bubblewrap based on a simple yaml config specific for the project path and it is enough for me.
I won't use anything requiring a login.
23 hours ago [-]
celrenheit 1 days ago [-]
I tried it and it worked great at first but I had multiple issues with it, the disk space usage was growing significantly, I need to login multiple times for each sandbox, it's closed source and not possible to customize to my need.
One other thing, I want to be able to handle multiple repos in the same sandbox and have a standard workflow around worktrees (one worktree per repo, all the worktree mounted in the VM).
I did try to use those for some stuff:
* Login requirement is something else
* It's closed source last I checked
* Pretty slow/unstable
There are many better namespace/container based options, VMs may be moderately more secure but when you more or less trust your agent and code you can do with lesser containment. And with the recent CVEs in kvm honestly there isn't a huge deal of difference vs namespaces.
(I'm building https://xbin.dev/ for some time now for managing my personal code/apps, a project which started specifically after Docker Sandboxes broke on me some time ago)
arscan 1 days ago [-]
Not a substantive comment on content but hopefully constructive feedback on presentation:
Holy moly, on mobile I was trying to read the example console screenshots/snippets and then it would just unexpectedly change. Took me a little while to figure out it’s some kind of carousel for the examples, and not more screenshots/snippets loading and pushing down content (or me going crazy). Please don’t do this on mobile sites, just let me scroll through the examples!
mikesir87 1 days ago [-]
Thanks for the feedback! Will pass it on to the web team to get this more mobile-friendly.
dorongrinstein 22 hours ago [-]
we at Control Plane (https://controlplane.com) allow you to run sandboxes anywhere - any cloud (our AWS, GCP, Azure, OCI accounts) or your cloud or bare metal hardware. What sets Control Plane sandboxes apart is:
- They can securely consume ANY service of ANY cloud without needing credentials
- They can securely communicate to any VPC or private network resource
- When you're ready to go to prod - you simply deploy to the Global Virtual Cloud (GVC) which can run in one region, multi-region, hybrid, any number of regions and clouds and data centers.
So Docker is finally adding native support for a microVM backend? I wonder how it well it will compare to using the Kata Containers 4 runtime with the KVM/QEMU backend.
mikodin 1 days ago [-]
Does anyone have a solution for iOS development?
I was all in on sandboxes and safehouse for my agents but the moment I got into iOS development it felt like my hand was forced to just run Claude / codex / pi directly on my machine because nothing else could do the dev loop.
It’s been a painful reality for me, I’m going against core pieces of how I feel I should be interacting with agent harnesses and yet, I need to get the work done so
followercode 1 days ago [-]
Hey, I work at Docker and my team works on mcp integration with sbx. A solution I've been trying is this:
1) Enable the xcode mcp server: https://developer.apple.com/documentation/xcode/giving-exter...
2) Add the xcode mcp server to sbx: `sbx mcp add xcode --command xcrun --args mcpbridge`
3) When you create the sandbox, use `--static-mcp xcode`. For example: `sbx create --static-mcp xcode claude .`
Make sure you have at least v0.38.0 of sbx. This makes a bridge from inside the sandbox to the xcode tools on your host, so be aware that it can run whatever tools you give it on the host. But the agent itself is still sandboxed.
notsirius 1 days ago [-]
Also had this pain point as an sbx user. Given the risk this adds to the host, would be great if there were more docs on how to setup kits to make it safer (e.g. disable yolo mode).
So essentially you can get latest Pi/Node pulling from that image:
`sbx run -t ghcr.io/shaftoe/sbx-template-pi:latest shell`
Like others here I'm also saddened by the login requirement but at the moment this is the best UX I could find for running sandboxed agents, the "kit/mixin" concepts are neat and I make use of them too: https://github.com/shaftoe/sbx-template-pi#stacking-the-extr...
zmmmmm 2 days ago [-]
it's running a full VM so the agent can eg: run docker commands safely etc
1 days ago [-]
outof 2 days ago [-]
Like many people, I suspect, I used Claude to write my own agent sandbox that suits my needs very well. Investing my time in a propietary product has become a hard sell.
matheusmoreira 2 days ago [-]
I did the same thing. It was my first "vibecoded" project. I've been using it every day and it's great. I'm writing a custom Rust network stack for it right now. Gonna replace the current nftables firewall with it.
As for Docker Sandboxes, I'll just ask Sol literally right now to see what it does better than my virtdev, and then I'll improve virtdev instead of using Docker.
zingar 2 days ago [-]
Were you following any patterns/standards/advice on what you needed to protect against? Anything you can point the rest of us to?
matheusmoreira 2 days ago [-]
> Were you following any patterns/standards/advice on what you needed to protect against?
Just the general knowledge that sharing a kernel with untrusted software is too dangerous, that hardware virtualization is an infinitely smaller attack surface and that the entire industry will be in deep shit if people or AI breaks hypervisors.
Initial threat model was supply chain attacks but eventually grew to include AI harnesses as well. Not very worried about them hacking me, more about accident prevention.
So that means each VM must be running a completely independent kernel that's fully isolated from the host's file system. They must also have fail closed network filtering built in.
In summary, it's a QEMU VM orchestrator with a base OS image and project specific delta images. VM lifecycle is managed by systemd. System level isolation is already pretty good and it already solves the "AI wiped out my $HOME" problem. I'm currently working on a custom network stack to replace the nftables based firewall.
embedding-shape 2 days ago [-]
You want to prevent the agent/others from reaching your home directory and other things. As long as you don't mount/sync directories/files from/to the container, so no mounting like "-v $(pwd):/app", but instead copy in, then when done, copy out.
And of course, instead of doing the "copy in > copy out" process manually, get your local agent to write a bash script that does that for you, given what directory you're in, and you're basically G2G.
tjoff 2 days ago [-]
What is the advantage of copying rather than a bind-mount?
furst-blumier 1 days ago [-]
"Oops I deleted everything under $FOLDER – that mistake is on me" doesn't kill it on your host system
tjoff 1 days ago [-]
Sure, but all projects are version controlled? You only mount the project dir so you can only loose your current changes - which is the same if you copy...
hvb2 2 days ago [-]
What specifically are you looking for? If you start from the premise that it runs as you right now, then that's something you can easily improve upon.
Start by mounting just your repo and passing in the keys for the agent. Take it from there, it's like software engineering, you iterate.
When you run into issues you expand the tools in the container available to it.
rvz 2 days ago [-]
Why developers will never pay for their tools.
KolibriFly 2 days ago [-]
[dead]
dSebastien 2 days ago [-]
The one thing I wonder about is how you enforce the usage of Docker Sandboxes vs running the agent on the host directly, apart from scanning machines for binaries
GZGavinZhao 2 days ago [-]
I'm confused:
1. If I run this on Mac, then inside the sandbox / microVM, am I still running MacOS or some Linux distribution?
2. If the only thing that's mounted from the host is the $PWD, how does it guarantee that it has all the system libraries that I have installed on my host system? e.g. my `/opt/homebrew` libraries or `sudo apt install libfoo-dev` headers
akdev1l 1 days ago [-]
Docker uses VMs in non-Linux OS to provide a Linux where containers can actually exist
franze 1 days ago [-]
Here is my solution which uses the Apple Virtualization Framework
als has lots of agents + and typical dev packages (node tooling, python tooling, ....) preinstalled
aliasxneo 1 days ago [-]
There must literally be hundreds of, "Looks cool, but I built <X>" in this thread. It leaves me with mixed emotions. If you're prone to analysis paralysis - this is an unfortunate time to be alive.
KolmogorovComp 23 hours ago [-]
How do you solve the issues of private key sharings that are stored in cwd .env? I haven’t found a satisfactory way to preserve them while letting the agent have access.
pamcake 23 hours ago [-]
> How do you solve the issues of private key sharings that are stored in cwd .env?
Stop putting sensitive stuff there.
panarky 23 hours ago [-]
[dead]
1 days ago [-]
fergie 2 days ago [-]
I use it (sbx), but I don't 100% trust that it actually works, and I would prefer something open source where the limits of the sandboxing could be tested and explored.
Maybe we should just ssh into separate development machines to ensure real and verifiable sandboxing? (as was totally standard before Docker became a thing)
LeBit 1 days ago [-]
You should research bubblewrap and nono.
cv_h 2 days ago [-]
I wrote a CLI tool that uses QEMU's microvm machine type under the hood. It can take any docker image and build a microvm.
I use it regularly to run Claude/Codex with permission checks disabled.
I hardly see how this matters, when Apple and Microsoft already have their own in box solutions for the same problem.
Better sandboxing for AI agents is exactly the main reason for containers improvements on macOS and Windows, with a few talks at WWDC, and BUILD.
Not sure how much they would get from Linux users then.
pixard 2 days ago [-]
Ah let's see, do they still want you to LOGIN, in order to use a local dev tool? Yes, yes they do. No thanks Docker. You can keep your buzzword reasoning as to why this is needed.
Doesn't everyone do this now? It's hardly a new idea. Yet every time someone proposes the idea, people fawn over it and proclaim it the best thing ever.
Yes, you can inject tokens via a proxy. What else is new?
nopurpose 2 days ago [-]
Who is doing it as first class feature with at least adequate UX?
I have skimmed alternatives offered in comments to this post (vibepod-cli, code-on-incus, opencode-docker, sandboxy, smolvm, amazing-sandbox) and none of them seem to do credentials injection at the proxy level.
LeBit 1 days ago [-]
nono.
Also fnox now does credentials proxying.
nopurpose 1 days ago [-]
Thanks, bookmarked nono to have a look later.
pkulak 1 days ago [-]
I’m sure they fixed this, but since everyone runs docker containers as root… is every file this thing writes going to be root owned? Does it have root access to any resource to give it visibility to?
killerstorm 1 days ago [-]
Are we sandboxing AI agent harness process, or the environment it executes commands in?
Ideally, they should run in _different_ sandboxes.
The environment might corrode the harness (e.g. rogue npm/pip packet would manipulate agent harness config).
I planed to do exactly this, with podman instead of docker, volume support.
Like:
$ podman run -it --rm -v .:/workspace local-dev-ia /usr/bin/oc
Configured with a .env file.
Hope to do it hopefully before the end of the week.
codethief 1 days ago [-]
> exactly this
This is nowhere near "exactly this". Docker Sandboxes uses micro VMs, you just use regular containers which have completely different security properties.
akdev1l 1 days ago [-]
podman run --annotation=run.oci.handler=krun -dp 8080:8080 -t --rm server-without-wasm
codethief 14 hours ago [-]
Yes, you can run Podman with different OCI runtimes, in the same way as you can run Docker with different OCI runtimes, and some of these OCI runtimes are microVM-based.
This is not what the person I was responding to is doing, though.
As for differences between the krun OCI runtime and Docker Sandbox (which also uses libkrun), let's please continue the discussion here: https://news.ycombinator.com/item?id=49240662 .
Draiken 1 days ago [-]
Bubblewrap plus some whitelisting of domains/sockets is all you need.
Docker is always a pain to use and this way I don't have to re-install everything a billion times for every different project.
llimllib 1 days ago [-]
This is what I currently do, but my software uses docker and docker mounts act as a bypass for the file system restrictions, plus docker processes started outside the sandbox allow network proxy escape.
Currently, I don't allow the agent access to docker, start docker myself, and then do short-lived sandbox-free sessions when the agent needs to do things that interact directly with docker; but that's annoying.
3371 2 days ago [-]
I used this for a while then decided to build my own suites that pack individual harness and respective host state (config, plugins, skills, etc.) into an image. Works better and much flexible in my opinion.
I wish they solved the issue happening for years on MacOS where Docker keeps up eating all available free space and ends up requiring restart of the whole machine, instead of Gordon and other useless shit.
TekMol 2 days ago [-]
So this is a VM by Docker?
For those who do not trust
docker run --rm -it -v "$(pwd)":/work -w /work myaiimage /bin/bash
AND do not want to use some other, free VM for some reason?
woadwarrior01 2 days ago [-]
Better yet, use Apple's container CLI if you're on a Mac, instead of the docker bloatware.
container run --rm -it -v "$(pwd)":/work -w /work myaiimage /bin/bash
pulse7 2 days ago [-]
Hasn't Docker always been just a thin layer of duct tape over existing solutions?
khanhnguyen8386 1 days ago [-]
Finally, a way to run --dangerously-skip-permissions without having a mild heart attack every time the agent decides to rm -rf a mystery directory.
TZubiri 1 days ago [-]
'adduser agent'
'su agent'
'curl domain/install.sh | sh'
'runagent'
Cameri 1 days ago [-]
I was going to try it but signing commits with GPG using a Yubikey is not supported.
The option left is to use SSH to sign commits which is a no-go for a different reason.
2 days ago [-]
notsirius 2 days ago [-]
been using this for a while - works great! Has also had a lot of updates over the past year so worth checking out again if you tried it a while ago
zingar 2 days ago [-]
Do the agents come preinstalled in the images? Or do they somehow use whatever I’ve installed locally? The former makes sense to me but then I’m wondering whether the sandbox images stay up to date with new releases of each image.
Has anyone started proving their sandboxes in Lean (or Coq, etc.)?
root-parent 1 days ago [-]
A Docker container is not a strong security boundary.
venatiodecorus 1 days ago [-]
these are firecracker microvms iirc, not containers
runtime_lens 2 days ago [-]
TO me, that's the important distinction: sandboxing limits what the agent can do but it doesn't necessarily enforce that the agent must run inside the sandbox. You need a separate control layer to enforce that boundary.
nezhar 2 days ago [-]
You design the sandbox so the agent starts in that layer. The next thing you can do is to limit the network access, this is what I'm working on right now.
Or do you mean something else?
runtime_lens 14 hours ago [-]
[dead]
AmazingTurtle 2 days ago [-]
So it's basically a container with a fancy name, innit?
Not sure what you mean by "just". Containerisation is generally understood to mean something like what Docker does, which includes sandboxing but a whole lot more on top, like image management etc. Bubblewrap is just sandboxing without the rest of containerisation.
mihaelm 2 days ago [-]
Docker Sandboxes is using microVMs, not containerization.
globular-toast 11 hours ago [-]
That's an implementation detail. They do that because less capable OSes don't have direct support for sandboxes.
angry_octet 11 hours ago [-]
No, that's not why.
blueaquilae 2 days ago [-]
Docker management will fail their tech at every opportunity.
I don't want this that bad. I want the agent to have open access to my system because it actually does important administrative things for me. It is THAT convenient and powerful.
Here's what I want: REALTIME OBSERVABILITY/POWERPOINT.
I don't want to just see what command it ran. I need graphics... what part of the file system it is touching, what network entities it is contacting. If it's running SQL I want the parsed query handed to me in a syntax highlighted and well formatted interface. Imagine that star trek computer presenting automated infographics while someone is doing a presentation, you know what I'm talking about? It's like a automated powerpoint as the agent does it's thing.
I need to understand my agent and what it typically does so I can dangerously wield it. I treat the agent like a gun in a live shooting scenario. That's how I want to use the LLM.
Sandboxes have their purpose. Just like how shooting ranges have their purposes. But I need to fire my gun in the real world and real world is a warzone.
quantumwoke 1 days ago [-]
Just a small meta note: most of the comments in this thread appear to be posting their own codebase (typically AI-generated) that accomplishes the same goal. It's interesting that this problem is simultaneously in high demand and yet considered trivial enough to vibe code per-user solutions to it.
cryptoz 2 days ago [-]
The linked page implies there is no linux support, I wonder why. It's there in the docs if you hunt for it.
...or...just hear me out now...we could limit it in the harness.
Don't give it shell access, just predefined tools.
mikesir87 1 days ago [-]
Disclaimer - on the Docker DevRel team
One of the demos I run is how easy it is to circumvent the harness limits. For example, I can configure a harness not to access file `secrets.txt`. But, then I can immediately have it create a Python file that can read any file and have it read `secrets.txt`.
At the end of the day, "please" isn't security. You want to know that the agent can only do and access the things it should access.
killerstorm 2 days ago [-]
What if it puts malicious code into test file and you allow `npm run test `?
guluarte 1 days ago [-]
the problem with this is... now i trust the agents more than myself lol
tenner_agent 1 days ago [-]
[flagged]
GautamTalksDev 1 days ago [-]
[flagged]
BorisBinyaminov 14 hours ago [-]
[flagged]
krupkinmaxim 1 days ago [-]
[flagged]
lubo92 1 days ago [-]
[flagged]
aegisora_ai 1 days ago [-]
[dead]
genshro 1 days ago [-]
[flagged]
2 days ago [-]
wpdevant 1 days ago [-]
[flagged]
aegisora_ai 1 days ago [-]
[flagged]
1 days ago [-]
claud_ia 1 days ago [-]
[flagged]
beernet 2 days ago [-]
[flagged]
saadyousfi 1 days ago [-]
[flagged]
KolibriFly 2 days ago [-]
[dead]
Esabelle 2 days ago [-]
[dead]
lorreyfum 1 days ago [-]
[dead]
songhonglei1985 2 days ago [-]
[dead]
liquid_space 1 days ago [-]
[dead]
pullrun 1 days ago [-]
[dead]
zuzululu 1 days ago [-]
[dead]
kmeh 2 days ago [-]
[dead]
solarengineer 2 days ago [-]
There is a name collision on MacOS where MacOS also provides containers [1]
Customize the exact environment of your container from the ground up (harnesses, tools, base image, packages, mounts, etc) and enter with a single command.
One correction: this isn't containers. Each session is a microVM with its own kernel on the platform's native hypervisor: Hypervisor.framework, WHP, KVM. We wrote a new VMM (not Firecracker) to make it more effective across platforms.
Explained a bit more here about the architecture and why those choices were made: https://www.docker.com/blog/why-microvms-the-architecture-be...
I stopped using Docker on macOS because host file system performance was so slow, even with all of the caching hacks piled on top of it, that it made the whole thing effectively unusable for development.
Directionally the post shared sounds great, but it seems "too good to be true" that we'd have a performant microVM for macOS.
Recently I've also been working on a VM stack for an agentic platform using pre-build images with some cloud injection scripts that simplifies the deployment of a private agentic cluster - in the end I went with full VM with a 4vCpu/8gb for the main agent and 2vCPU/4Gb - only the main agent had docker-in-docker, the rest rootless docker but I agree it's still an elevated risk.
I'll definitely have to give this a spin and see if I can simplify it to one larger box with this solution.
Concept is great - works quite well - I often have multiple short lived sandboxes running at once.
Docs [1] on overriding auth are incorrect. Sbx ignores inject[].username for basic auth and instead the stored secret needs to be the complete Authorization header. This should be made clear, or fixed.
Having to log in every couple of days SUCKS!! Opening the browser so I can login (which we shouldn't have to do) interfers with my scripts that create and destroy sandboxes as I need them.
I miss the old worktree functionality - I dislike the new clone concept - So I've created my own scripts that create a worktree for a feature, and run sbx create/run from there.
[1] https://docs.docker.com/ai/sandboxes/customize/kit-reference...
Read more about the secrets handling here - https://docs.docker.com/ai/sandboxes/security/credentials/
Kits provide the ability to also define new credentials and how to inject them into new services (connect to internal systems, etc.).
Is it based on libkrun?
For the people upthread who asked about on customization: templates (like snapshotting a running sandbox) and kits (YAML applied at creation like install steps, files, network and credential rules, or define a new agent outright) are the supported path now. It's early but take a look here: https://docs.docker.com/ai/sandboxes/customize/
On MCP, since credential handling was mentioned here: the sandbox sees one gateway endpoint, and OAuth tokens stay in the host credential store rather than in the VM. https://docs.docker.com/ai/sandboxes/mcp-gateway/
All this is early. We're looking at more based on feedback from users like running sandboxes in the background for long-horizon work and a lot more (including what you all raised in the thread here). Keep them coming.
https://www.docker.com/products/docker-sandboxes/
Brew is notoriously developer-unfriendly.
Now, whether you should you be using the Homebrew Python is a completely different question. YMMV for other platforms managed via Homebrew.
I've traditionally used MacPorts for dev tooling and Homebrew for everything else, but with more aggressive adoption of tooling like uv an nvm I'm not sure the different really matters for me anymore.
In the end I realized that Brew is a package manager for consumers, and as a professional i should’nt keep fighting it.
Python through brew is the one I expect to be the latest one I use for one-off scripts.
Python in particular is well known to not be a stable target. For anyone. By design. If you expect long term use of a specific version of code, use a different language. It is not at all homebrew's fault that they're how many people discover that.
Is this open source? Can I install this on a non Ubuntu system?
See repo `docker/sbx-releases`. The `.rpm` there has Rocky Linux in the name but works on Fedora.
I looked at the docs last week and didn't see anything about that.
https://learn.microsoft.com/en-us/virtualization/api/hypervi...
https://www.qemu.org/docs/master/system/whpx.html
On Windows 11, too? At least for hardware virtualization in VMWare one would have to disable Windows Device Guard & Credential Guard for that.
I run it with superset and then each git worktree is mounted in a sandbox that is configured for each repo i work in.
Closest open source I have seen is https://earendil-works.github.io/gondolin but the DX is not as polished. https://exe.dev/ would be perfect but it does not come with outbound firewall.
Does anyone have a better alternative?
https://github.com/kstenerud/yoloai
Outbound firewall is `--network-isolated`: egress is denied except the agent's own API endpoints plus domains you allow, enforced sandbox-side (working on host-side enforcement now). `--network-none` if you want nothing.
Credential brokering works the way you describe (currently Claude-only, I'll add more as time allows). The API key stays on the host, a local proxy injects it into the outbound request, and the sandbox never holds anything worth stealing. Other agents' credentials currently arrive as read-only file mounts instead (weaker, and something I'll fix soon). Generalising the injector is the obvious next thing.
One difference from your setup: yoloAI copies your worktree instead of mounting it. The agent works on the copy, you `yoloai diff`, and `yoloai apply` replays the commits into your real repo. That's deliberate. Docker's own security docs talk about the dangers of bombs being left behind in a live-mounted dir (git hooks, package.json scripts, Makefiles, IDE task config), which diff/apply avoids.
Isolation is per-sandbox rather than fixed: runc, gVisor, or Kata VMs (QEMU or Firecracker) on Linux; Seatbelt or full macOS VMs via Tart on a Mac.
Cloudflare/artifact-fs does lazy shallow git clones with a FUSE filesystem. https://github.com/cloudflare/artifact-fs
Would that be faster?
Re: sandboxing methods like Clawk, Amla sandbox, bwrap, agentvm, ARM64 MTE with wasmtime-mte: https://news.ycombinator.com/item?id=48893850
A few months ago now I started adding seccomp sandboxing to jinja2rs and then liboverlayfs support to ansiblers (which are early Rust ports).
Haven't finished that, but I started working on a VM format that stores signed machine state into an OCI container repository, using the hypervisor migration support of KVM/QEMU.
Though this is not safe yet if ever, VM migrations are probably another way to sandbox and deploy en masse.
It has network filtering + placeholders for secrets.
OSS, no logins needed
Documention at this moment in an early stage.
Overall, its a great project but for me was simpler just use Virtual Machine Manager (libvirt GUI).
I wish all luck to the maintainers, but probably DX-wise I will prefer to have more granular or predictable controls (eg micro cloud from Canonical).
(Not affiliated with them, just tried it out last week.)
Yep - similar in some ways but headed towards different directions.
I am building a virtual machine to simplify/replace container infra. Ex. we run containers inside of linux VM's even in the `cloud`, resulting in managing both the vm, and the containers.
But smol machines is a lightweight, portable VM that you can package into a single portable .smolmachine file to be rehydrated on any platform, kind of like how containers are used for today.
Sandboxing happens to be a feature of virtual machines, so we are alike in being used for sandboxing.
I'm using it as my main driver since months.
Support for running agent harnesses in unprivileged podman containers is on my feature list. :-)
Only real alternative is to use microvms. My goto solution for this are apple/containers.
For the vast majority of developers, containers are enough, which is why they are ubiquitous while vms are less common. Ofc that ubiquity has led to lazy configuration, which is how the jailbreaking can occur. Knowing what you are doing with containers is a requirement to use containers as an AI sandbox.
I'm not sure where "vast majority" cuts in but I would say a huge number of developers use docker and it is inconvenient at best if your AI harness can't actually run and test the infra it is building against
Why microVMs? I never ever run a container, AI harness or other, in something else than a full on VM. I could use a microVM but in any case I really don't see why I'd run a container on one of my bare metal OS: the place of a container is inside a VM (or microVM).
Especially for AI harnesses where the threat of an escape is very real: the more defense in depth, the better.
And If I can use rootless Podman instead of "rootfull" Docker, the better. Most of my containers are Podman btw.
> My goto solution for this are apple/containers.
To each his own: my goto solution is an actual server on my LAN with shitload of cores and memory and plenty of scripts to provision VMs etc.
I really don't understand why people are YOLO'ing containers on their bare metal OS.
You know about the people who do not bother with the container? Quite a few make “No problem so far” comments on HN discussions.
https://github.com/NVIDIA/openshell
I have a dedicated container for that. It can run its own Docker daemon and other system services if needed.
Apart from the Claude login token, it has no SSH keys or other credentials. I push everything I need to it from the local machine. And I pull the Claude generated outputs from it.
Of course, this kind of setup requires a stack which can run or at least be tested without any credentials.
What I don‘t have compared to sbx is an outbound firewall, but my VM does not have any personal/interesting data, only a vanilla Fedora installation and the project dir with open source code, so I do not care much about exfiltration.
https://GitHub.com/jgbrwn/vibebin
It also provides a nice TUI for network policy management and prebuilt sandboxes which have claude code, codex, etc.
Not necessarily better but OpenSandbox[0] by Alibaba seems similar.
[0]: https://github.com/alibaba/OpenSandbox
Has egress and ingress filtering, egress can be bound to host/internet/subnet or even better to internal apps (which are each separate netns) meaning you can do your own firewall/vpn/whatever per sandbox. Plus you control what other components in the sandbox env the app can communicate with.
Really not built for day-to-day dev work though, more like automating your company/life / getting rid of SaaS (e.g. for technical Founders / Sales etc, not exactly useful for dev work)
sbx run --name "yoursandbox" -- --cd "$PWD"
This boots a sbx session in the worktree directory.
For Claude there is no --cd so it's more hacky, but I solved it by creating a sbx kit with entrypoint script that reads a flag (e.g --cwd) from the terminal preset command and then inside the sandbox cd's there and starts claude.
Def makes integrating easier but I try to avoid using direct mode for security (exposes .git folder, though there's probs a better way to protect it by disabling hooks or something)
https://docs.docker.com/ai/sandboxes/workflows/#host-worktre...
It's an OS-level sandbox, though. It doesn't launch VMs or containers for sandboxing purposes; it uses whatever sandboxing features your host kernel offers.
We run cloud sandboxes, and have some experimental local sandbox support that is fully OSS.
Main thing for amika.dev is you can control the sandboxes and agents interchangeabley by SSH, web, or API, and can expose the services the agent is working on over signed URLs
Entire sandbox config is a TOML file
We're going to improve the local OSS sandbox mode and add better network controls over the next couple weeks.
Ultimately, what we're building kind of like if Tailscale and Firecracker had a baby, with a messaging protocol for remote controlling any sandboxed agent
It's free to try out. Still a lot to build, so we really appreciate any and all feedback about what we should focus on
Mine uses containers, and makes only the git/jj workspace read/write, hiding all credentials that are in my home dir.
Internally uses a single VM + Incus containers, supports docker/Kubernetes in each sandbox, has integrated worktree management.
thanks
So I wrote a wrapper around Gondolin which allows me to do that and a few other things: https://github.com/codethief/tuor
(Warning: Still very much experimental / underdocumented.)
The definition for your cloud sandboxes is just a TOML config in your repo
We also have an API and CLI to let users message the agent from outside or across sandboxes
We're still building a lot, so if you have any time to try it out (amika.dev) and give feedback, that is worth gold to us!
It seems with gondoling i need to explain the agent to run commands in the sandbox, but then where does the agent run itself?
[0]: https://earendil-works.github.io/gondolin/workloads/
Their example implementation with pi uses a pi extension so that pi runs on the host but the read/write/bash/etc tools run in the guest. Doesn’t have to be that way though.
And that's just an easy case - really if it has any network access at all it can come up with a clever way to route a request through the network such that the key comes back somewhere in the request. If you scan for it inbound too, the machine can obfuscate it.
Our agents are trained to be so intensely helpful and they have such intricate knowledge of how things work that they will do some incredibly clever tricks to do what you ask them to do.
edit: You may want to look into tokenizing proxies as the general application of this concept.
The main reason a "proxy-managed" env var is set is because most CLI tools assume if the env var is set, auth is set. If the env var is unset, it will assume auth needs to occur. Fortunately, most don't do a pattern matching on what the value actually is.
There's also an ability to create kits where you can setup credential injection into other services as well.
https://github.com/pjlsergeant/byre -- slightly different security model, but lazer-focused on developer experience; my daily driver and I love it not just because I wrote it. The TUI is great for configuring and setting up instant boxes just how you want
https://pleasedonotescape.com/ -- a list of every other agent jail I could find, filterable by open-source and whatever else you want
TLDR: You're agent will get a isolated v8 runtime (chrome's sandboxed javascript runtime)
Basically, it's a plain QEMU/KVM VM on a stock Debian cloud image: device model stripped down to a virtio disk, a virtio NIC and a serial console, nested virt off, no passwordless sudo in the guest. It also ships a containment check that scans outward from inside the guest, so the network boundary is something you can verify.
Wrapping the whole environment rather than a single agent session puts supply chain attacks inside the boundary too. A poisoned npm or PyPI package, or a compromised editor extension, lands in the VM instead of on the host. That was the original reason I set this up; agents just made it more urgent.
There's no per-domain egress allowlist; the policy is "internet yes, private addresses no". Secret injection isn't built in either, though Infisical's agent-vault on the host as an egress proxy covers that part.
Wrote the whole setup up here, in case it's useful:
https://karamatli.com/posts/network-isolated-kvm-sandbox-ai-...
> So rather than pick one, this post advocates layering both, in the spirit of defense in depth: a sandbox VM wraps your containers along with the whole toolchain, and that sandbox reaches the internet but has no route to anything private.
Yup it's the only proper way.
And that is true not just for AI harnesses/agents (that shall try to escape), but also for stuff like Plex/Jellyfin/Immich/private pastebin etc.
If you care about security, there really simply is zero reason to run containers on the bare metal.
What's a "microVM" and what's the security model here compared to using real virtual machines with actual constraints on breakouts?
Is it marketing fluff?
Incus/LXD has had VM's for a long time now.
Firecracker is designed to start a VM in under 125ms and 5MB. Netbsd advertises that you can direct-boot a MICROVM kernel configuration in under 10ms.
I think this is one of the big issues with today's MCP servers. Most are based on Node.js and take a lot more memory than they should, compared to the complexity that the job requires. Just run a few MCP servers locally, and all your RAM is gone...
https://en.wikipedia.org/wiki/Memory_ballooning
There is a need for a Docker desktop like GUI for this market.
The world didn't fit around me, so I made "me" fit around "the world". I bet those stubborn dinosaurs can learn a new trick or two also, if management lets them do it during working ours...
Also, WSL (Windows Subsystem for Linux) has been baked into Windows for a long time and makes it very easy to play with Linux, as does using the Hyper-V VM system. Any developer unfamiliar with Linux because they use Windows, has little excuse.
Modern .NET did not went cross platform by accident, and Java development has always been "develop on Windows deploy on UNIX", in corporations where Mac tends to have little presence.
That's my experience anyway, there's a lot of restrictions once you need to do some real basic things. For basic prompts maybe but interacting with a full stack ehh.
So bit hesitant to call firecracker a real VM myself.
Packer and Incus setup:
Then: The reality of these things are that eventually you will want to do something useful or different with them and the flexibility simply isn't there compared to a real vm, if you desperately need boot times then there's plenty of other options here, especially with packer. Some of my prompts are often hitting 60+ minutes so it's not really something I think about.I use something in between - a single Ubuntu VM, into which I spawn multiple Incus LXC containers for the agents. The containers only use 50 MB or so per instance (separate systemd, ...). This way I pay the VM RAM tax only once, and the agents are still contained inside the VM if they manage to escape the LXC containers.
This way I don't pay tens of VMs "tax" but basically only three (plus one or two VMs I use for testing enhancements to my VMs provisioning / optimization / securing setup, when I work on that).
I don't use LXC (I could) but regular containers, inside VMs.
I'll look into the stripped down "micro" VMs but then I don't spend my days launching VMs/shutting them down so it's not a big deal.
I find solutions like this to be a like trying to patch a leaking boat on a lake with duct tape. It will help, but it's not a proper solution.
Also, often the tasks you want the AI to perform are in the outside world. Like "connect to my servers, and figure out X and Y".
The proper way is permission isolation. I run a small k8 cluster in the homelab and I have 3 types of pod/agent combinations for my AI agents. Read only, one that can change my gitops but it needs to create PRs that admin approves, and admin.
Likewise with code. I have a forgejo git instance where agents have ability to create feature branches and so on, but merging is gated.
Those things require "GH enterprise features".
In fact more and more things we do at home will require "enterprise features". Why? Because a person with AI is basically a small team, but some of team members behave like Chimps on crack... So security must be top notch.
I don't think "proper" permissions are possible without dramatically limiting the way people use these tools.
https://docs.docker.com/ai/sandboxes/security/credentials/
> That runs the codex OCI in a qemu microvm.
AFAIU it's actually the other way around: krun spawns a libkrun-based (not QEMU-based) VM inside a crun container. Source: https://github.com/libkrun/libkrun/discussions/538#discussio...
So with your solution you get the additional security benefit of containerizing the hypervisor on the host.
(it's in the doc, but kinda strange to not see some instructions on the main page, probably distro related)
https://github.com/swelljoe/flar
I’ve been using this pretty extensively for a few months on Mac and Linux and have been super happy with it.
The nails in the coffin were 1) login was required 2) login was broken because they "didn't consider" it would be run in a headless environment [0] and 3) they shipped with hardcoded binary paths and root requirements [1].
I moved on and use Incus directly with small helper scripts, smolvm, or a full fat VM running desktop Claude or ChatGPT if I (or someone I mentor) really needs the full app. I'm not yoloing every new claw agent in --dangerously-destroy-my-things mode, so network restrictions are best effort, though filesystem access stays tight.
Either way, docker sandboxes really didn't seem to be it, and the company didn't seem interested in trying to be anything beyond an enterprise solution.
0. https://github.com/docker/sbx-releases/issues/186#issuecomme... 1. https://github.com/docker/sbx-releases/issues/48
I took a similar approach with https://runmachine.dev/ but later switched to OrbStack for iOS development.
It blocks network access by default, mounts only what you specify, and you can add a customization layer. This is all done in the container itself (srt for network blocking). It doesn't implement a central point for secret sharing, MCP exposure, etc. So it might not have enough features for some but it works well for my needs.
I just found through this thread yoloai which has an apple container backend, so its quite similar using that. My main issue would be that network access is allowed by default. https://github.com/kstenerud/yoloai
Several other projects listed here use libkrun which is an alternate implementation that works with Mac's HVF. smolvm, microsandbox, podman (with likrun backend), gondolin.
https://github.com/iodize6399/ai-devcontainer/tree/main/.dev...
I quite like the 'features' layer system, adding extra tools to container in a declarative plugin-like way
Being able to 'safely' run with skip permissions has been a gamechanger
I especially like the firewall it has.
And even if that works, this is a very short list. As soon as you reach for Go, Rust tooling etc nothing works. So you need to manually maintain this list which is nothing but painful trial and error.
https://www.mrafayaleem.com/blog/sandboxing-claude-cli-with-...
Haven't tested it yet, but it seems to address the same issue as Docker Sandboxes, but in a different way.
But it’s not as well sandboxed for sure.
(I'm sure you can spend time to come up with a proper bubblewrap configuration that allows go build to succeed, but it's probably not worth the effort.)
Eg, if used with Colima in macOS, it means I can run a devcontainer in an isolated VM and Nono inside the devcontainer can restrict a lot what can and cannot be done.
You get credentials proxying and network outbound limits.
How is Docker Sandbox better sandboxed?
Both went through stuff the sandbox was set up to allow.
eg: https://josecastillolema.github.io/podman-wasm-libkrun/#libk...
Obviously, there's a reason why Docker released Docker Sandbox as a separate product:
- Barely anyone bothers to configure Docker/Podman with a different OCI runtime like krun. Heck, most people don't even know about OCI runtimes in the first place. Case in point: Most people here in this HN discussion are proposing using "standard" containers (with the default OCI runtime) for sandboxing. This is what I was trying to get at.
- A sandbox for agent needs tighter network control.
As for differences between the krun OCI runtime and Docker Sandbox (which also uses libkrun), let's please continue the discussion here: https://news.ycombinator.com/item?id=49240662 .
(I am testing this now as a backend for my pet project which currently supports firecracker and gvisor. No network.)
But, in any case, why put in effort doing something people don't expect or ask for? We can assume everyone running agents is either a) using their own sandbox, or b) doesn't care. I think we can guess which category most people fall into. You could maybe argue about responsibility, but I don't think you can argue about "serious engineering".
Exactly. It's not as though it's difficult. It never occurred to me to not do this from day one, and it astonishes me that anyone runs this stuff bare metal. Since then, I've brought several other people on board, and that's all they've ever seen: I don't think they'd know how to run outside a sandbox, and that's just fine.
Also if your thing doesn't work with `pi` out of the box, then low effort
{ "allowedHosts": [ ".anthropic.com", ".claude.com", ".pi.dev", "npm.org", ".npmjs.org", ".github.com", ".githubusercontent.com", ".pypi.org", ".pythonhosted.org" ], "baseImage": "docker.io\/library\/node:22", "displayName": "Pi", "environmentVariables": [ "IS_SANDBOX=1" ], "installCommands": [ "npm install -g --ignore-scripts @earendil-works/pi-coding-agent", "npm install -g global-agent" ], "launchCommand": [ "pi" ], "mounts": [ { "containerPath": "\/root\/.pi", "hostPath": "~\/.pi", "readOnly": false } ] }⏎
Can a user blindly trust this sandbox, because that's how people will treat it based on the marketing. Sounds like it could be useful for far more than just AI though.
Docker Sandbox spawns a micro VM, not a standard container isolated by host kernel mechanisms (Linux namespaces etc.)
The VM sits on its own libvirt network in a dedicated firewall zone, and specific directories are shared via filesystem passthrough. To keep the agent from accessing anything related to Git, the actual gitdir is stored on a separate path outside the mount point.
I review the git diff manually and commit it from the host.
Then again, I guess GUI is a pain in Docker. I tend to operate through Zed and an ACP harness though, so my GUIs are sort of “inside the container” anyways.
Currently considering using ACP for codex so that I can do more of the driving from my editor (emacs, over ssh) and something similar for claude code (it doesn't seem to be as good as codex at supporting re-attachable sessions).
One concern is making sure my editor's ACP client doesn't enable/support fancy terminal stuff, because that would basically void all the benefits of using a VM sandbox.
Would suggest Hashicorp Packer or cloud-init for deterministic images, not hard to setup or use. LLM's have little problem with them either I find.
If you need quicker environment rebuilds consider using something smaller like alpine as the base, though once you setup a golden image even heavy things like debian are fine.
https://github.com/pkhamre/opencode-docker
All the open GH issues about supporting Linux that I subscribed to have gone unresponded to.
OpenShell looks like a good alternative, but it still has "Do not use in production" plastered all over the website, which doesn't fill me with confidence yet
https://docs.docker.com/ai/sandboxes/#get-started has instructions for Ubuntu.
I wrote a tool to use `bubblewrap` to containerize any agent (at least all the agents I've used a couple of times), and bind mount the system stuff read-only, so the agent has your "usual" environment, but they can only see the project. Their history persists (either through a bind mount or a "shadow" copy of the history that only the wrapped agent sees), the agent can still create and manage containers of its own using podman's rootless mode, etc. It's nearly instant to start because it's just a namespace (plus a few copied files for the container support and session history); no container needs to be built/fetched/updated/whatever. bubblewrap is extremely well-tested as it is used by flatpak and several other large projects, so I trust it quite a bit (more than I trust Docker).
https://github.com/swelljoe/flar
But, Docker is rarely the right way to manage containers on Linux, IMHO.
Regardless, I'm hoping something that isn't behind a login screen is going to win out.
There is no reason to require a login for creating local mini sandboxes.
If you’re on Apple, native solutions like “container-machine init” come built in and are pretty good, if you’ll only be on Apple hardware.
Though, they need to remove the login requirement.
Still early and the local sandboxes are experimental right now
https://github.com/gofixpoint/amika
For clarity I’m actually using podman, not Docker.
Now I'm curious to know how hardened the Docket Sandbox orchestration interface is. I guess we can assume they have run Mythos against it for a few weeks maybe? It's unclear.
Unless we're talking 0-day/CVE, running an unprivileged container is as trustable as a VM. The only difference is how strictly you want to hold the memory/CPU bar. Infact on linux, containers are more lightweight than VMs.
So yeah, not "vulnerable".
Containers have access to the kernel ABI, and as shown in the latest kernel exploits, all the memory handling surface that exposes. The virtualisation interface, offering fewer services, is significantly harder.
Containers are obviously lighter than VMs, both to start and to schedule, but firecracker is pretty fast. gVisor pays overhead per syscall vs at startup.
So yeah, more vulnerable.
On Linux, you can run docker directly on the host, but you can also very easily setup a vm with incus and run docker from there.
Example: sbx run claude ./ ../another-project:ro
But sbx is a bit annoying to use with OpenCode for instance (which has zero sandboxing by default, unlike codex CLI or Claude Code). You cannot easily change ~/.config/opencode/opencode.jsonc AFAIK.
[0]: Black Hat OpenAI-Hugging Face incident: https://www.youtube.com/watch?v=87DyyMV0kCY&t=1021s
Still obviously you should run all untrusted code in a sandbox, but extreme actions like that would be very unusual with the model that shipped.
My startup (https://github.com/gofixpoint/amika) copies agent configs into local or cloud sandboxes
We run OpenCode currently, but need to improve the setup for users, so would like understand more of the issues you have sandboxing it, if you can share
cgroups are a mechanism designed for hierarchical organization and resource distribution. Against a malicious and capable actor, and that is how we have to treat AI agents, cgroups will not withstand.
Also, the kernal is an interface too big for what an AI agent needs and is therefore offering a gigantic attack surface completely unnecessarily.
This is not my main point though. Both are based on cgroups and cgroups are the wrong tool for the job.
The same is not applicable for any kind of sandboxing for two reasons:
1. The boundary is in the kernal’s own code, enforced by the thing you are trying to be protected from. -> Use a VM
2. The kernal is a gigantic attack surface -> Use gVisor
[1] https://docs.cloud.google.com/kubernetes-engine/docs/resourc...
Say I want to use pi inside a container. If I wrap pi within a bubblewrap or within nono, how is that less secure than using a vm?
Also, I think most people run containers inside VMs anyway and not directly on their hosts (on Mac and windows you have to use a vm anyway).
It doesn’t change the fact a malicious process is still attacking the same kernel , but it can reduce what it can do to that vm.
Automating browsers with LLM agents properly requires a lot more work than Process.Start into powershell, but the advantages can be immense once you have achieved integration this way. Incrementally maintaining this integration is generally easy because human users cannot tolerate rapid changes either.
It's a hell of a lot easier to convince management to adopt a robot that looks and acts like a human employee than one that looks like a combine harvester. The combine is far more efficient, but it is also totally indiscriminate. Nothing constrains its appetite except for the invisible fence imposed by GPS. The amount of infrastructure required to keep farm equipment from running astray is incredible. In the context of agriculture, the added complexity is definitely worth it. We don't want to have to recreate the same thing with our technology if it can be avoided. Sandboxes and security isolation boundaries are not things to aspire to. These are costs to be paid for admission to something more valuable.
Also, what took them so long?
Anyway, I decided to try it in a VM. Got: "You are not authenticated to Docker. Starting the sign-in flow..." (Just to try it.) Joke's on me.
It goes like this: - bash script parser + interpreter (with hooks for things like file open, execute etc.,) - wasm executor for execution. - wasm implementations of common tools like coreutils, grep, sed etc., from the uutils project. - wasm implementation of python by a VMware backed project. - entirely virtualized filesystem using Go's io/fs.FS. (tmp dirs can be implemented using any backend)
Works like a charm for the limited usecase I have. There are definitely some drawbacks with threading and especially with preopens in wasm. But a cheap sandbox for simple file explorations and minimal computations.
I won't use anything requiring a login.
One other thing, I want to be able to handle multiple repos in the same sandbox and have a standard workflow around worktrees (one worktree per repo, all the worktree mounted in the VM).
These were some of the reasons that led me to build: Clawk - https://github.com/clawkwork/clawk
I've been using sbx for a bit now, and there have been some old versions that had this problem, but haven't had this problem in a while when using secrets https://docs.docker.com/ai/sandboxes/get-started/#authentica...
There are many better namespace/container based options, VMs may be moderately more secure but when you more or less trust your agent and code you can do with lesser containment. And with the recent CVEs in kvm honestly there isn't a huge deal of difference vs namespaces.
(I'm building https://xbin.dev/ for some time now for managing my personal code/apps, a project which started specifically after Docker Sandboxes broke on me some time ago)
Holy moly, on mobile I was trying to read the example console screenshots/snippets and then it would just unexpectedly change. Took me a little while to figure out it’s some kind of carousel for the examples, and not more screenshots/snippets loading and pushing down content (or me going crazy). Please don’t do this on mobile sites, just let me scroll through the examples!
- They can securely consume ANY service of ANY cloud without needing credentials - They can securely communicate to any VPC or private network resource - When you're ready to go to prod - you simply deploy to the Global Virtual Cloud (GVC) which can run in one region, multi-region, hybrid, any number of regions and clouds and data centers.
our website is https://controlplane.com
I was all in on sandboxes and safehouse for my agents but the moment I got into iOS development it felt like my hand was forced to just run Claude / codex / pi directly on my machine because nothing else could do the dev loop.
It’s been a painful reality for me, I’m going against core pieces of how I feel I should be interacting with agent harnesses and yet, I need to get the work done so
1) Enable the xcode mcp server: https://developer.apple.com/documentation/xcode/giving-exter... 2) Add the xcode mcp server to sbx: `sbx mcp add xcode --command xcrun --args mcpbridge` 3) When you create the sandbox, use `--static-mcp xcode`. For example: `sbx create --static-mcp xcode claude .`
Make sure you have at least v0.38.0 of sbx. This makes a bridge from inside the sandbox to the xcode tools on your host, so be aware that it can run whatever tools you give it on the host. But the agent itself is still sandboxed.
On another topic, can't help but notice that "leading coding agents" somehow does not include Pi.
To work around that limitation I came up with this https://github.com/shaftoe/sbx-template-pi
So essentially you can get latest Pi/Node pulling from that image:
`sbx run -t ghcr.io/shaftoe/sbx-template-pi:latest shell`
Like others here I'm also saddened by the login requirement but at the moment this is the best UX I could find for running sandboxed agents, the "kit/mixin" concepts are neat and I make use of them too: https://github.com/shaftoe/sbx-template-pi#stacking-the-extr...
As for Docker Sandboxes, I'll just ask Sol literally right now to see what it does better than my virtdev, and then I'll improve virtdev instead of using Docker.
Just the general knowledge that sharing a kernel with untrusted software is too dangerous, that hardware virtualization is an infinitely smaller attack surface and that the entire industry will be in deep shit if people or AI breaks hypervisors.
Initial threat model was supply chain attacks but eventually grew to include AI harnesses as well. Not very worried about them hacking me, more about accident prevention.
So that means each VM must be running a completely independent kernel that's fully isolated from the host's file system. They must also have fail closed network filtering built in.
> Anything you can point the rest of us to?
I have published my virtdev's design document.
https://github.com/matheusmoreira/virtdev/blob/master/DESIGN...
Yes, it is AI generated.
In summary, it's a QEMU VM orchestrator with a base OS image and project specific delta images. VM lifecycle is managed by systemd. System level isolation is already pretty good and it already solves the "AI wiped out my $HOME" problem. I'm currently working on a custom network stack to replace the nftables based firewall.
And of course, instead of doing the "copy in > copy out" process manually, get your local agent to write a bash script that does that for you, given what directory you're in, and you're basically G2G.
Start by mounting just your repo and passing in the keys for the agent. Take it from there, it's like software engineering, you iterate.
When you run into issues you expand the tools in the container available to it.
https://apps.apple.com/app/aifcc-ai-first-computer/id6782364...
als has lots of agents + and typical dev packages (node tooling, python tooling, ....) preinstalled
Stop putting sensitive stuff there.
Maybe we should just ssh into separate development machines to ensure real and verifiable sandboxing? (as was totally standard before Docker became a thing)
I use it regularly to run Claude/Codex with permission checks disabled.
https://github.com/cvhariharan/mvm
Better sandboxing for AI agents is exactly the main reason for containers improvements on macOS and Windows, with a few talks at WWDC, and BUILD.
Not sure how much they would get from Linux users then.
Yes, you can inject tokens via a proxy. What else is new?
I have skimmed alternatives offered in comments to this post (vibepod-cli, code-on-incus, opencode-docker, sandboxy, smolvm, amazing-sandbox) and none of them seem to do credentials injection at the proxy level.
Also fnox now does credentials proxying.
Ideally, they should run in _different_ sandboxes.
The environment might corrode the harness (e.g. rogue npm/pip packet would manipulate agent harness config).
Like:
$ podman run -it --rm -v .:/workspace local-dev-ia /usr/bin/oc
Configured with a .env file. Hope to do it hopefully before the end of the week.
This is nowhere near "exactly this". Docker Sandboxes uses micro VMs, you just use regular containers which have completely different security properties.
This is not what the person I was responding to is doing, though.
As for differences between the krun OCI runtime and Docker Sandbox (which also uses libkrun), let's please continue the discussion here: https://news.ycombinator.com/item?id=49240662 .
Docker is always a pain to use and this way I don't have to re-install everything a billion times for every different project.
Currently, I don't allow the agent access to docker, start docker myself, and then do short-lived sandbox-free sessions when the agent needs to do things that interact directly with docker; but that's annoying.
For those who do not trust
AND do not want to use some other, free VM for some reason?container run --rm -it -v "$(pwd)":/work -w /work myaiimage /bin/bash
'su agent'
'curl domain/install.sh | sh'
'runagent'
The option left is to use SSH to sign commits which is a no-go for a different reason.
Other than the login problem, it’s a decent option.
├── bin
│ └── sbx
├── libexec
│ ├── containerd-shim-nerdbox-v1 <- Nerbox integration for ContainerD
│ ├── mkfs.erofs
│ ├── mkfs.ext4
│ ├── nerdbox-kernel-arm64
│ └── nerdbox-rootfs-arm64.erofs
More info about Nerdbox is here https://github.com/containerd/nerdbox
https://engine.build/lab/agent-sandboxes
The open source section specifically.
Or do you mean something else?
Here's what I want: REALTIME OBSERVABILITY/POWERPOINT.
I don't want to just see what command it ran. I need graphics... what part of the file system it is touching, what network entities it is contacting. If it's running SQL I want the parsed query handed to me in a syntax highlighted and well formatted interface. Imagine that star trek computer presenting automated infographics while someone is doing a presentation, you know what I'm talking about? It's like a automated powerpoint as the agent does it's thing.
I need to understand my agent and what it typically does so I can dangerously wield it. I treat the agent like a gun in a live shooting scenario. That's how I want to use the LLM.
Sandboxes have their purpose. Just like how shooting ranges have their purposes. But I need to fire my gun in the real world and real world is a warzone.
The other url is their marketing page.
Yes, Linux is supported.
Don't give it shell access, just predefined tools.
One of the demos I run is how easy it is to circumvent the harness limits. For example, I can configure a harness not to access file `secrets.txt`. But, then I can immediately have it create a Python file that can read any file and have it read `secrets.txt`.
At the end of the day, "please" isn't security. You want to know that the agent can only do and access the things it should access.
[1] https://github.com/apple/container