> I wanted specifically to find a minimal set of restrictions to run untrusted code.
I don't think we should consider containers to be a security boundary. Even full VMs can be escaped, and have been, many times.
The fact that this is possible in the first place makes me think we need a much better approach.
As far as I know, Firecracker, gVisor, and Kata Containers are the solution here. They use VM primitives (x64_64 and ARM64 extensions) and have lighter codebases
https://firecracker-microvm.github.io/
But I don't have any direct experience with any of them. I'd be curious what people who have built on top of them think
edit: OK it looks like Kata can use Firecracker, so as far as isolation, it's either Firecracker or gVisor. And Firecracker is the VMM I mentioned, but gVisor is quite different -- it's more like a user space kernel that emulates syscalls.
Depends on the threat model. Security is not black-and-white.
Containers protect against "I don't trust this curlpipe to not crap all over my dotfiles," rather than, "there might be a sandbox escape attack in this random file I downloaded."
If a VM is not sufficient for your threat model, I'm curious what is?
Podman supports using KVM backed virtualization for containers via libkrun: `podman run --runtime=krun` . Still not the end-all-be-all security boundary, but better I think.
I definitely wouldn't trust standard Linux style containers that expose a shared Linux kernel at the moment, there's been far too many LPE and container breakout vulnerabilities this year. It's possible that in future if the kernel gets a lot more hardened, that could change but things like Firecracker are a better bet from a security standpoint.
They are a security boundary, but like everything else, not perfect.
I think until something hardware-based like CHERI becomes widely deployed (which seems extremely unlikely in the near to mid term given), we're going to keep seeing VM escape CVEs pop up indefinitely.
We don't have any "security boundaries" by this definition, just "security make-it-harder"s. I.e. "security boundaries" always have a relative strength associated with them, not a guarantee they keep the thing secure without any doubts.