Repository navigation
Docker Security Profiles (seccomp, apparmor, etc) #17142
Description
Activity
[RFC] Docker Security Profiles
The profile would be passed on docker run, we can reuse the flag we already have --security-opt.
Something like docker run ... --security-opt native:/path/to/config.toml ...
Obviously doesn't have to be toml since that's super hipster :p
Assumptions
- no one is going to sit and write out all the syscalls/capabilities their app needs
- automatic profiling would be super cool but like
aa-genprofit is never
perfect, leads to pain or removing the profile altogether, and an
unmaintainable config file (we can always attempt this later)
Goals
- maintainable config
- readable by humans and not a linux syscall/cap nerd
- something an app developer would want to write
- someone who did not write the config should be able to understand, at
least at a high level, what is restricted
Inspiration
Grouping into categories
High level things you would want to configure should be generic and limited
to (for example):
- Networking
- Filesystem (Disk)
- Runtime (CPU/Memory operations)
- User Operations
- Misc
Defining Permissions
The cool thing about tame
I think we should implement are what they refer to as "flags". It's a set of
syscalls that they allow for a common goal, such as TAME_RW will allow all
the syscalls for i/o operations but TAME_RPATH only allows the syscalls that
will enable read-only effects on the filesystem.
We can have this same concept and define them w syscalls and capabilities.
We would need to discuss what these were and find the most common use cases for
them.
Behaviors
- If one permission denies a syscall and another allows it, the
deny should always override the allow. - Passing an empty config drops everything and nothing is allowed
Super super super alpha example
Kinda like jfrazelle/bane but better.
[Networking]
Flags = [
# this will allow sendto(2), recvfrom(2), socket(2), connect(2)
"dns",
# adds CAP_NET_RAW
"ping"
]
[Filesystem]
Flags = [
# will allow lstat(2), chmod(2), chflags(2),
# chown(2), unlink(2), fstat(2) on /tmp
"tmp"
]
# filepaths where you would like to log on write
LogOnWrite = [
"/etc/**",
"/root/**"
]
# read-only filepaths
ReadOnly = [
"/sys/**"
]
[Runtime]
Flags = [
# allows getentropy(2), madvise(2), minherit(2),
# mmap(2), mprotect(2), mquery(2), munmap(2)
"malloc"
]
[User]
Flags = [
# allows getuid(2), getgid(2), setuid(2), setugid(2)
"create"
]Backends
Will use whatever is installed on the system so if they have apparmor but no seccomp, then it will use apparmor (which can technically do all the syscall, cap, and filesystem privileges).
- AppArmor
- Seccomp
- Capabilities
File Globbing
Taken from apparmor profiles file globbing.
| Glob Example | Description |
|---|---|
/dir/file |
match a specific file |
/dir/* |
match any files in a directory (including dot files) |
/dir/a* |
match any file in a directory starting with a |
/dir/*.png |
match any file in a directory ending with .png |
/dir/[^.]* |
match any file in a directory except dot files |
/dir/ |
match a directory |
/dir/*/ |
match any directory within /dir/ |
/dir/a*/ |
match any directory within /dir/ starting with a |
/dir/*a/ |
match any directory within /dir/ ending with a |
/dir/** |
match any file or directory in or below /dir/ |
/dir/**/ |
match any directory in or below /dir/ |
/dir/**[^/] |
match any file in or below /dir/ |
/dir{,1,2}/** |
match any file or directory in or below /dir/, /dir1/, and /dir2/ |
More Goodness
- I think we should allow people to define their own
flags(or whatever we end up calling them). It could be cool to have a way to do it with adefineintext/templateI believe this is possible if it is implemented the way I am thinking ;)
This is great. Some thoughts on the config file:
It would be cool to take the way apparmor does file permissions (e.g. from the chromium profile,
... {
/lib/@{multiarch}/libgcc_s.so* mr,
/lib{,32,64}/libm-*.so* mr,
/lib/@{multiarch}/libm-*.so* mr,
...
}
but maybe allowing something more user-friendly for permission names (like read, write, etc). Default deny would be preferable, but that might not be the best option (maybe something the user can set, like AccessPolicy: whitelist or AccessPolicy:blacklist?).
It'd also be cool to have the logging support a similar scheme, so that something like
Log = [
"/etc/something/config:read",
"/var/run/something/**:write"
]
ah yay @kisom I definitely like the idea of something like :ro or :rw much like how it works for volumes
@jfrazelle having a shorthand is good for people who write a lot of these, but also having long names is easier for people to remember; both could probably be supported if a leading char is used distinguish short form from long form. Something like :lrw v. "lock read write".
yes for sure that makes sense
How can it do the file permissions with the apparmor way? for example,how to limit to write a file?
it will generate an apparmor profile, this is not just seccomp config it
will be a generic security profile with backends
On Tue, Oct 27, 2015 at 6:23 PM, keloyang [email protected] wrote:
How can it do the file permissions with the apparmor way? for example,how
to limit to write a file?—
Reply to this email directly or view it on GitHub
#17142 (comment).
👍 to whatever @jfrazelle says.
@anusha-ragunathan you should check this out.
for phase 1 see: #17989
42 remaining items
I like the idea, but I also like the idea of allowing the image to request more access, which could at least allow the docker client to report to the user that this container will not run without the following capabilitiy, or requires this syscall, or requires SELinux to be disabled.
right, any assistance to let user know some features are required so he can check what they are about and decide to enable them would be nice, to avoid security newbies to just enable everything by default
Yes, oddly I was having a discussion about this earlier today, and it was something that has come up before.
I was thinking of perhaps prototyping it by defining a security metadata schema that could define the necessary things, and then having a tool to read that and construct the run command.
I am not so sure about raising privileges, as a message saying "this container needs --cap-add SYS_ADMIN" to run might be abused to encourage people to run things with escalated privileges. Using it to drop privileges that are not needed seems ok though if the container has been labelled that way, eg the nginx image may just need NET_BIND_SERVICE and can drop all the other default capabilities.
Yes I love the idea of having the image run with less privs, but also preventing:
docker run ...
permission denied
Followed by
docker run --privileged ...
Success
And the user goes off running the containers without any security forever.
In both case this would indeed encourage users to run with extra privileges without taking care.
So need to make it clear about the risks.
"this container needs --cap-add SYS_ADMIN.
Your default configuration do exclude this capability, please use with care blah blah blah"
"
Another option would be to offer a link which explain each capability / seccomp role (good luck docker documentation team) and make it clear about the potential security risk.
@rhatdan @ndeloof: How about a nagging flag such as --allow-insecure for using dangerous options such as
- privileged
- cap_sys_admin
- user=root
(Whether supplied to docker run or provided in the container image).
docker could exit with an error and a message such as:
You are attempting to run the container with (dangerous flags). Please add
--allow-insecureto confirm you want to run the container without the default security. To read more about the secure use of docker, please visit http://... .
Funny indeed, I had that conversation with @justincormack. I think it'd make sense to allow the image-maintainer to specify what capabilities / profile is needed for the image to run, but it should not automatically apply those (the person running the image should be the one deciding if the container actually gets those permissions).
Perhaps;
docker run --security-opt seccomp:embedded
to run the image with the seccomp-profile that's embedded in the image.
Possibly even think of disallowing --cap-add and --security-opt, and only allowing running images with the embedded profile? (Using a whitelist of images / trusted sources). Haven't given it much thought yet, so needs more thinking :D
There was talk in #22109 of allowing Seccomp profiles to be layered, permitting more than one to be used at a time. I think this would be an ideal way to add image-specific Seccomp profiles without requiring users to opt into using just the profile embedded in the image, or just the global profile baked into the daemon. In some cases, applying both profiles will have no benefit (the image profile could well block every syscall the global profile does). Still, loading both doesn't require fully trusting the security profile baked into an image, which might be more insecure than the default profile.
This doesn't help in cases where the image requires a syscall blocked by the default filter, but the baked-in filter only restricts a few high-impact syscalls. I'd say this should be handled similarly to the suggestions above for handling images that want to add capabilities instead of remove them. Requiring a flag or similar seems like a good idea.
Layering wouldn't really work with other things one might embed in an image's security profile, though. You can't exactly layer SELinux or Apparmor labels on the same file, for example.
Well a good packager could have his PID1 do a lot of what we are talking about, drop caps for example. Problem is most container packagers don't control PID1 code, they just stick in something like httpd.
I am not crazy about blocking options from the user like --cap-add or --security-opt, worried about unexpected consequences.
It makes sense to have an embedded security profile into the container metadata but would be great to prioritize the docker daemon settings because, as a sec guy, you maybe want to enforce a minimum profile and allow someone to run a container with a "better" profile but deny the use of profiles allowing things the default don't allow.
This can work in three ways:
--seccomp:embedded -> try to run the container with the embedded profile comparing it with the default profile. if the embedded profile is less secure, stop.
--seccomp:merge -> try to run the container merging (with the layering proposed in #22109) the embedded profile with the default profile prioritizing the default opts creating a more secure profile.
--seccomp:force-embedded -> run the embedded profile ignoring the default. This only can occur if the docker daemon is explicitly configured to allow this.
Current proposal for that can be found here : #32801
I think I want to close this in favor of the (somewhat more concrete) #32801.
As mentioned in our ROADMAP.md, we'd like to progress toward seccomp support in Docker 1.10.
As a phase 1, I propose allowing the Engine to accept a seccomp profile at container run time. In the future, we might want to ship builtin profiles, or bake profiles in the images: design work about that future would be a plus.
Ping @jfrazelle who's interested to look into that!