I am pleased to announce the first release of the caps library, which
allows you to use capability types in your OCaml programs (and
libraries).
The main idea is that after you used this library, your functions
can have signatures like
val foo: < Cap.network; Cap.stdout; Cap.random; .. > ->
int -> float
meaning this function requires the network, stdout, and random
capabilities to work. The signature reveals the internal effect
this function has and the kind of system calls it internally does (or its callees),
I designed this library while working at Semgrep on the semgrep codebase
and it was useful to sandbox or control parts of the codebase so
that young engineers would not call dangerous functions in certain parts.
It is I think even more useful in the new coding-agent era to control
in the signature the code generated by AI.
Having myself abused phantom types in the past to check stuff statically in a couple of projects, I have mixed feelings with these capability approaches from an ergonomic point of view.
E.g. if I temporarily want to printf debug something locally in a nested context I don’t want to have to change all my capability sets, or worse need to add a new argument, up to the top-level, just to be able to do that temporary change.
Do you actually have design insights to share on the kind of capability granularity that is desirable? I feel that having a capability per system call, as you seem to suggest, it overkill.
I agree. This is why for example I consider the Logs library “ambient authority”, that is does not require capabilities (well it’s an external library, so it’s not covered by the semgrep rules imposed on the project so it’s easier, but even if it was a local library I could decide that it does not require capabilities). In fact for some times I had Cap.Console.stdout a capability but not Cap.Console.stderr because I thought we could be more laxist about stderr.
So in the different projects where I’m using capability I try to find the right balance. For instance, at the beginning I wanted to have fine-grained capability about the filesystem with Cap.FS.tmp, Cap.FS.homedir, etc. and I ended up with just Cap.open_in and Cap.open_out.
It all depends what you want to “sandbox/control” in the codebase. You don’t even have to use all the capabilities provided by Cap.ml and if you look at Cap.mli you’ll see all the syscalls are not there, just what I considered the most important for my projects.
There is also some Cap.tmp_caps_UNSAFE : unit -> < Cap.tmp > (and similar functions for other caps) that provides temporary solution to get capabilities outside the Cap.main function, but it should be just used as a last resort, and temporarily.
In this webapp implementation I distinguish sub services that can change the session and those that can’t at the type level because I wanted to make it easier to distinguish services that fiddle with that security aspect and those that do not (as can be seen in the service tree here).
Do you think that’s the kind of stuff you could express with a capability ? It seems a bit tricky here because it affects the signature of the service namely function from HTTP requests to response vs. from HTTP requests to session responses. I’m just wondering because assuming I’d be using your capabilities it would be nice to try to fit that in that framework too.
I don’t fully understand the Webapp world, but I guess you could have a
type cap_modify_session
abstract type representing the ability to modify a session and Session API that would required to be passed this capability if they would want to modify the session. Then when you would pass around sessions, if some code require to modify the session then you will clearly see in its signature the cap_modify_session being passed.
Note that one issue with object type and objects is that it’s not super easy to extend in a general way; So it’s hard to extend Cap.ml outside Cap.ml … For instance in Semgrep we use authorization token but Cap.ml knows nothing about that so we need to pass 2 arguments to some functions, the caps, and the token; you can define types that encapsulabe both, but in OCaml you can’t define generic function such as 'here is an object with 3 capabilities and more, and return an object with those 3 capabilities, and an auth token capabilities, and those “more”. Still, it is possible for some limited cases to combine 2 set of capabilities designed in separate libs.
In the same way currently there is a coarse grained Cap.exec capability representing the ability to call external programs; I would like at some point to refine this with a more precise cap_exec_git, cap_exec_find, etc. so you clearly see the actual program being called in the signature. Same for Cap.network where one might want to also constrain in the signature the actual host one could connect to. It is possible to define such capabilities outside Cap.ml, you just define new abstract types and the builder of this abstract type can be required to be passed the Cap.exec capability, making the new capability a kind of subtype. But I lack experience doing those things for now. I dunno the limitations it would have and “ergonomic” issues I could get. This is why for now I limit myself to coarse-graind syscalls, which has been already super useful to document, understand, and sandbox parts of the codebase.
it could be interesting to distribute caps not as a library but as a “template for a library”: something you vendor into your source tree, into which you modify an extra_capabilities file (newline separated valid identifiers, empty when distributed), and it generates the main mli and ml at build time which include the caps capabilities and then also your own.
this way you can use caps for the generic unix/sys capabilities and you can also add your project’s logic capabilities (session management, ro and rw access to some databases, etc.). you are responsible for adding the capability args in your local mlis.
It reminds me a long time ago (7 years ago) working on cstruct_cap (see the interface, the pull-request and the issue). orec is also a nice example of how to abuse phantom types. One fair conclusion I can draw from this experiment is that even those most convinced of the merits of capabilities do not use cstruct_cap. I’m not sure what to make of this, but I can relate to:
Regarding the “template for a library”, good idea. Well it’s also already available like that, you can just git clone the ocaml-caps repo and with dune things are easy to compose.
Yes Romain, at Semgrep lots of the other developers were also skeptical and didn’t like the friction added by capability types. I think you and Hannes also were not super convinced about the use of capability in the EIO library. Personally for me it’s a no brainer … I don’t even understand how I was programming without this library before It’s like when people complained about using types 20 years ago and preferred to use PHP/Ruby/Javascript … it was a no brainer to me static types were far better for most programming (especially on large codebase and with other developers where refactoring is so much safer with static types). Now all those dynamic languages have gradual typechecker (actually usually written in OCaml ). I expect the same will happen with capability types, especially in the coding-agent era where they are really really useful.
I think you’re missing the most crucial point: usability. The point isn’t to say that capabilities are bad (and you might be surprised by our view on this, given that you’ve included @hannes - and there is a reason why I suggested cstruct_cap), but rather to work out how to use and integrate them without causing friction. There are several ways to achieve this (and they’re not limited to simply providing an interface).
I agree usability is super important, but as I said in the talk, I used them in 4 different projects and found that their use didn’t require that much extra annotations, and while there are times where the error message is not super clear and fixing the capability typing error takes a little bit of time, I’ve found overall really useful to document and control a codebase. In fact it’s exactly the same thing with types … they add some friction, but they also document and constraint the code and prevent bugs (or bad practice).
I don’t think this really solves the problem. If you want that to be useful you’d need to be able do this in a compositional way (e.g. the library that provides you with sessions for web apps also provides the capabilities to control them).
I think that framing that as if it’s the typed vs unityped debate under another form misses the mark. The debate here is rather how precise should your types be and how this precision interacts with abstraction. There is a delicate and non-obvious balancing act to make here.
For example I had my share of let’s express capabilities via phantom object types in the browser APIs provided by js_of_ocaml. Cool hack but unusable.