# \[mystery solved\] Compiling with continuations: flambda and performance

**URL:** <https://discuss.ocaml.org/t/mystery-solved-compiling-with-continuations-flambda-and-performance/6621>\
**Category:** Ecosystem\
**Tags:** lwt, monads, flambda, concurrency, async\
**Created:** [October 16, 2020, 7:21pm UTC](https://discuss.ocaml.org/t/mystery-solved-compiling-with-continuations-flambda-and-performance/6621 "2020-10-16T19:21:23Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![Chet\_Murthy](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ocaml.org/chet_murthy/32/1501_2.png) [@Chet\_Murthy](https://discuss.ocaml.org/u/Chet_Murthy)\
**Post date:** [October 16, 2020, 7:21pm UTC](https://discuss.ocaml.org/t/mystery-solved-compiling-with-continuations-flambda-and-performance/6621/1 "2020-10-16T19:21:23Z")

</div>

Aha. Needed a little more care with the placing of abstractions for continuations.

Erasing post, b/c now it all works as desired.

Restoring the deleted post:

> I was recently chewing over monadic concurrency, and got to wondering what the performance difference is, between writing in direct style, and writing in the -simplest- kind of monadic style I could imagine: just continuations in tail position.
> 
> So I wrote a benchmark: [event/tests/test\_ab.ml at master · chetmurthy/event · GitHub](https://github.com/chetmurthy/event/blob/master/tests/test_ab.ml)
> 
> It’s a simple test with two modes: “direct” and “kont”. In both modes, it loops thru a buffer, fetching a character and bitwise-or-ing it into an accumulator (to simulate -using- the fetched byte, and hopefully forcing the optimizer and hardware to wait for that fetch). And I’m using the 4.11.1+flambda compiler, with “-O3 -unbox-closures”, so hopefully I’m getting excellent inlining and allocation-avoidance.
> 
> And yet, direct-style is 5x faster than CPS on this (admittedly micro)benchmark. I wonder why … It’s been decades since I went down into the assembler … does anybody have any ideas on what might be going wrong here?
> 
> ```auto
> ./test_ab.opt direct 1024000
> 258MiB/sec: 1024000 read in 0.003792 secs
> 
> ./test_ab.opt kont 1024000
> 50MiB/sec: 1024000 read in 0.019679 secs
> 
> ```
> 
> P.S. I guess, for such a simple benchmark, I was hoping that the compiler could inline enough to discover the loop. Ah, well.

---

<div class="post-metadata">

**Author:** ![c-cube](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ocaml.org/c-cube/32/1727_2.png) [@c-cube](https://discuss.ocaml.org/u/c-cube)\
**Post date:** [October 16, 2020, 8:55pm UTC](https://discuss.ocaml.org/t/mystery-solved-compiling-with-continuations-flambda-and-performance/6621/2 "2020-10-16T20:55:35Z")

</div>

In my experience, even for algorithmic code that is more naturally expressed in an imperative way (say, with a lot of integer arrays), loops beat direct tail-recursive functions, even without any CPS encoding. ie. `while … ` or `for …` with local references tends to be faster than `let rec loop …`. That makes me a bit sad because OCaml doesn’t offer a lot of imperative control flow so the code tends to be uglier this way.

---

<div class="post-metadata">

**Author:** ![keigoi](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ocaml.org/keigoi/32/2451_2.png) [@keigoi](https://discuss.ocaml.org/u/keigoi)\
**Post date:** [October 16, 2020, 9:44pm UTC](https://discuss.ocaml.org/t/mystery-solved-compiling-with-continuations-flambda-and-performance/6621/3 "2020-10-16T21:44:05Z")

</div>

Could you please recover the original post? I remember that it described flambda’s optimisation of a (continuation?) monad. I am also interested in optimising function monads like this (e.g. a state monad  
for linearity checking [https://github.com/keigoi/linocaml](https://github.com/keigoi/linocaml) ). Thanks!

---

<div class="post-metadata">

**Author:** ![Chet\_Murthy](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ocaml.org/chet_murthy/32/1501_2.png) [@Chet\_Murthy](https://discuss.ocaml.org/u/Chet_Murthy)\
**Post date:** [October 16, 2020, 9:51pm UTC](https://discuss.ocaml.org/t/mystery-solved-compiling-with-continuations-flambda-and-performance/6621/4 "2020-10-16T21:51:53Z")

</div>

Argh, I’m sorry, I didn’t save a copy, and I don’t know if this discussion board software has a way to recover deleted text. But in any case, the benchmark I wrote is here: [https://github.com/chetmurthy/event/blob/master/tests/test\_ab.ml](https://github.com/chetmurthy/event/blob/master/tests/test_ab.ml)

I’m experimenting, trying to figure out the most-efficient form of monadic I/O, for this simplest example, before moving on to more-complex things.

---

<div class="post-metadata">

**Author:** ![dbuenzli](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ocaml.org/dbuenzli/32/18_2.png) [@dbuenzli](https://discuss.ocaml.org/u/dbuenzli)\
**Post date:** [October 16, 2020, 9:56pm UTC](https://discuss.ocaml.org/t/mystery-solved-compiling-with-continuations-flambda-and-performance/6621/5 "2020-10-16T21:56:18Z")

</div>

> [@Chet\_Murthy](#):
>
> recover deleted text

Recovered it from my emails in your original post.

---

<div class="post-metadata">

**Author:** ![vlaviron](https://avatars.discourse-cdn.com/v4/letter/v/ec9cab/32.png) [@vlaviron](https://discuss.ocaml.org/u/vlaviron)\
**Post date:** [October 16, 2020, 10:04pm UTC](https://discuss.ocaml.org/t/mystery-solved-compiling-with-continuations-flambda-and-performance/6621/6 "2020-10-16T22:04:07Z")

</div>

For your information: your CPS code is bad because it allocates closures all the time. To prevent this, you should eta-expand your definitions:

```auto
(* bad: [read1 n k] first computes [read1 n], then applies the result to [k] *)
let read1 n : int Kont.comp =
  Kont.return (Char.code (Bytes.get buffer (n mod 1024)))

(* good: [Kont.return] can be inlined, and [read1 n k] is a single direct application *)
let read1 n k =
  Kont.return (Char.code (Bytes.get buffer (n mod 1024))) k

```

This is particularly important for the recursive `readn`, which can be specialised for a given continuation if eta-expanded but not otherwise.  
So small experiments on my side show read speeds of around 1GB/s both with direct and cps styles (with flambda, and with some patches to speed up the loop bodies themselves).

---

<div class="post-metadata">

**Author:** ![Chet\_Murthy](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ocaml.org/chet_murthy/32/1501_2.png) [@Chet\_Murthy](https://discuss.ocaml.org/u/Chet_Murthy)\
**Post date:** [October 16, 2020, 10:05pm UTC](https://discuss.ocaml.org/t/mystery-solved-compiling-with-continuations-flambda-and-performance/6621/7 "2020-10-16T22:05:06Z")

</div>

Yes, this is what I figured-out. AKA: “writing in CPS form is hard to get right”. _grin_

---

<div class="post-metadata">

**Author:** ![Levi\_Roth](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ocaml.org/levi_roth/32/2268_2.png) [@Levi\_Roth](https://discuss.ocaml.org/u/Levi_Roth)\
**Post date:** [October 16, 2020, 10:22pm UTC](https://discuss.ocaml.org/t/mystery-solved-compiling-with-continuations-flambda-and-performance/6621/8 "2020-10-16T22:22:02Z")

</div>

> [@Chet\_Murthy](#):
>
> Argh, I’m sorry, I didn’t save a copy, and I don’t know if this discussion board software has a way to recover deleted text.

In the upper right-hand corner of the post, by the timestamp, there’s a little orange pencil icon showing that it was edited. If you click that, you can step through the edit history to see previous versions.
