# Deterministic hashing

**URL:** <https://discuss.ocaml.org/t/deterministic-hashing/18308>\
**Category:** Learning\
**Created:** [June 27, 2026, 2:35pm UTC](https://discuss.ocaml.org/t/deterministic-hashing/18308 "2026-06-27T14:35:12Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![kentookura](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ocaml.org/kentookura/32/5187_2.png) [@kentookura](https://discuss.ocaml.org/u/kentookura)\
**Post date:** [June 27, 2026, 2:35pm UTC](https://discuss.ocaml.org/t/deterministic-hashing/18308/1 "2026-06-27T14:35:12Z")

</div>

I have been using `Hashtbl.hash` to generate unique ids from a certain datatype in my program. The problem with this is that between runs of the program the same input yields a different hash. This causes cram tests to fail even when the program hasn’t changed.

Is it possible to use `Hashtbl.seeded_hash` for this purpose? The docs say that the function “is **further** parametrized by an integer seed” (emphasis mine), so assume I will see the same nondeterministic behavior.

I have already come up with an alternative solution that is particular to my application but from an architectural perspective using a hash-like function is much cleaner.

Eager to hear any suggestions. Thanks!

---

<div class="post-metadata">

**Author:** ![K\_N](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ocaml.org/k_n/32/2793_2.png) [@K\_N](https://discuss.ocaml.org/u/K_N)\
**Post date:** [June 27, 2026, 3:06pm UTC](https://discuss.ocaml.org/t/deterministic-hashing/18308/2 "2026-06-27T15:06:48Z")

</div>

This is surprising because, according to the [code](https://github.com/ocaml/ocaml/blob/trunk/stdlib/hashtbl.ml#L541-L546) `Hashtbl.hash` calls `Hashtbl.seeded_hash` with a fixed seed of `0`.  
But looking at [caml\_hash](https://github.com/ocaml/ocaml/blob/trunk/runtime/hash.c#L258-L278) everything (for a fixed seed) should be deterministic, unless the data structure you hash contains closures. In this case the code pointer is part of the hash and may yield different results for different runs. Or if you are using objects, then their internal id is used for the hash, but their internal id may change depending on creation order.

Edit : the following shows that the issue might indeed be code pointer used in hash + ASLR :

```ocaml
let f x y = x + y

let () = Format.printf "%d\n%!" (Hashtbl.hash f)

```

```shell
$ ocamlopt foo.ml
$ ./a.out; ./a.out; ./a.out
545623849
393380115
243345998
$ setarch -R ./a.out ; setarch -R ./a.out ; setarch -R ./a.out 
1014361679
1014361679
1014361679

```

---

<div class="post-metadata">

**Author:** ![kentookura](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ocaml.org/kentookura/32/5187_2.png) [@kentookura](https://discuss.ocaml.org/u/kentookura)\
**Post date:** [June 27, 2026, 4:14pm UTC](https://discuss.ocaml.org/t/deterministic-hashing/18308/3 "2026-06-27T16:14:39Z")

</div>

Converting to a string and using the [digest module](https://ocaml.org/manual/5.3/api/Digest.html) might suffice? The data structure is not the issue.

---

<div class="post-metadata">

**Author:** ![kentookura](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ocaml.org/kentookura/32/5187_2.png) [@kentookura](https://discuss.ocaml.org/u/kentookura)\
**Post date:** [June 27, 2026, 4:33pm UTC](https://discuss.ocaml.org/t/deterministic-hashing/18308/4 "2026-06-27T16:33:55Z")

</div>

A faulty assumption on my end may be that the data structures I am passing in are really the same between runs of the program. It is hard to verify if there is nondeterminism here, but this seems increasingly likely to me.

---

<div class="post-metadata">

**Author:** ![conroj](https://avatars.discourse-cdn.com/v4/letter/c/7c8e57/32.png) [@conroj](https://discuss.ocaml.org/u/conroj)\
**Post date:** [June 28, 2026, 4:27pm UTC](https://discuss.ocaml.org/t/deterministic-hashing/18308/5 "2026-06-28T16:27:33Z")

</div>

> [@kentookura](#):
>
> This causes cram tests to fail even when the program hasn’t changed.

Since you raise this question in the context of automated testing, I assume that you’ll want the hashes to be consistent across revisions to the standard library? In general, a hash function interface that doesn’t promise to be stable for all eternity won’t be.

A good [fingerprint function](https://en.wikipedia.org/wiki/Fingerprint_(computing)) is one that has a fixed specification, and the hashes found in the Digest module seem like they qualify. [Duff](https://github.com/mirage/duff) also appears to implement Rabin’s algorithm.
