# Recommendations for parsing structured text files with specific annotations/commands

**URL:** <https://discuss.ocaml.org/t/recommendations-for-parsing-structured-text-files-with-specific-annotations-commands/10272>\
**Category:** Ecosystem\
**Tags:** markdown, library\
**Created:** [August 7, 2022, 10:16am UTC](https://discuss.ocaml.org/t/recommendations-for-parsing-structured-text-files-with-specific-annotations-commands/10272 "2022-08-07T10:16:42Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![dmentre](https://avatars.discourse-cdn.com/v4/letter/d/c4cdca/32.png) [@dmentre](https://discuss.ocaml.org/u/dmentre)\
**Post date:** [August 7, 2022, 10:16am UTC](https://discuss.ocaml.org/t/recommendations-for-parsing-structured-text-files-with-specific-annotations-commands/10272/1 "2022-08-07T10:16:42Z")

</div>

Hello,

I would like to parse structured text files with some specific commands inside them to obtain an OCaml Sum type that I will use in my program. For now, I’m using text files in markdown format with additionally LaTeX-like commands (e.g. `\mycommand{arg1, arg2}`).

For parsing markdown I plan to use [omd](https://github.com/ocaml/omd) with some regexp to parse my commands.

Do you have any recommendations regarding libraries/frameworks to use for such a job? As anybody done similar things?

I’m not specially tied to Markdown or Latex-like commands and will happily switch to other formats if a library already provides necessary parsing.

Best regards,  
david

---

<div class="post-metadata">

**Author:** ![lindig](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ocaml.org/lindig/32/532_2.png) [@lindig](https://discuss.ocaml.org/u/lindig)\
**Post date:** [August 7, 2022, 10:36am UTC](https://discuss.ocaml.org/t/recommendations-for-parsing-structured-text-files-with-specific-annotations-commands/10272/2 "2022-08-07T10:36:51Z")

</div>

I would look into Angstrom. It is very flexible and combines traditional scanning and parsing. Classic scanning/parsing using lex/yacc only works with languages that are designed that way. Many real-world languages don’t separate scanning and parsing well enough.

---

<div class="post-metadata">

**Author:** ![darrenldl](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ocaml.org/darrenldl/32/291_2.png) [@darrenldl](https://discuss.ocaml.org/u/darrenldl)\
**Post date:** [August 7, 2022, 11:28am UTC](https://discuss.ocaml.org/t/recommendations-for-parsing-structured-text-files-with-specific-annotations-commands/10272/3 "2022-08-07T11:28:39Z")

</div>

I tend to use MParser if you want to retain line number information, but Angstrom is perfectly good indeed.

---

<div class="post-metadata">

**Author:** ![Chet\_Murthy](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ocaml.org/chet_murthy/32/1501_2.png) [@Chet\_Murthy](https://discuss.ocaml.org/u/Chet_Murthy)\
**Post date:** [August 7, 2022, 9:18pm UTC](https://discuss.ocaml.org/t/recommendations-for-parsing-structured-text-files-with-specific-annotations-commands/10272/4 "2022-08-07T21:18:54Z")

</div>

If you can arrange for prefix/suffix strings to come from some well-understood fixed and small set, and that they are not present in the text of your specific commands, I’d go with regexp to pull out your specific commands as strings, and then … well, anything you want to parse them.

If your example `\mycommand{arg1,arg2}` is an example of what you need to parse, then it’s harder, b/c “}” is too common. But perhaps if you just scan for parenthesis-matching, you can use that to know where the end of the command is.

What I’m saying is: it can be a lot of trouble to parse your entire file using some parser, only to extract some little bits. In your case, that would be running a markdown parser. So if there’s a way to regard the file as just text, and find your particular strings in some other way, that can be effective.

---

<div class="post-metadata">

**Author:** ![keleshev](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ocaml.org/keleshev/32/3779_2.png) [@keleshev](https://discuss.ocaml.org/u/keleshev)\
**Post date:** [August 8, 2022, 12:47pm UTC](https://discuss.ocaml.org/t/recommendations-for-parsing-structured-text-files-with-specific-annotations-commands/10272/5 "2022-08-08T12:47:59Z")

</div>

At some point I had a similar idea of of porting my blog to Scribble language that looks like this:

[1 Getting Started (racket-lang.org)](https://docs.racket-lang.org/scribble/getting-started.html#%28part._first-example%29)

```auto
#lang scribble/base
 
@title{On the Cookie-Eating Habits of Mice}
 
If you give a mouse a cookie, he's going to ask for a
glass of milk.
 
@include-section["milk.scrbl"]
@include-section["straw.scrbl"]

```

Of course, to do that, I started with writing a parser for Scribble in OCaml. And to do that I—of course—started with making a parser combinator library… I got quite far, but abandoned it in favor of using Pandoc.

Here’s my incomplete Scribble parser, if you’re looking for an inspiration: [new.keleshev.com/scribble at master · keleshev/new.keleshev.com · GitHub](https://github.com/keleshev/new.keleshev.com/tree/master/scribble)

---

<div class="post-metadata">

**Author:** ![dmentre](https://avatars.discourse-cdn.com/v4/letter/d/c4cdca/32.png) [@dmentre](https://discuss.ocaml.org/u/dmentre)\
**Post date:** [August 28, 2022, 4:57pm UTC](https://discuss.ocaml.org/t/recommendations-for-parsing-structured-text-files-with-specific-annotations-commands/10272/6 "2022-08-28T16:57:40Z")

</div>

Thank you @Chet_Murthy for pointing out that the commands should be different enough of the regular text to be easy to parse! I think this is my case but I’ll keep that in mind.

---

<div class="post-metadata">

**Author:** ![dmentre](https://avatars.discourse-cdn.com/v4/letter/d/c4cdca/32.png) [@dmentre](https://discuss.ocaml.org/u/dmentre)\
**Post date:** [August 28, 2022, 4:59pm UTC](https://discuss.ocaml.org/t/recommendations-for-parsing-structured-text-files-with-specific-annotations-commands/10272/7 "2022-08-28T16:59:56Z")

</div>

Thank you @lindig, @darrenldl and @keleshev for your suggestions. It seems to me a bit overkill for now but I’ll keep that in mind (Angstrom, MParser) in case my first approach does not work.
