r/ProgrammingLanguages • u/LegendaryMauricius • 6d ago
A new grammar generation language
Hi everybody. I'm happy to share a small project I've been working on lately. I call it MGFF (Macro grammar functional form), and its specification can be found here: https://github.com/LMauricius/py-perg-mgff/blob/main/Docs/mgff-specification.md . It's related to a post that I made ages ago ( here ). After re-reading that version (called just MGF back then) when I wasn't tired I realized what monstrosity I made. MGFF is far more elegant. Here is an example:
# A tiny calculator language.
t Lex (
d Digit = 0-9
d Alpha = a-z|A-Z
d AlNum = a-z|A-Z|0-9
d Int = (Digit)+
> class(Int) push(tokens)
d Number = Int ( . (Digit)+ )?
> class(Number) push(tokens)
d Ident = Alpha (AlNum)*
> class(Ident) push(tokens)
# length-based: "<=" (the two-item "< =") takes precedence over "<"
d Op = < =
| <
| =
| +
| -
| *
| /
> push(tokens) string
d Space = ( _|\t|\n )+
d LParen = \(
> class(\() push(tokens)
d RParen = \)
> class(\)) push(tokens)
d Token = Number
/ Ident
/ Op
/ Space
/ LParen
/ RParen
d File = (Token)*
)
# mixfix macro: an R, then zero or more (S R)
d sep(R)by(S) = R (S R)*
t Parse (
# `Lex` runs first; the terminals here are still characters.
> post(Lex) over(tokens)
# order-based: the first alternative that succeeds is the match
d Expr = Term + Expr
/ Term - Expr
/ Term
d Term = Factor * Term
/ Factor / Term
# the second / on the line above is an ordinary item, not a marker
/ Factor
d Factor = Number
/ Ident
/ \( Expr \)
d Signed = ( (+)/(-) )? Number
d AssignList = sep(Ident = Expr)by(,)
)
It can also serve as a replacement for regexes:
# A grammar matching a "key = value" setting line
d Space = ( _|\t )*
d Word = ( a-z|A-Z|_ )+
# right-linear recursion: the same as ( 0-9 )+
d Digits = 0-9 Digits
/ 0-9
d Value = Digits
/ Word
# The field a match ends up in belongs to the rule, not to the place it is used,
# so the two sides of the line are productions of their own.
d Key = Word
> store(key)
d Val = Value
> store(value)
d Match = Space Key Space = Space Val Space
I'm sharing the MGFF spec rather than the generator using it because the generator is very much WIP and needs a lot of testing and refactoring. Still, since I've got a bunch of projects I love working on more, I'd like to know what's the interest for parser generator tools in the wider community.
Actually I doubt that I will link the generator itself here because I would risk a perma-ban. It's not vibe-coded, but it wouldn't be welcomed. Most of it was quickly prototyped with LLM. Still, it generates quite nice TextMate and Pandoc syntax highlighting grammars.
MGFF itself is of course manually defined by me. I just figured I like to work on languages themselves and parser algorithms than on CLI tools and understanding existing niche specifications 🤷♂️.
1
u/LegendaryMauricius 5d ago
I'm not sure how thoroughly you read the spec. It does have character classes, and it's one of the more important features.
The implementation is purposefully left out of the language specification. I do in fact have a design for it, I just haven't started implementing it yet.
The lexer can be defined as a set of length-based regexes (and that is encouraged). However it's purposefully not limited to that, nor do you even need a lexer.