Overview
GLRMask is a grammar-constrained generation library for high-throughput LLM decoding. It is optimized for fast next-token mask computation with extremely low tail latency, even for complex grammars.
GLRMask achieves this by doing as much constraint work as possible ahead of generation. It compiles a grammar against a model vocabulary into a Constraint. A compiled Constraint is immutable and reusable, but compilation takes time: tens of milliseconds for a moderately complex JSON Schema, hundreds of milliseconds for a very complex one, and up to a few seconds for a full programming-language grammar.
GLRMask also has a dynamic mode. DynamicConstraint is built from the same grammar and model vocabulary, but defers more of the constraint work until generation, so startup is much faster.
GLRMask accepts JSON Schema and general context-free grammars in EBNF, Lark, and GLRM. GLRM also supports composition from separately compiled subgrammars, including tokens that cross parent/child boundaries.
Benchmarks
TBM (time between masks) is the per-token cost of computing each constraint mask during generation. TTFM (time to first mask) includes constraint setup and the first mask.