Grammar reference
Grammar formats and syntax.GLRMask accepts JSON Schema, GLRM, Lark, and EBNF.
Constraint constructors
| input | Python constructor | typical use |
|---|---|---|
| JSON Schema | Constraint.from_json_schema(...) | structured JSON, tools, typed API responses |
| GLRM | Constraint.from_glrm_grammar(...) | native grammar/lexer control, reusable subgrammars |
| Lark | Constraint.from_lark(...) | existing Lark grammars |
| EBNF | Constraint.from_ebnf(...) | conventional CFG descriptions |
DynamicConstraint has matching constructors.
JSON Schema
GLRMask supports a subset of JSON Schema. Unsupported features may be rejected.
GLRM
GLRM is GLRMask’s native grammar format. A grammar begins with glrm 1; and a start declaration. Rules use = and end with ;.
glrm 1;
start value;
t WS = /[ \t\r\n]+/;
ignore WS;
t NUMBER = /-?(0|[1-9][0-9]*)/;
nt value = NUMBER | "null";
Use explicit eps for epsilon. Regex terminals use full-match semantics. Unsupported or non-regular regex constructs are rejected.
External subgrammars
Declare a child grammar by name and bind a compiled constraint in Python:
extern grammar payload;
payload = glrmask.Constraint.from_json_schema(payload_schema, vocab)
document = glrmask.Constraint.from_glrm_grammar(
grammar,
vocab,
subgrammars={"payload": payload},
)
Inline and externally bound subgrammars have the same semantics:
g inner = {
start value;
nt value = "null";
};
Special tokens
To use a special token in a GLRM grammar, declare it by name:
glrm 1;
start message;
extern token TOOL_CALL;
nt message = TOOL_CALL call;
nt call = "lookup()";
Bind it outside the grammar:
constraint = glrmask.Constraint.from_glrm_grammar(
grammar,
vocab,
bindings={"TOOL_CALL": tool_call_token_id},
)
Direct finite automata
A terminal or nonterminal can use an explicit finite automaton body:
t WORD = fa {
start begin;
accept done;
begin -> middle: "a";
middle -> done: "b";
};
End tokens
End tokens are handled by the decoder. If the constraint is accepting, generation may stop without committing another token.