Getting started
GLRMask is a grammar-constrained generation library for high-throughput LLM decoding. It is optimized for fast next-token mask computation with extremely low tail latency, even for complex grammars.
Installation
python -m pip install glrmask
Usage
GLRMask compiles a grammar and vocabulary into a Constraint. The resulting Constraint can be serialized and cached for reuse across requests.
At runtime, call constraint.start() to initialize a ConstraintState. In the decoding loop, run state.mask() in parallel with the model’s forward pass so the mask is ready in time for sampling. Then apply the mask to the logits, sample a token, and call state.commit_token(token_id) to advance the state.
state = constraint.start()
while generating:
# Run these in parallel.
logits = llm.forward(...)
mask = state.mask()
logits = apply_mask(logits, mask)
token_id = sample(logits)
state.commit_token(token_id)
For constraints that will not be reused enough to justify full compilation, DynamicConstraint is a drop-in replacement for Constraint that starts much faster, but leaves more work in the token loop and hence generates masks more slowly.
Python quickstart
python -m pip install glrmask llama-cpp-python torch
import numpy as np
from llama_cpp import Llama
from torch import from_numpy
from torch.distributions import Categorical
import glrmask
llm = Llama(model_path="model.gguf", logits_all=True)
vocab = glrmask.Vocab.from_llama_cpp(llm)
end_token_ids = vocab.llama_cpp_end_token_ids
end_tokens = set(end_token_ids)
get_logits = lambda: llm.scores[llm.n_tokens - 1]
sample = lambda logits: Categorical(logits=from_numpy(logits)).sample().item()
prompt = "Classify this review: The story dragged badly. Sentiment: "
input_tokens = llm.tokenize(prompt.encode())
MAX_OUTPUT_TOKENS = 64
Without constraints
llm.reset()
llm.eval(input_tokens)
generated = []
for _ in range(MAX_OUTPUT_TOKENS):
logits = get_logits()
token = sample(logits)
llm.eval([token])
generated.append(token)
if token in end_tokens:
break
print(llm.detokenize(generated).decode())
With GLRMask
schema = '{"type":"string","enum":["positive","negative","neutral"]}'
constraint = glrmask.Constraint.from_json_schema(schema, vocab)
llm.reset()
llm.eval(input_tokens)
state = constraint.start()
generated = []
for _ in range(MAX_OUTPUT_TOKENS):
logits = get_logits()
mask = state.mask(llm.n_vocab())
if state.is_accepting():
mask[end_token_ids] = True
logits[~mask] = -np.inf
token = sample(logits)
llm.eval([token])
generated.append(token)
if token in end_tokens:
break
state.commit_token(token)
print(llm.detokenize(generated).decode())
Rust quickstart
Constraint is the normal compiled Rust type:
use glrmask::{Grammar, Constraint, Vocab};
let vocab = Vocab::new(vec![
(0, b"\"yes\"".to_vec()),
(1, b"\"no\"".to_vec()),
]);
let schema = r#"{"type":"string","enum":["yes","no"]}"#;
let constraint = Constraint::compile(Grammar::json_schema(schema), &vocab)?;
let mut state = constraint.start();
let mask = state.mask();
state.commit_token(0)?;
if state.is_accepting() {
// The current prefix may validly end here.
}
if state.is_rejected() {
// No valid continuation remains.
}
# Ok::<(), glrmask::Error>(())
DynamicConstraint::compile(...) also accepts a Grammar and returns a DynamicConstraint. Its start() method returns a DynamicConstraintState with the same decoding methods as ConstraintState.
Grammar::bind_grammar(...) binds a child source grammar before a vocabulary is chosen:
let grammar = Grammar::glrm(
"glrm 1; start start; extern grammar payload; nt start = payload;",
)
.bind_grammar("payload", Grammar::json_schema(r#"{\"type\":\"null\"}"#))?;
let constraint = Constraint::compile(grammar, &vocab)?;
# Ok::<(), glrmask::Error>(())
For exact token IDs or compiled child constraints, build a ConstraintSpec:
use glrmask::{ConstraintSpec, Grammar, Constraint, Vocab};
let vocab = Vocab::new(vec![
(0, b"{".to_vec()),
(1, b"}".to_vec()),
(2, b"null".to_vec()),
]);
let child = Constraint::compile(
Grammar::json_schema(r#"{"type":"null"}"#),
&vocab,
)?;
let source = r#"
glrm 1;
start document;
extern token CONTROL;
extern grammar payload;
nt document = CONTROL "{" payload "}";
"#;
let spec = ConstraintSpec::builder(Grammar::glrm(source), &vocab)?
.bind_token("CONTROL", [32001])?
.bind_grammar("payload", &child)?
.build()?;
let constraint = spec.compile()?;
let dynamic_constraint = spec.compile_dynamic()?;
let mut state = constraint.start();
# Ok::<(), glrmask::Error>(())
ConstraintSpecBuilder::bind_grammar(...) accepts a Grammar, ConstraintSpec, Constraint, or DynamicConstraint child.
If the parent grammar is expensive and the child changes frequently, compile the parent with its extern grammar left unresolved, cache that Constraint, and bind compiled children later:
let mut parent = Constraint::compile(
Grammar::glrm(
"glrm 1; extern grammar payload; start document; nt document = payload;",
),
&vocab,
)?;
let child_a = Constraint::compile(Grammar::json_schema(schema_a), &vocab)?;
let child_b = Constraint::compile(Grammar::json_schema(schema_b), &vocab)?;
let with_a = parent.bind_grammar("payload", &child_a, &vocab)?;
let with_b = parent.bind_grammar("payload", &child_b, &vocab)?;
# Ok::<(), glrmask::Error>(())
The parent remains reusable. It can also be saved and loaded before binding; late binding consumes loaded parser automata and pooled weights directly from their packed representation, while small composition metadata is decoded lazily. A child passed to Constraint::bind_grammar(...) must already have all of its own external grammars bound.
Grammar formats
Unfortunately, there is no universally accepted EBNF dialect. In keeping with this tradition, GLRMask includes its own.
GLRM is GLRMask’s native grammar format. A grammar begins with glrm 1; and a start declaration:
glrm 1;
start value;
t NUMBER = /-?(0|[1-9][0-9]*)/;
nt value = NUMBER | "null";
Declarations use =, and epsilon is written as eps. Terminals and nonterminals can use fa { ... } bodies. Regexes use full-match semantics and reject unsupported or non-regular constructs. GLRMask also accepts Lark and EBNF.
Reusing compiled subgrammars
Declare a compiled child with extern grammar name; and bind it by name:
payload = glrmask.Constraint.from_json_schema(payload_schema, vocab)
document = glrmask.Constraint.from_glrm_grammar(
'''
glrm 1;
start document;
extern grammar payload;
nt document = "{" payload "}";
''',
vocab,
subgrammars={"payload": payload},
)
Inline g name = { ... }; and externally bound extern grammar name; use the same language semantics, including scope-local ignores.
Special tokens
Special tokens are declared by name and bound to their model token IDs outside the grammar:
grammar = '''
glrm 1;
start message;
extern token TOOL_CALL;
nt message = TOOL_CALL call;
nt call = "lookup()";
'''
constraint = glrmask.Constraint.from_glrm_grammar(
grammar,
vocab,
bindings={"TOOL_CALL": tool_call_token_id},
)
Bind the tool-call special token used by your model. Lark and EBNF use @token(<id>) for special tokens.
End tokens are handled by the decoder. If the constraint is accepting, generation may stop without committing another token.
mask = state.mask(model_vocab_size)
if state.is_accepting():
mask[end_token_ids] = True
token = sample_with_mask(logits, mask)
if token in end_tokens:
stop_generation()
else:
state.commit_token(token)
Saving compiled constraints
A compiled Constraint can be serialized and loaded again:
blob = constraint.save()
constraint = glrmask.Constraint.load(blob, vocab)
Load an artifact only with the exact vocabulary it was compiled against. Constraint::load() currently does not verify a vocabulary supplied separately by the caller. Composed constraints are saved as one artifact, including their child constraints.