Python API
Python API reference.Vocab
Vocab.from_llama_cpp(llm)
Build the byte vocabulary used by a llama-cpp-python model.
vocab = glrmask.Vocab.from_llama_cpp(llm)
end_token_ids = vocab.llama_cpp_end_token_ids
EOG, control, unused, and empty-piece tokens are omitted from the byte vocabulary. Their IDs remain available in llama_cpp_end_token_ids for the decoder.
Vocab.from_id_to_bytes(id_to_bytes)
Build from an explicit {token_id: bytes} mapping.
Vocab.from_dict(token_to_id)
Build from a {bytes: token_id} mapping.
A compiled constraint is tied to these exact token IDs and byte realizations.
Constraint
Constraint is the normal compiled type and can be reused across requests.
Constraint.from_json_schema(schema, vocab)
Constraint.from_ebnf(ebnf_source, vocab)
Constraint.from_lark(lark_source, vocab)
Constraint.from_glrm_grammar(
glrm_source,
vocab,
subgrammars=None,
bindings=None,
)
subgrammars binds compiled children to extern grammar names. bindings maps an extern token name to one or more exact token IDs.
payload = glrmask.Constraint.from_json_schema(payload_schema, vocab)
document = glrmask.Constraint.from_glrm_grammar(
'''
glrm 1;
start document;
extern token CONTROL;
extern grammar payload;
nt document = CONTROL "{" payload "}";
''',
vocab,
subgrammars={"payload": payload},
bindings={"CONTROL": [32001, 32002]},
)
constraint.start()
Call constraint.start() to create a new ConstraintState for a generation run.
state = constraint.start()
constraint.save() / Constraint.load(data, vocab)
Serialize and reload a compiled artifact.
blob = constraint.save()
constraint = glrmask.Constraint.load(blob, vocab)
Load only with the exact vocabulary the artifact was compiled against.
constraint.mask_len()
Returns the number of packed u32 words required by fill_mask().
ConstraintState
ConstraintState is mutable per-sequence state.
state.mask(size=None)
Returns a NumPy Boolean array indexed by model token ID without advancing the state.
mask = state.mask(size=llm.n_vocab())
If size is omitted, the wrapper uses the token extent represented by the constraint.
state.fill_mask(bitmask)
Writes the allowed-token mask into a contiguous NumPy int32 array interpreted as packed u32 words.
state.commit_token(token_id)
Commits a model token and advances the state.
state.commit_bytes(data)
Advances the state by raw bytes. This is useful for integrations and tests.
state.forced()
Returns a forced token sequence when one can be determined.
state.is_accepting()
Returns whether the current prefix may end. An accepting prefix may still allow more tokens.
state.is_rejected()
Returns whether no valid parser/tokenizer state remains.
End tokens
End tokens are handled by the decoder. If the constraint is accepting, generation may stop without committing another token.
mask = state.mask(llm.n_vocab())
if state.is_accepting():
mask[end_token_ids] = True
token = sample_with_mask(logits, mask)
if token in end_tokens:
break
state.commit_token(token)
Do not invent empty bytes for EOS.
DynamicConstraint
DynamicConstraint uses the same source formats, but compiles faster and generates masks more slowly.
DynamicConstraint.from_json_schema(schema, vocab)
DynamicConstraint.from_ebnf(ebnf_source, vocab)
DynamicConstraint.from_lark(lark_source, vocab)
DynamicConstraint.from_glrm_grammar(
glrm_source,
vocab,
subgrammars=None,
bindings=None,
)
from_glrm_grammar(...) accepts the same subgrammars and bindings arguments.
constraint.start()
start() returns a DynamicConstraintState with the same methods as ConstraintState:
state.mask(size=None)
state.fill_mask(bitmask)
state.commit_token(token_id)
state.commit_bytes(data)
state.forced()
state.is_accepting()
state.is_rejected()
constraint.save() / DynamicConstraint.load(data, vocab)
Serialize and reload a dynamic artifact.
constraint.mask_len()
Returns the packed-mask length in u32 words.