Python API

Python API reference.

Vocab

Vocab.from_llama_cpp(llm)

Build the byte vocabulary used by a llama-cpp-python model.

vocab = glrmask.Vocab.from_llama_cpp(llm)
end_token_ids = vocab.llama_cpp_end_token_ids

EOG, control, unused, and empty-piece tokens are omitted from the byte vocabulary. Their IDs remain available in llama_cpp_end_token_ids for the decoder.

Vocab.from_id_to_bytes(id_to_bytes)

Build from an explicit {token_id: bytes} mapping.

Vocab.from_dict(token_to_id)

Build from a {bytes: token_id} mapping.

A compiled constraint is tied to these exact token IDs and byte realizations.

Constraint

Constraint is the normal compiled type and can be reused across requests.

Constraint.from_json_schema(schema, vocab)
Constraint.from_ebnf(ebnf_source, vocab)
Constraint.from_lark(lark_source, vocab)
Constraint.from_glrm_grammar(
    glrm_source,
    vocab,
    subgrammars=None,
    bindings=None,
)

subgrammars binds compiled children to extern grammar names. bindings maps an extern token name to one or more exact token IDs.

payload = glrmask.Constraint.from_json_schema(payload_schema, vocab)

document = glrmask.Constraint.from_glrm_grammar(
    '''
    glrm 1;
    start document;
    extern token CONTROL;
    extern grammar payload;
    nt document = CONTROL "{" payload "}";
    ''',
    vocab,
    subgrammars={"payload": payload},
    bindings={"CONTROL": [32001, 32002]},
)

constraint.start()

Call constraint.start() to create a new ConstraintState for a generation run.

state = constraint.start()

constraint.save() / Constraint.load(data, vocab)

Serialize and reload a compiled artifact.

blob = constraint.save()
constraint = glrmask.Constraint.load(blob, vocab)

Load only with the exact vocabulary the artifact was compiled against.

constraint.mask_len()

Returns the number of packed u32 words required by fill_mask().

ConstraintState

ConstraintState is mutable per-sequence state.

state.mask(size=None)

Returns a NumPy Boolean array indexed by model token ID without advancing the state.

mask = state.mask(size=llm.n_vocab())

If size is omitted, the wrapper uses the token extent represented by the constraint.

state.fill_mask(bitmask)

Writes the allowed-token mask into a contiguous NumPy int32 array interpreted as packed u32 words.

state.commit_token(token_id)

Commits a model token and advances the state.

state.commit_bytes(data)

Advances the state by raw bytes. This is useful for integrations and tests.

state.forced()

Returns a forced token sequence when one can be determined.

state.is_accepting()

Returns whether the current prefix may end. An accepting prefix may still allow more tokens.

state.is_rejected()

Returns whether no valid parser/tokenizer state remains.

End tokens

End tokens are handled by the decoder. If the constraint is accepting, generation may stop without committing another token.

mask = state.mask(llm.n_vocab())
if state.is_accepting():
    mask[end_token_ids] = True

token = sample_with_mask(logits, mask)
if token in end_tokens:
    break
state.commit_token(token)

Do not invent empty bytes for EOS.

DynamicConstraint

DynamicConstraint uses the same source formats, but compiles faster and generates masks more slowly.

DynamicConstraint.from_json_schema(schema, vocab)
DynamicConstraint.from_ebnf(ebnf_source, vocab)
DynamicConstraint.from_lark(lark_source, vocab)
DynamicConstraint.from_glrm_grammar(
    glrm_source,
    vocab,
    subgrammars=None,
    bindings=None,
)

from_glrm_grammar(...) accepts the same subgrammars and bindings arguments.

constraint.start()

start() returns a DynamicConstraintState with the same methods as ConstraintState:

state.mask(size=None)
state.fill_mask(bitmask)
state.commit_token(token_id)
state.commit_bytes(data)
state.forced()
state.is_accepting()
state.is_rejected()

constraint.save() / DynamicConstraint.load(data, vocab)

Serialize and reload a dynamic artifact.

constraint.mask_len()

Returns the packed-mask length in u32 words.