llguidance HATES This One Weird Grammar
A tiny grammar that exposes a strange edge case in llguidance's greedy lexer.
Consider this grammar:
%llguidance {"no_forcing": true}
start: STEM SUFFIX
STEM: "list" | "listen"
SUFFIX: "ed"
It accepts two strings:
list + ed = listed
listen + ed = listened
Lark accepts both. llguidance accepts listened.
But listed?
>>> b"listed" in next_token_mask
False
>>> consume_token(b"listed")
False
This is not because the model token crosses a grammar-terminal boundary. llguidance can normally handle that just fine.
The problem is greediness.
What happens after list?
While reading STEM, the lexer sees:
list accepting
liste live, non-accepting
listen accepting
After list, it could stop. But it could also keep going toward the longer match listen.
For the input listened, continuing is correct:
listen | ed
For listed, it is not:
list | ed
The annoying part is that the first byte after list is e. That byte is ambiguous. It can either continue STEM toward listen, or begin SUFFIX = "ed".
So llguidance continues the current lexeme:
list accepting
liste still possible
listed dead
The d finally proves that liste... was the wrong path. But by then the useful boundary after list is two bytes behind.
An ordinary maximal-munch lexer remembers the last accepting position while it scans. When the longer attempt dies, it can fall back to list, emit STEM, and continue with ed.
llguidance’s recognizer is shaped differently. It is being advanced incrementally while llguidance walks the model vocabulary trie. While the current lexeme can still continue, it keeps going. When continuation finally fails, its current lexer state is liste, which is not accepting. The earlier accepting list boundary is gone.
That is the whole bug.
The easy fix
Move the choice out of the lexer and into the parser:
%llguidance {"no_forcing": true}
start: stem SUFFIX
stem: "list" | "listen"
SUFFIX: "ed"
Lowercase stem is now a parser rule, not one greedy terminal.
And:
>>> b"listed" in next_token_mask
True
>>> consume_token(b"listed")
True
BEHOLD.
Why this almost never matters for JSON
This sounds worse than it is.
Most constrained generation is JSON, and JSON is unusually friendly to this style of lexer. String boundaries are explicit because of the closing quote. After true, false, and null, the next legal byte cannot also continue the same literal. Structural punctuation gives the lexer obvious places to stop.
Numbers contain paths like:
1 accepting
1e live, non-accepting
1e2 accepting
but the surrounding JSON grammar normally prevents the dangerous handoff. The byte that continues the number is not simultaneously a legal first byte of whatever comes after that number.
The pathological shape is more specific:
A lexeme has already matched, but the next byte can both continue that lexeme and begin the next legal lexeme.
That is exactly what happens here. After list, e can continue STEM and it can begin SUFFIX.
If a custom grammar hits this, the simplest repair is usually to expose that boundary choice to the parser rather than hiding it inside one terminal.
This is not fundamental to llguidance
There is another possible fix: remember earlier accepting lexer positions. If the attempted longer match later dies, return to the most recent accepting boundary and replay the bytes after it in the new parser context.
I implemented an opt-in version of that in redesign/greedy-lexeme-fallback-integrated-v3. The normal fast path stays intact; the fallback machinery matters only when one of these unresolved greedy boundaries is actually present.
But as a grammar author, you probably do not need any of that.
You just need to know that this innocent-looking grammar:
STEM: "list" | "listen"
SUFFIX: "ed"
contains a lexical boundary that llguidance cannot see soon enough.
And, for once, listed is more troublesome than listened.