llguidance HATES This One Weird Grammar (Handoff rule)
A tiny lexer failure and the simple boundary rule that predicts it.
Consider this grammar:
%llguidance {"no_forcing": true}
start: STEM SUFFIX
STEM: "list" | "listen"
SUFFIX: "ed"
It accepts listed and listened:
list | ed
listen | ed
Lark has no trouble with either. llguidance accepts listened. But it rejects listed as a single model token:
>>> b"listed" in next_token_mask
False
>>> consume_token(b"listed")
False
The failure becomes fairly obvious if you watch the recognizer one byte at a time.
After list, STEM has already matched. But the lexer cannot know that list is the final greedy match, because listen is also possible.
Then it sees e:
list accepting
liste live, non-accepting
The e keeps STEM alive, so llguidance continues it. The following d finally kills that path. The trouble is that the correct interpretation was not liste... at all. It was:
list | ed
The evidence arrives one byte too late.
There is a useful way to recognize grammars that can cause this.
The dangerous handoff
At a valid lexeme boundary, ask what the very next byte can do.
If that byte can continue the current lexeme and also begin a legal following lexeme, the boundary is hidden from a recognizer that only commits when continuation fails.
Our grammar is almost a deliberately minimal example:
current lexeme: STEM = "list" | "listen"
next lexeme: SUFFIX = "ed"
after "list":
e can continue STEM -> "liste..."
e can begin SUFFIX -> "ed"
That overlap does not prove every such grammar will fail. Parser context, terminal priorities, and the other active terminals can matter. But it tells you exactly where the cheap byte-at-a-time handoff has become suspicious.
Contrast an identifier followed by (:
IDENT = [A-Za-z_][A-Za-z0-9_]*
The ( cannot continue IDENT. As soon as it arrives, the lexer knows the identifier is over. The same byte can safely be reconsidered as the beginning of whatever comes next.
Our e is different. It does not kill the old lexeme, so there is no handoff yet.
The same terminal can therefore be safe in one grammar and awkward in another. Suppose we have a dotted name:
STEM = [a-z]+(\.[a-z]+)*
Before a colon it is comfortable:
STEM ":" VALUE
The colon cannot continue STEM, so it reveals the boundary immediately.
Now put the same terminal before an extension that begins with a dot:
STEM EXT
EXT = ".json" | ".txt"
The regex for STEM did not change. But . now has the same two jobs as the e in listed: it can continue the current dotted stem, or it can begin the following extension. Whether the boundary is easy is therefore a property of the terminal in parser context, not merely a property of its regex.
Fixing an affected grammar
The easiest fix is usually to make the parser own the ambiguous choice:
%llguidance {"no_forcing": true}
start: stem SUFFIX
stem: "list" | "listen"
SUFFIX: "ed"
stem is now a parser rule rather than one terminal. With that change:
>>> b"listed" in next_token_mask
True
>>> consume_token(b"listed")
True
The general lesson is not “split every terminal into tiny pieces.” That would wake the parser unnecessarily often and can make a grammar much slower. The useful place to split is a boundary where two interpretations genuinely need parser-level context.
For the dotted example, that might mean making the dots visible to the parser instead of burying all of them inside STEM. The exact rewrite depends on the language, but the principle is the same as lowercase stem: move the decision to the layer that has enough context to make it.
Why JSON is mostly safe
This is also why the issue is easy to miss in normal constrained generation. JSON has unusually explicit boundaries.
Strings end at ". That quote cannot continue the contents of the string. true, false, and null are followed by punctuation, whitespace, or end of input, none of which continues the literal. Arrays and objects are full of commas, colons, and brackets that make handoffs obvious.
Numbers look more worrying:
1 accepting
1e live, non-accepting
1e2 accepting
So JSON numbers do not satisfy the stronger rule that “once a terminal is accepting, every live continuation is also accepting.” But they usually satisfy the rule that actually matters here: the byte continuing the number is not simultaneously the first byte of a legal lexeme after that number in the current parser context.
The context qualification is important. Lexer safety is not purely a property of one regex. The same terminal can be harmless before one neighbour and awkward before another.
There is a stronger condition that is easier to check: once the combined lexer is accepting, require every continuation that keeps it alive to remain accepting. Our lexer fails because list is accepting but liste is not. Ordinary identifiers usually satisfy it. JSON numbers do not, which is why that condition is too conservative to use as a definition of safety. The handoff test is more useful because it asks whether the temporary non-accepting continuation conflicts with what the parser could legally do next.
Or fix the recognizer
The grammar does not fundamentally have to be rewritten. A recognizer can remember that list was accepting while it continues trying listen. If the longer path later dies, it can return to the saved boundary and replay the later bytes as the next lexeme.
I implemented an opt-in version of that in redesign/greedy-lexeme-fallback-integrated-v3.
For grammar authors, though, the compact diagnostic is probably more useful:
When a lexeme has already matched, can the next byte both keep it alive and begin what is allowed after it?
For STEM = "list" | "listen" followed by SUFFIX = "ed", the answer is e.
And that one byte is enough to make listed the weird case.