Search papers, labs, and topics across Lattice.
This paper introduces a method for parallel lexing that generalizes the concept of certified split points from single bytes to bounded windows, allowing for the recovery of token boundaries in complex input scenarios. By employing a conservative model of a maximal-munch scanner, the authors prove the soundness of their approach and demonstrate its effectiveness across a diverse set of token configurations. The results show that a significant majority of previously uncertified token sets can now yield certified windows, enhancing the robustness of parallel lexers in handling various input forms.
A staggering 91 out of 95 previously uncertified token sets now gain a witnessed window, revolutionizing how we handle complex input in parallel lexing.
A certified split point lets a parallel lexer cut unlexed input at a single byte with the serial token stream provably preserved, but several conventional token sets in the predecessor's controlled study certify no byte once string, comment, or whitespace-run forms are included (arXiv:2608.03473). We generalize from a byte to a bounded window: a byte string after which the position where the current token began is known, regardless of surrounding context. We certify the directly usable form of that recovery: the token covering the window's final byte begins at the reported origin. The certificate is conditional on occurrence and may be vacuous; every applicability figure counts only windows carrying an asserted completely tokenizable occurrence witness. We give a conservative model of a maximal-munch scanner's possible histories across a window, prove it sound, and decide reachability in that model exactly by exhausting a finite quotient of its reachable configurations, so every answer of the unbudgeted procedure is either a certified window with its origin or a proof that the model admits none. Within the stated flat, non-nullable, completely-tokenizable scope, model-positive answers are semantic certificates; negatives are relative to the conservative model, which deliberately refuses some windows a greedy scanner would allow. In a sample of 400 random token sets, 91 of the 95 non-nullable sets certifying no byte gain a witnessed window, with zero inconclusive searches, and every exact-empty row of the predecessor's study gains a witnessed window of two to four bytes. Rewind-stress rows exercised 1,079,392 executions that scanned through the window and contained at least one rewind, with zero disagreements against the shipped scanner. The analysis runs once after automaton construction, using only the compiled tables and no input.