Skip to content

This site describes solve-engine as it is on main: 2.43.0, which npm does not have yet. npm installs 2.40.0, so a page may show an answer that version does not give yet.

ExpressionLexer

Defined in: packages/engine/src/lexer/ExpressionLexer.ts:587

Character-by-character tokenizer for expression text.

Scans a raw line/expression string into a stream of typed tokens (numbers, identifiers, operators, units, keywords, …), handling markdown-line classification (classifyLine), inline s`...` solve spans, and package-contributed vocabulary (registered via registerVocabulary/unregisterVocabulary, keywords operators, and units a package wants recognized as their own token types rather than falling through to generic identifiers).

Most consumers should use the higher-level Lexer wrapper, which adds streaming next()/peek() access over this class’s scan results.

new ExpressionLexer(localeCode?, _lookup?): ExpressionLexer;

Defined in: packages/engine/src/lexer/ExpressionLexer.ts:736

ParameterTypeDefault valueDescription
localeCodestring'en'Locale whose keyword table seeds the lexer.
_lookup?TokenLookupundefinedIgnored. The lexer built its keyword, unit and phrase tables from the locale and the registered packages and never read the lookup it was handed; the parameter stays so existing callers compile.

ExpressionLexer

_inlineSolveSpans: InlineSolveSpan[] = [];

Defined in: packages/engine/src/lexer/ExpressionLexer.ts:674

Inline solve spans collected during the most recent tokenization pass. Populated by tokenizeInto() and consumed by scanDocument().

iterator: Generator<Token, void, undefined>;

Defined in: packages/engine/src/lexer/ExpressionLexer.ts:1526

Iterate the tokens of the current input.

Runs tokenizeInto first, so the tokens are the same as before, but a fault part way through the line is thrown from the first next() rather than at the token it occurred on. A caller that wants the tokens read before a fault calls tokenizeInto with its own array.

Generator<Token, void, undefined>


classifyLine(lineText): LineClassification;

Defined in: packages/engine/src/lexer/ExpressionLexer.ts:2417

Classify a single line of markdown text.

ParameterType
lineTextstring

LineClassification


findInlineSolves(lineText): InlineSolveSpan[];

Defined in: packages/engine/src/lexer/ExpressionLexer.ts:2445

Find all inline solve markers in a line with precise coordinate mapping.

ParameterType
lineTextstring

InlineSolveSpan[]


getKeywords(): Record<string, string>;

Defined in: packages/engine/src/lexer/ExpressionLexer.ts:2438

Every keyword this lexer currently recognizes, locale keywords (pi, sqrt, convert, …) merged with any plugin-contributed ones from registerVocabulary() (e.g. a package’s custom keywords), mapped to the token type they lex to. A snapshot copy, not a live reference mutating the return value has no effect on the lexer.

Record<string, string>


registerVocabulary(plugin): void;

Defined in: packages/engine/src/lexer/ExpressionLexer.ts:767

Register a plugin to extend the lexer with custom tokens.

All registrations are additive, built-in patterns still work. Keywords, operators, and units from the plugin are merged with existing ones. Calling multiple times adds more entries.

Note: multi-word phrases are now handled by the TokenNormalizer (see IEnginePackage.normalizerRules), not the lexer.

Built-in tokens CANNOT be overridden. Throws a EngineError if the plugin attempts to register a keyword, operator, or unit that conflicts with a built-in one.

ParameterType
pluginLexerVocabulary

void


reset(input, from?): void;

Defined in: packages/engine/src/lexer/ExpressionLexer.ts:996

Point the lexer at a new input.

ParameterTypeDefault valueDescription
inputstringundefinedThe text to tokenize.
fromnumber0Where tokenizing starts. Offsets and columns stay those of input, so a caller that starts past a marker (a blockquote’s > , a list’s - ) still gets spans on the line as written, the way scanDocument does for a list item.

void


scanDocument(text): ScanLineResult[];

Defined in: packages/engine/src/lexer/ExpressionLexer.ts:1032

Scan a full document text in a single pass, classifying each line and tokenizing non-skipped lines.

Replaces the separate classifyLine() + findInlineSolves() + per-line reset() + tokenizeAll() pattern with a single character-by-character walk through the entire document. Key benefits:

  • Single reset(): this.pos, this.len, this.line, and this.lineStartPos are set once for the whole document, not per-line.
  • Single classification: classifyLine() runs once per line inline; skipped lines are jumped over without tokenization.
  • Shared tokenization: Non-skipped lines are tokenized using the existing state machine, yielding Token[] without per-line reset().
  • Inline solve detection: findInlineSolves() is called only for lines that classifyLine() marks as having inline solves.

Tokenization is scoped to each line by temporarily restricting this.len to the line end position, so tokenizeInto() naturally stops at the line boundary. After tokenization, this.len is restored and this.pos advances past the newline.

ParameterTypeDescription
textstringThe full document text (with newlines).

ScanLineResult[]

Array of ScanLineResult, one per line, in document order: at least one, since an empty document is one empty line, and a document ending in a line break has an empty line after it.


tokenizeAll(): Token[];

Defined in: packages/engine/src/lexer/ExpressionLexer.ts:1224

Tokenize an expression string into an array of Tokens.

Runs tokenizeInto into a fresh array.

Optimizations:

  • CHAR_CLASS jump table (Uint8Array) → switch on small integers
  • Direct character-code dispatch (c0 cached pattern)
  • Mathematical digit parsing (integer math, not slice+parseFloat)
  • Inline operator tokenizer with two-char peek-ahead
  • Whitespace eliminated in-lexer (never emitted)
  • 0-char and 1-char fast paths

Token[]


tokenizeInto(out): void;

Defined in: packages/engine/src/lexer/ExpressionLexer.ts:1252

Scan the current input and append every token to out.

This is the scanner itself: tokenizeAll and the iterator are thin wrappers over it. It pushes into an array the caller owns rather than yielding, because a generator paid a resume per token and Array.from a second pass on top, and because a caller that catches a tokeniser fault (highlighting a line with an unterminated quote) still holds the tokens read before it.

IMPORTANT: this.len is read ONCE on entry (const len = this.len). scanDocument() relies on this to scope a pass to a single line by restricting this.len to the line end before calling. Do not re-read this.len mid-loop without also updating scanDocument().

ParameterTypeDescription
outToken[]The array to append to. Left as it was if the input is empty.

void


unregisterVocabulary(plugin): void;

Defined in: packages/engine/src/lexer/ExpressionLexer.ts:935

Unregister a plugin, removing its custom tokens from the lexer.

This is the inverse of registerVocabulary(). All keywords, operators, and units registered by the plugin are removed. After unregistration, those tokens will revert to their default behavior (e.g., keywords become IDENT, operators become ERROR).

Calling unregisterVocabulary with a plugin that was never registered is safe, it simply has no effect.

ParameterType
pluginLexerVocabulary

void