Skip to content

This site describes solve-engine as it is on main: 2.43.0, which npm does not have yet. npm installs 2.40.0, so a page may show an answer that version does not give yet.

NormalizerRule

Defined in: packages/engine/src/normalizer/NormalizerRule.ts:184

A pluggable normalization rule registered with the TokenNormalizer.

Each rule has a name, priority, and match function. The match function receives the current token stream and a position, and returns a NormalizerMatch on success or null on failure.

Higher priority rules are tried first at each position. This allows long phrases (priority 100, e.g. “to the power of”) to match before shorter fragments (priority 80, e.g. “power of”).

  • Must be pure (no side effects, no mutation of input tokens)
  • Must return null for any position that doesn’t match
  • Consumed tokens must be consecutive starting at pos
  • Replacement tokens must be valid for downstream parsing
// A phrase fusion rule that converts "to the power of" into CARET
const phraseRule: NormalizerRule = {
name: 'phrase:to the power of',
priority: 100,
match: (tokens, pos) => {
if (pos + 4 > tokens.length) return null;
const phrase = tokens.slice(pos, pos + 5)
.map(t => t.value.toLowerCase()).join(' ');
if (phrase === 'to the power of') {
return {
consumed: 5,
replacement: [createFusedToken('CARET', 'to the power of', tokens.slice(pos, pos + 5))],
};
}
return null;
},
};
readonly name: string;

Defined in: packages/engine/src/normalizer/NormalizerRule.ts:189

Human-readable name for debugging and diagnostic display. Convention: "category:description", e.g. "phrase:to the power of".


readonly priority: number;

Defined in: packages/engine/src/normalizer/NormalizerRule.ts:199

Priority for ordering rules. Higher values are tried first. Recommended ranges:

  • 100: Long multi-word phrase fusion (e.g., “to the power of”)
  • 80: Short phrase fusion (e.g., “power of”, “times by”)
  • 50: Implicit operator insertion (e.g., implicit multiply)
  • 20: Domain-specific transformations

readonly optional shape?: readonly RuleSlot[];

Defined in: packages/engine/src/normalizer/NormalizerRule.ts:250

The rule’s leading shape: what the tokens from the match position onward may be, one RuleSlot per position.

This generalises startTokenTypes, which constrains only the first token. Constraining the first token alone is not enough to separate the rules that matter: every rule firing on a bare NUMBER declares the same start type, so they all remain candidates at every number in the document. What distinguishes them is the token after it, NUMBER COLON being a clock time and NUMBER SLASH a network address, and that fact is only usable by an index if the rule states it rather than hiding it inside match().

Depth is the rule’s choice, not the interface’s. The normalizer builds one lookup plane per declared slot and intersects them, so a rule that declares three positions is filtered on three. It may also index fewer planes than were declared, which stays correct for the reason given on RuleSlot: a shallower filter admits more candidates, and each surviving rule still runs its own match().

Prefer this to startTokenTypes in new rules. When both are given, this wins; startTokenTypes: ["IDENT"] means exactly shape: [{ types: ["IDENT"] }].

// 9:00am, 16:00, a clock time is a number followed by a colon
shape: [{ types: ["NUMBER"] }, { types: ["COLON"] }]
// sha256("hi"), a known word followed by an opening parenthesis
shape: [{ types: ["IDENT"], values: HASH_NAMES }, { types: ["LPAREN"] }]

readonly optional startTokenTypes?: readonly string[];

Defined in: packages/engine/src/normalizer/NormalizerRule.ts:216

The token types this rule can match at position pos, as a performance hint. When present, the normalizer only tries the rule at a position whose token type is in this list, and skips it everywhere else.

This MUST be exact: a rule whose match() begins if (tokens[pos].type !== "IDENT") return null can only ever match an IDENT, so declaring ["IDENT"] changes nothing (the rule would have returned null at every other position anyway) while letting the normalizer skip it for the many number/operator tokens in a document. Omit it, or list every type the rule can match, when the first token is not fixed; an omitted hint means “always try”, the original behaviour. Declaring a type the rule cannot actually start on is harmless; declaring too FEW (missing a type it can match) makes the rule silently unreachable there, which is a bug.


readonly optional unshapedReason?: string;

Defined in: packages/engine/src/normalizer/NormalizerRule.ts:272

Why this rule cannot declare a shape, for the few that genuinely cannot.

A rule with neither a shape nor a startTokenTypes hint is tried at every position of every line, so it raises the cost of the whole document rather than only its own feature. Registering one logs a warning naming the rule, which is how a package author finds out before their users do.

Some rules really cannot be described by a leading shape: an unbounded forward scan, a greedy match against a table the host mutates at runtime. Setting this states that case, silences the warning, and leaves the reason in the code where the next person will read it. It is deliberately a sentence and not a boolean, because “why” is the part worth keeping.

unshapedReason: "Scans forward an unbounded number of NUMBER UNIT pairs, so no fixed leading shape describes it.",
match(
tokens,
pos,
environment?
):
| NormalizerMatch
| null;

Defined in: packages/engine/src/normalizer/NormalizerRule.ts:285

Attempt to match a pattern starting at position pos in the token stream.

ParameterTypeDescription
tokensToken[]The current token stream (may be partially normalized from prior passes)
posnumberThe current position to attempt matching from
environment?NormalizerEnvironmentWhat the rule may read about the engine running it; see NormalizerEnvironment. A rule that reads nothing of the engine’s ignores it.

| NormalizerMatch | null

A NormalizerMatch if the pattern is found, or null if no match