Skip to content

This site describes solve-engine as it is on main: 2.43.0, which npm does not have yet. npm installs 2.40.0, so a page may show an answer that version does not give yet.

Testing

Terminal window
npm run test:ci

The suite is 37,388 tests across 890 spec files. The figures are imported from docs/src/data/testStats.json, which is derived from the last full run and checked in continuous integration, so they cannot be older than the suite. Tests run against source rather than the built package, so a failure points at a line you can edit.

jest.config.js is the fast configuration, the one npm test, npm run test:ci and therefore npm run verify use. It leaves out four suites that are slow or memory-hungry: heavy/MemoryLeak, LexerFuzz, LexerVocabularyFuzz and LongDocumentRobustness.

jest.full.config.cjs puts them back. npm run test:full runs it single-threaded with a raised heap and writes the report the test-count stats are read from. It is what continuous integration runs, and what npm run verify:ci runs. A change that touches the lexer vocabulary, a unit, or document-scale behaviour should be checked with it before it is pushed: those are the suites the fast run cannot see, and the ones that have gone red in continuous integration after a green verify.

packages/engine/jest.config.cjs is the third, and it is deliberately independent: it is what the engine would use as a repository of its own, so it cannot import the root config. cd packages/engine && npm test uses it, and so do most editor test runners.

Independence means the two are kept in step by hand, and hand-syncing has failed three times: on the editor-library mocks the language tests need, on the temporal-polyfill transform without which three Temporal suites cannot load, and on ignoring a git worktree checked out under .claude/. Every one of them stayed green in continuous integration, which runs the root config, and every one of them looked like fewer tests rather than a failure.

Two guards now cover that. npm run lint:jest-configs compares the two configs on everything that decides whether a test can run, allowing only for the different root each addresses files from. And npm run lint:stats refuses a report in which any suite ran no tests at all, because a suite that fails to load contributes zero failed tests: the run stays green and the total quietly shrinks. No suite here legitimately holds no tests, so that shape is always a broken run rather than a smaller one.

jest.coverage.config.cjs runs the same full suite with coverage measured, and fails when it falls under the floor the config declares: a few points beneath the value measured when the floor was set, so ordinary churn never trips it and a real drop does. It runs daily in its own workflow rather than on every pull request, because the per-file coverage conversion makes the run several times slower than the suite itself. npm run test:coverage is the same check locally; run it for a change that removes tests.

Coverage is measured in shards. scripts/run-coverage.mjs starts three Jest processes side by side, each running a third of the spec files with --shard, and each writes its raw counts without checking any threshold, because a third of the specs covers only part of the code. The script then adds the three sets of counts together and checks the floor once, on the whole suite’s figures, so the number is the same one a single run would give. SOLVE_COVERAGE_SHARDS changes the count (a machine with less memory can use one or two). A single run had grown to 80 to 90 minutes, against the workflow’s 90-minute limit, and about half the daily runs were cancelled before the floor was measured (#890).

No pull request shows a scheduled run’s result, so the coverage workflow and the nightly fuzz soak each end in a report job: when the run does not succeed, including when it is cancelled at its time limit, the job reopens (or opens) one issue for that workflow and comments with the run’s link, and for the fuzz soak the start seed and the seed of each finding the log names. A manual run is watched by whoever started it, so it does not report.

A test that makes a real network call proves the integration works rather than only its mocked shape, and it also fails whenever the service is slow or rate-limits the caller, which is no fault of the change under test. So those calls run only with SOLVE_LIVE_NETWORK=1 set: daily in the live-network workflow, and locally through npm run test:live. Without the flag each live case reports itself skipped and asserts what it can offline, and the client’s error branches (a failed status, an empty result, a request that stalls until its timeout) are driven through a stubbed fetch in the gates. A new data source that calls a service follows the same pattern.

Tests live beside the area they cover: lexer, parser, virtual machine, each package, and a set of integration suites that drive a real engine end to end.

There is also a directory of regression tests named after the specific defect each one prevents, which is the most useful documentation of past mistakes the repository has.

A whole-document feature, one that reads or re-runs other lines, is tested through every entry point in __tests__/integration/CrossPathDocumentFeatures.spec.ts, because each entry point behaves differently and a form that works through one and misbehaves through another is exactly the drift a per-feature test misses.

A test of a whole note shows the answer a reader sees, and that is the test a change is judged by. It has two limits: it can only pass the helpers inside the engine the arguments a note happens to produce, and when it fails it points at the whole pipeline rather than the part that broke. So a change also tests its parts directly. Every helper, class or module function it adds or alters (a budget check, a pool of reusable values, a formatter, a normaliser rule, a parselet) gets unit tests of its own, calling it with ordinary, boundary and hostile arguments. A fix whose cause was one function has a test of that function, as well as of the note that exposed it.

An ordinary test checks that a feature does what its author meant, for the input its author had in mind. An adversarial test does the opposite: it tries to break the feature, with the input nobody had in mind. Most of the defects found in this engine’s recent reviews were of that kind. A check that compared exact values as doubles, a trace that missed a section, a sweep that could exhaust memory: each passed its own tests and failed the first time something unexpected reached it. So every bug fix and every feature ships adversarial tests beside its ordinary ones, in the same change.

They attack from three sides:

  • Security: what a hostile document could carry. A word such as constructor or __proto__ wherever a word the reader typed is looked up, since a plain object inherits properties with those names; input sized to exhaust time or memory; characters that look like one thing and are another, such as a zero-width space or a digit from another script; and text shaped like markup or an injection, which must be read as text.
  • Realistic breakage: what real readers and hosts do. A typo, a unit that does not fit, a value from another line rather than a literal, the feature meeting the others (a check over it, a what-if through it, a trace of it), the same document through the other entry point, an edit, a snapshot round trip.
  • Edge cases: the boundaries. Zero and negative zero, 2^53 and the 34-digit decimal limit, the largest doubles, the quotients with no finite answer, empty and blank lines, a stray carriage return or a trailing newline, a change of clocks, a leap day, a month end.

The engine’s promise is that an answer is right or an honest refusal that names the problem. So, whatever the input, these are failures: a confident wrong number; a raw JavaScript error such as a TypeError; an internal name in what the reader sees, such as [object Object]; an unexplained NaN; a hang; the two document passes disagreeing; and a change to Object.prototype, which would affect every object in the host’s process.

The shared kit is packages/engine/tools/adversarial.ts. It holds the corpora of hostile and edge inputs, fill() to run one form over a whole corpus, and the checks that say what honest means:

import { NUMERIC_EDGES, PROTOTYPE_WORDS, expectHonestLine, expectPrototypeUntouched, fill } from "@tools/adversarial";
test.each(fill("round(X, 2)", NUMERIC_EDGES))("rounding stays honest: %s", (line) => {
expectHonestLine(line);
});
test.each(fill("5 as X", PROTOTYPE_WORDS))("an inherited word is an unknown word: %s", (line) => {
expectPrototypeUntouched(() => expectHonestLine(line));
});

__tests__/hardening/AdversarialFeatureSweep.spec.ts runs every form the engine reads over those corpora, single lines and cross-line documents alike, and a new form gets a template there in the same change. The honesty checks cannot know the right answer to an arbitrary input, so they are paired with ordinary assertions wherever the right answer is known.

A known open bug found this way is filed as an issue and pinned as a test.failing with exactly one assertion, named after the issue. When the fix lands the test starts passing, which fails it, and it moves into the ordinary set. It is never deleted or weakened to make a run green.

Every example in the syntax reference is executed when the suite runs, and compared against the documented result. There are 3,157 of them.

The format is a fenced block tagged solve, with the expected result after a comment marker:

50% of 200 // 100

Because the marker is the language’s own comment syntax, every documented line is valid input a reader can paste unchanged. A blank line starts a fresh engine, so examples cannot leak variables into one another.

If you change behaviour, the documentation fails the build rather than going quietly out of date.

Excluded from every test run because they are slow and timing-sensitive. They have their own workflow, which compares a pull request against its merge base on the same runner; see performance.

The comparison job runs the merge base, then this branch, one after the other in a single job, and divides. That order is worth knowing about when you read the result: the second pass runs on a machine the first has been working for several minutes, and a branch that adds benchmark cases makes its own pass longer still.

A ratio near one means nothing changed. A large ratio on a case whose timed function constructs an engine (new ExpressionEngine(...) inside the loop) is usually the machine rather than the code, because those cases are dominated by construction and construction is where the drift lands.

So one pass over the limit is not yet a failure. When the first comparison puts a case or a suite over its limit, the job measures again: it re-runs only the suites that hold an over-limit case, on the merge base (kept in a worktree of its own, with its own install) and on this branch, for three rounds, switching which side goes first each round so neither always gets the warmer machine. An item fails only if the median of its re-measured ratios is still over the limit. The job summary and the pull-request comment show both measurements side by side, so a confirmed regression reads confirmed regression and a cleared one reads “cleared: noise on the first measurement”. The limits and the reference are the same for both passes; a real regression still fails, a few minutes later than before.

A suite file is the smallest unit the harness can run, so the re-measure runs whole suites, and judges only the items that were over.

When a result matters, measure it directly rather than trusting one run: build both revisions, load the two bundles into one process, and interleave the passes so drift hits each equally. That method reversed the verdict on the pipeline wave-2 branch, which the sequential comparison called 1.4x to 1.8x slower on several cases and which was, measured interleaved, faster on the same ones.

How noisy is it, exactly? A branch changing no engine code at all was put through the comparison to find out. Per case the spread reached 1.95x and 0.51x on identical code. Across a whole suite, the geometric mean stayed between 0.943x and 1.065x.

That is why the gate is built the way it is: a single case has to be half again as slow before the build stops, while a suite only has to average a quarter slower. A uniform slowdown that leaves every individual case under its own limit is the shape worth catching, and the suite mean is what catches it. If you are tempted to tighten the per-case limit, measure first, on a branch that changes nothing.

The benchmark specs also carry plain toBeLessThan bounds, and those are a different instrument from the comparison. The comparison is the regression gate; the bounds are a smoke bound underneath it, there to catch a collapse rather than a few per cent.

So each is set to at least four times the slowest median the shared runner has been measured delivering, rounded up to a readable number, and each test’s name states its own number so the two cannot drift apart.

A bound set close to the measured figure does not watch the engine, it watches the runner. One was: the pipeline suite’s variable chain constructs an engine and parses five lines, and it was bounded at 2ms on a runner delivering medians between 1.61 and 2.10. It failed on the merge base pass, which left the comparison with nothing to compare, so pull requests that had changed nothing near it went red. Raising it lost nothing, because the comparison was already watching that case more closely than the bound ever could.