Skip to content

This site describes solve-engine as it is on main: 2.43.0, which npm does not have yet. npm installs 2.40.0, so a page may show an answer that version does not give yet.

Security

A calculator’s whole job is to evaluate whatever somebody typed, so hostile or malformed input is the normal case here rather than an edge case. In an editor integration it arrives one keystroke at a time from a person who is still mid-thought, and the process the engine is running in is the editor. Losing that process is not an acceptable outcome for a mistyped sum.

Two things follow. The engine has no capability it does not need, and it refuses rather than fails when an expression asks for too much.

Every claim below is checkable against the source, and where one command settles it, the command is here too.

The block below is a real engine running in your browser, the same component the front page uses. Every line in it is something the engine declines, and the answer column is the refusal it gives back.

It is editable. Change a number, paste in the worst thing you can think of, hold a key down. The refusals follow what you type and the page carries on, which is the whole demonstration.

sum(x, 1:100000000)
5 kg in m
sqrt()
2 +* 3
f(n) = f(n-1)
f(5)
1 km > 500 m

Line by line, and none of these is a special case written for this page:

  • sum(x, 1:100000000) asks for a hundred million values from twenty characters. A range costs nothing until something expands it, so this is the cheapest way there is to ask for an enormous allocation. It comes back as COLLECTION_TOO_LARGE, naming the setting that refused it: This collection has 100000000 elements, past the limit of 100000 (see the engine’s vm.maxCollectionSize setting).
  • 5 kg in m is not a conversion. Mass and length do not measure the same thing, and the engine says so in words, a mass cannot be converted to a length, rather than handing back 5 kg as though the request had been honoured.
  • sqrt() gets sqrt() takes 1 argument, but was given none, from a check against a table of builtin arities rather than from whatever happens inside sqrt when its argument is missing.
  • 2 +* 3 is a parse error, reported as one and confined to its line.
  • f(n) = f(n-1) defines a function with no base case, and f(5) calls it. The recursion guard stops it at a nesting depth of 50 rather than at whatever depth the JavaScript call stack happens to give out, which is not a failure a host can catch.
  • 1 km > 500 m is true. It is here because the interesting half of a safety story is the part where the engine still answers: units are unified for the comparison rather than the numbers 1 and 500 being compared.

Every one of those is recoverable. The engine instance is still usable, and the next line is evaluated normally, which is why the last line still answers after the six above it.

A word the engine looks up in a table, a converter (as hex), a unit, a timezone (in Tokyo), a function name, is matched only against entries that table actually holds. A name every JavaScript object carries on its prototype, constructor or __proto__, is unknown like any other unrecognised word: 5 as constructor is UNKNOWN_AS_CONVERTER, not the operand handed back, and map(constructor, [1,2,3]) is an unknown function rather than an internal fault. The tables are never written through such a name, so nothing an expression types reaches the prototype chain.

One honest note while you are experimenting, because it is the first thing you will notice. An unterminated " is something the lexer cannot get past at all, so it fails the scan of the whole document rather than of one line, and the notepad clears the column rather than leaving yesterday’s answers beside changed text. Close the quote and everything comes back. Blank is the correct thing to show there, for the same reason the engine reports a conversion it cannot do instead of returning the number unconverted: an error, never a guess.

There is no eval, no new Function, and no code generation anywhere in packages/engine/src. An expression is lexed, normalised, parsed and compiled to a fixed bytecode instruction set, then run on a virtual machine that can only do what its opcodes do. There is no path from an expression to arbitrary JavaScript, because there is nothing at the end of the pipeline that turns text into code.

Terminal window
grep -rn "eval(" packages/engine/src # nothing
grep -rn "new Function(" packages/engine/src # nothing

The instruction set is a closed list. A package can add opcodes, and those are ordinary TypeScript functions the host chose to register, not something an expression can reach on its own. The bytecode virtual machine covers the shape of it.

The engine reads no files, spawns no processes and opens no sockets. There is no import of node:fs, node:child_process, node:net, node:http or node:https anywhere in the source, which is one grep to confirm:

Terminal window
grep -rEn "from [\"'](node:)?(fs|child_process|net|http|https)[\"']" packages/engine/src

The solve command does read files, which is why it is a package of its own (packages/cli) rather than part of the engine: it depends on the engine, and the engine’s package, source and promise are unchanged by it.

The MCP server (packages/mcp) lets an AI tool call the engine, which means evaluating text a model wrote. Its defaults are chosen for that, and a tool call can change none of them: the network is off unless the person starting the server passes --network on; every call gets a fresh engine, cleared when the call ends, so no call can read a variable, a unit or a cached value another call left; and its engines are built without solve-global-variables, whose global :name is the one form that reaches a store shared across engines. An input the engine refuses as a whole, such as a document past its line limit, costs that call a coded refusal and leaves the next call unaffected. The server depends on the engine and nothing else: it answers the protocol’s messages itself rather than through an SDK, so it adds no tree of its own, and the engine’s single runtime dependency, below, is unchanged.

That is also why it runs unchanged in Node, in a browser and in a worker: there is no host facility it expects to find.

There are four fetch calls in the source, in two files, and both of those files are live-data features rather than engine machinery:

FileEndpointReached by
uom/CurrencyExchange.tsapi.frankfurter.deva currency conversion, such as 100 USD in GBP
uom/CurrencyExchange.tsapi.coingecko.coma crypto pair, such as 1 BTC in USD
packages/weather/OpenMeteoClient.tsgeocoding-api.open-meteo.coma place name in a weather expression
packages/weather/OpenMeteoClient.tsapi.open-meteo.comthe forecast itself

Be clear about what that means, because it is the one place where “no I/O” needs a qualification. Currency and weather are both in the default package set, so an engine constructed with no arguments will make those requests. Nothing is sent until an expression asks for a live figure: 2 + 2, 15% of 2400 and 100 cm + 2 m issue no requests at all. What leaves the process is what the expression contained, which is a currency code or a place name.

Stocks and knowledge are the other two live-data packages. Both take a fetching function supplied by the host rather than an API key, and neither is registered by default, so unconfigured they return an honest “not configured” value rather than a guess.

A host that wants no outbound traffic at all switches the network off:

import { createEngine } from "solve-engine";
const engine = createEngine({ config: { network: { enabled: false } } });

No async resolver runs, so no request is started. A currency conversion then answers No exchange rate for USD to GBP: live data is switched off for this engine (network.enabled is false), and a weather lookup says the same. The error names the setting rather than an outage, which is the behaviour the engine aims for everywhere: an error, never a guess. Rates the host primes by hand keep converting. The one thing the switch cannot undo is a request a host-supplied plugin function starts on its own before returning a promise; the built-in packages all fetch through async resolvers, which the switch stops before they run. See switching live data off.

Leaving the two packages out is the other way to the same place, and removes the code paths as well as closing them:

import { ExpressionEngine } from "solve-engine";
import {
BUILTIN_PACKAGES,
CURRENCY_PACKAGE,
WEATHER_PACKAGE,
} from "solve-engine/packages";
const offline = BUILTIN_PACKAGES.filter(
(pkg) => pkg !== CURRENCY_PACKAGE && pkg !== WEATHER_PACKAGE,
);
const engine = new ExpressionEngine({ packages: offline });

Two packages fewer, and no code path that can reach the network.

@tanstack/query-core, which caches and deduplicates async resolution. That is the whole list, and packages/engine/package.json is the whole audit.

Everything else is in this repository: the lexer, the normaliser, the parser, the virtual machine, the unit table, the currency and date handling, and the computer algebra. A supply chain you can read in an afternoon is a deliberate choice rather than an accident of scope.

Untrusted input must not be able to hang or kill the host process. Every limit below is checked, and exceeding one raises a recoverable error that names what was refused, what the ceiling was, and which setting to change.

SettingDefaultWhat it countsError
validation.maxExpressionLength2,000characters in one expression, before lexingEXPRESSION_TOO_LONG
validation.maxComplexity500tokens, plus function calls x 5, plus deepest parenthesis nesting x 10EXPRESSION_TOO_COMPLEX
validation.maxNestingDepth50recursive-descent depth in the parserNESTING_DEPTH_EXCEEDED
vm.maxInstructions50,000opcodes executed for one expressionINSTRUCTION_LIMIT_EXCEEDED
vm.maxStackDepth200slots on the value stackSTACK_LIMIT_EXCEEDED
vm.maxCollectionSize100,000elements in one expanded range or matrixCOLLECTION_TOO_LARGE
vm.maxAllocatedElements2,000,000elements one evaluation may materialise in totalALLOCATION_LIMIT_EXCEEDED
vm.maxFunctionCalls10,000user-defined function calls in one evaluation, however they nestFUNCTION_CALL_LIMIT_EXCEEDED
vm.maxLineRunsPerPass1,000,000line runs one pass over a document does across its lines: what-if and sweep re-runs, goal-seek probes, and span aggregates’ reads at sixteen to a runPASS_WORK_BUDGET_EXCEEDED
vm.maxRetainedElements10,000,000elements a document’s answers keep, counted each pass: a list or matrix its cells, text one to eight characters, any other answer oneDOCUMENT_ELEMENT_LIMIT_EXCEEDED
performance.maxDocumentLines100,000lines in one document, checked before it is scannedDOCUMENT_TOO_LARGE
date.maxOffsetYears / minOffsetYears100 / -100how far a workday offset may walk, in yearsDATE_OFFSET_LIMIT_EXCEEDED

The defaults live in constants/Configuration.ts, where each one is documented with the reasoning behind its value.

Every row but two counts work inside a single expression. The two that count a whole document are there because the per-line limits could only cap one line at a time.

vm.maxLineRunsPerPass counts the work that reaches across lines: twenty sweep lines, each inside its own cap, once made a single pass re-run 2,000,000 lines. It is one count per pass over the document, and the line whose work would cross it is refused while the lines above keep their answers.

vm.maxRetainedElements counts what the document keeps. Each line’s answer stays in memory for as long as the note is open, and vm.maxAllocatedElements only stops one line making more than two million elements: 500 lines of :m = map(10*x, 0:99999), 13 KB of text, kept 50,000,000 elements and 385 MB of heap between them. The line whose answer would take the note past the ceiling is refused, and a name it assigned is let go, so the value is not kept through a variable either. A line that waits on a live value keeps nothing until the value lands; its re-run is then charged against what the whole note keeps, and an answer no larger than the one it replaces is always kept.

Both counts start again on every pass, and both document passes charge alike. The incremental evaluator counts a line it does not re-run at what it recorded when it last ran, so a live editor refuses the same line a fresh pass does.

Two details the table cannot show. The complexity score reaches its ceiling before the parser’s nesting depth does at the default settings, so a deeply nested expression is usually refused by maxComplexity and the parser limit is the backstop for a host that raised it. And COLLECTION_TOO_LARGE arrives as an error value in the result rather than as a thrown error, which is the same information by a different route: a value that knows it is an error, and that stays an error through any operation it takes part in rather than degrading into a number.

Two further limits are not EngineConfig fields at all, and putting them in the table as though they were would be overclaiming. Nested user-function calls are capped at a depth of 50 (FUNCTION_RECURSION_LIMIT_EXCEEDED), which is a parameter of createVM rather than a config field, so an ExpressionEngine always uses the default. The normaliser refuses to emit more than 10,000 tokens from one line (NORMALIZED_TOKEN_LIMIT_EXCEEDED), and refuses a rule chain that is still changing the line after 100 passes (NORMALIZER_PASS_LIMIT_EXCEEDED); both are options on the normaliser itself.

Pass a config override to the constructor. It is merged field by field over the defaults, so a field you do not mention keeps its default:

import { ExpressionEngine } from "solve-engine";
const engine = new ExpressionEngine({
config: { vm: { maxCollectionSize: 1_000 } },
});

The config field is an EngineConfigOverride, a per-section deep partial: you name only the one field you are changing, and the other four in vm (and every other section) keep their defaults. No spread of DEFAULT_CONFIG is needed.

The ceiling you set is the one that appears in the refusal, so sum(x, 1:5000) on that engine answers This collection has 5000 elements, past the limit of 1000 (see the engine’s vm.maxCollectionSize setting).

Ordinary expressions are nowhere near any of this. The longest range in the test suite is a thousand elements, and five levels of function composition is sixteen calls against a ceiling of ten thousand. The limits are set where a document cannot reach them and a denial-of-service attempt cannot avoid them.

This is the part worth understanding, because it is the one a reader is most likely to get wrong when building something similar.

Per-operation limits do not compose. Each of the three lines below passes every limit in the table above. The first two build vectors of 1,501 elements, a fraction of maxCollectionSize. The third multiplies them, and the result is not the size of either operand, it is their product:

:a = map(1*x, 0:1500)
:b = transpose(a)
b * a

The answers on the first two lines are the vectors themselves, clipped to the width of the column. The third line is the point: Evaluating this expression would materialise 2,253,001 matrix cells, past the limit of 2,000,000 elements for one evaluation.

Before 1.0.0 this shape of expression killed the process. The same three lines with a twenty thousand element vector ask for 400,040,001 cells, which aborted with a heap fatal that no try in the engine, the host, or the test runner could contain. A cap on collection size, on matrix size, on anything measured per site, passes all three lines, because the fatal quantity is not any input. Three properties fix it, and all three are needed:

  • It is a total, not a per-site cap. Twenty-five individually legal collections in one expression are charged against one budget rather than each being waved through on its own.
  • It is consulted before allocating, wherever the size is knowable in advance. A matrix product works its size out from the two shapes, so the refusal above happens instead of the allocation rather than after it.
  • It is reset only by the outermost evaluation. The instruction counter is not: executeBytecode re-enters itself for function bodies and map or reduce transforms, and each reentrant call gets a fresh instruction count, so recursion refreshes its own allowance on the way in. A budget with that property would bound nothing. This one does not refresh, which is why maxFunctionCalls sits beside it.

The same total covers values that grow without being a collection at all. A string joined to itself, a bigint multiplied by itself and the exact decimal behind same-currency money each double every time the operator touches them, so a one-line function that squares or concatenates its argument, called a few dozen deep, once built a value of hundreds of megabytes from under a hundred characters while every per-operation limit passed. Each is now charged where it is made, against this same budget, so the doubling chain is refused with ALLOCATION_LIMIT_EXCEEDED at a couple of megabytes rather than at the process ceiling.

One more property of that reentrant machinery is worth stating for a host that runs programs on one machine through executeBytecode directly: a program that fails leaves the value stack at the depth it found it. A function body or a plot sample that had pushed operands before it threw used to leave them behind, so sixty-four failed samples of plot x + f(x) with f undefined crossed vm.maxStackDepth and reported the undefined function as STACK_LIMIT_EXCEEDED; now the first failed sample reports UNDEFINED_FUNCTION, and the stack is as the caller left it. In the same spirit, an opcode the machine has no handler for is refused at that instruction with MALFORMED_BYTECODE_UNKNOWN_OPCODE (its context carries the offset), rather than running as a no-op and failing a few instructions later with a stack underflow that names nothing. The OpRegistry a host can still hand to createVM has never been consulted by the dispatch loop, so an opcode registered there is refused the same way; the registry, sharedOpRegistry and createVM’s first parameter are deprecated and go in the next major, and a package that wants the machine to run its code declares pluginFunctions.

The cost of carrying the counter is an integer add and a compare at the sites that allocate, measured at roughly 3.5 percent on the virtual machine benchmark suite, which is at the edge of what run-to-run noise on one machine can resolve.

  • Element-wise matrix arithmetic and builtins that return a matrix (transpose, inv, det) are charged after allocating rather than before. Their output cannot be larger than their input, so the first such allocation is never the fatal one and the running total refuses the next.
  • A value that doubles under an operator (a string joined to itself with +, a bigint multiplied by itself, the exact decimal behind same-currency money) is charged on birth, the same after-the-fact backstop as the matrix builtins above. One doubling cannot be larger than twice an operand that was itself charged, so the first is never the fatal one and the running total refuses the next, well before V8’s own string or BigInt ceiling. Bigint exponentiation and shifts are refused before allocating instead, by a fixed bit ceiling, because either can ask for an arbitrarily large integer from a short line.
  • The per-document line cache keeps each line’s answer, and what those answers hold is bounded by vm.maxRetainedElements. The entries themselves (a line’s text, its compiled program and what it reads) grow with the document, which is bounded by performance.maxDocumentLines.
  • A refused line still runs. vm.maxRetainedElements bounds memory, not time: a line is refused once its answer is known, so the work of making it is already done, within the per-line limits above. A note of 500 such lines takes as long to evaluate as it did; it no longer keeps what they made.

A test suite checks what somebody thought to assert. A fuzzer checks the input nobody thought of, which for an engine like this one is most of the input it will ever see.

The fuzzer is seeded, so a finding reproduces exactly. It shrinks a finding to a minimal reproducer automatically, and commits that reproducer to a corpus that replays on every ordinary test run, so a fixed bug cannot come back quietly. Two generators feed it:

  • The expression grammar, drawing its vocabulary from a live engine, so a package added next month is covered without editing the fuzzer.
  • The bytecode virtual machine, generating and mutating opcode streams directly. executeBytecode is a public export from solve-engine/vm, which makes malformed bytecode a real caller surface rather than a hypothetical one, so it is fuzzed as one.

Three invariants are asserted, and only one of them can be observed from inside the process being tested. The runner therefore supervises a heap-limited child from outside:

  1. The process never dies. An out-of-memory abort is uncatchable, so it is observed as a child exit code.
  2. Nothing hangs. A wedged synchronous loop cannot time itself out, so it is observed as a heartbeat file that stopped advancing.
  3. Every failure is a well-formed EngineError, never a raw JavaScript exception reaching the host.

The 1.0.0 hardening run executed 2.6 million cases with no process death. It also found things review had not: a missing entry in the operand-width table that desynchronised every bytecode scanner after a date literal, nine raw exceptions reachable through the public vm export, and a nine-character expression that looped forever inside a single opcode where no limit could interrupt it.

Run it yourself:

Terminal window
npm run fuzz # random seeds, both generators
npm run fuzz -- --minutes=10 # a longer soak
npm run fuzz -- --seed=12345 # a specific run, reproducibly

A suite answers “does what I asserted still hold”. Before a release the question is the other one: “did anything change that I did not intend”, and no assertion can answer that, because the changes worth finding are the ones nobody thought to write down.

tools/differential/ runs a corpus through the last published build and the candidate build and classifies every disagreement by shape, so a reviewer makes one judgement per kind of change rather than one per row. The corpus comes from the documented examples, every string literal in the test suite, the recorded fuzz corpus and the grammar-aware generator. The clock, the timezone, Math.random and fetch are all pinned before the engine is imported, and each build is probed twice so that anything still unstable is dropped rather than reported as a difference.

The baseline is an installed package rather than a git checkout, because the tarball is the artefact a user receives and two working trees compare two things nobody ever ran. The 1.0.0 run compared 40,892 expressions against 1.0.0-beta.6. The candidate suffered zero process deaths against the baseline’s 68, which is the safety limits on this page doing their job in the only way that counts.

Report privately through GitHub’s advisory form rather than opening a public issue. A way around any bound on this page, or a way to make the engine consume unbounded time or memory, is a legitimate report.