This site describes solve-engine as it is on main: 2.43.0, which npm does not have yet. npm installs 2.40.0, so a page may show an answer that version does not give yet.
Statistics
Packages:
MATHPHRASES_PACKAGE(averages, spread, comparisons) andSTATISTICS_PACKAGE(correlation, regression, percentile, z-score, and the probability distributions). Both registered bycreateEngine(); for a slimmer engine, register them explicitly (see choosing packages).
Statistics are the numbers that sum up a list of values: where its middle is,
how spread out it is, and how two lists move together. The average (the
mean) adds the values and divides by how many there are; the median is the
middle value once they are in order, which a single very large or very small
value cannot drag away. Write the question the way it is said, with the values
after of:
average of 10, 20, 30 // 20median of 1, 5, 3 // 3The spreadsheet spellings
Section titled “The spreadsheet spellings”The same questions go by other names in a spreadsheet or another notepad, and
those names are read too. sum of is total of, mean of and avg of are
average of, and min of, max of and product of give the least value,
the greatest, and all of them multiplied together.
sum of 1, 2, 3 // 6mean of 1, 2, 3 // 2min of 4, 2, 9 // 2max of $5, $7 and $3 // $7.00product of 2 m, 3 m // 6.00 m²A spreadsheet writes the values in brackets after the name instead, and that works as well:
sum(1, 2, 3) // 6average(4, 8) // 6mean(1, 2, 3) // 2median(1, 5, 3) // 3stdev(1, 2, 3) // 0.82stdev(...) is the population standard deviation, the same as stdev of and
standard deviation of on this page. A spreadsheet’s STDEV is the sample
form, which is written sample stdev of (see spread and shape).
Two other readings of the same brackets are kept. A call that names each
element and then gives the list is map-reduce
(sum(x, [10, 20, 30])), and a call over lines is a
line range (average(line 1 : line 4)). mean,
median and stdev stay ordinary names wherever no bracket follows them, so
mean = 4 is still a variable; a function of your own under one of those three
names is refused by name, since the call would never reach it.
A range is not one of the readings. 1:3 is the run of whole numbers 1, 2 and 3
only as the list of sum, prod, map or reduce (see
ranges); everywhere else a colon
between two numbers is a clock time, so average(1:3) would be the average of
1:03 AM. Rather than answer that, or ask for the line reference it was never
given, the call says what the colon is here and how to write the values. sum
is the one aggregate that reads a range, and total, its synonym, reads one
the same way when the range is its only argument.
average(1:3) // ERROR: In average(...), 1:3 is a clock time, not a range, and a time cannot be averaged: a colon between two numbers is a range only as the list of sum, prod, map or reduce. To average numbers, list them with commas, as in average(1, 2, 3).average(1, 2, 3) // 2sum(1:3) // 6total(1:3) // 6A list that carries units
Section titled “A list that carries units”A list of quantities answers in a unit rather than as a bare number, and the unit is the first one written. Everything after it is read in that unit, so a set spelling one measure two ways adds up rather than needing to be retyped.
total of $4.99, $12.50, $3.20 // $20.69total of 1.2 km, 3 km, 800 m // 5.00 kmtotal of 800 m, 1.2 km, 3 km // 5,000.00 mThe last two lines are the same three distances and the same answer, shown in whichever unit the list opened with.
A list mixing measures has no unit both halves can be read in, so it is refused
rather than answered: average of 5 kg, 3 m reports mass and length cannot be
averaged. Adding a mass to a length is not a harder sum, it is a different
question, and a number here would be confidently wrong.
A value that is not a number at all is refused the same way, by every form on
this page: a piece of text, a date or a clock time, a bracketed list, a colour.
Each of those used to be read as some number it happened to convert to, most
often zero, so total of "Travel" reported nothing spent. A quoted name is text,
not a set of lines; to add up the lines under a heading, name it as a section,
total of section "Travel" (see sections), or tag the lines
and total the tag (see category tags). A bracketed list
is one value, so write its members out with commas instead. count of counts
anything, text included.
total of "Travel" // ERROR: Text cannot be added: only numbers and quantities can. To gather the lines under a heading, write total of section "Travel"; to gather tagged lines, use "total of #tag".total of [1, 2, 3] // ERROR: A bracketed list cannot be added: only numbers and quantities can. List the values with commas instead, as in "total of 1, 2, 3".total of 1, 2, 3 // 6The same rule holds wherever a set is named, so a column of money totals to
money whether the lines are gathered by position with
total above, by name with a
category tag or a section, or by
a line range.
1.2 km3 km800 mtotal above // 5.00 kmThe boundary is a bare number sitting in a list of quantities. It contributes
its magnitude, which is what a count written beside a column of measurements has
always done, so total of 1 km, 500 is 501.00 km. And count of counts, so it
carries no unit at all.
A list of percentages
Section titled “A list of percentages”A percentage is a share of something, such as a tax rate or a discount, and a set of shares of the same whole adds up to a share as well: a 10% rate and a 20% rate are 30% between them, and their average is 15%. So every aggregate on this page answers a set of percentages with a percentage, whether the values are written with commas or brackets, gathered by line, by heading or by tag.
sum(10%, 20%) // 30.00%total of 10%, 20% // 30.00%average of 10%, 20% // 15.00%max(10%, 20%) // 20.00%median of 10%, 20%, 40% // 20.00%stdev of 10%, 20% // 5.00%A product of percentages is a share of a share, so it is a percentage too, and every one of these answers is worked out from the decimals as written, so it equals the percentage it shows (see multiplying and dividing percentages):
product of 10%, 20% // 2.00%sum(10%, 20%) == 30% // true10% // 10.00%20% // 20.00%total above // 30.00%average above // 15.00%A percentage beside a plain number or an amount is refused by name, because the
two have no reading in common: 100 + 10% is 110, a tenth more than 100, but a
list has no order that says the percentage is a share of the number, and adding
the percentage’s fraction (0.1) to 100 would be confidently wrong. The refusal
names the percentage and the fraction it stands for, so the value can be
rewritten either way.
sum(10%, 100) // A percentage (10%) and a number cannot be added together: a percentage is a share of an amount, not an amount of its own. Write every value as a percentage, or write 10% as the number 0.1; to raise an amount by a percentage, write it as 100 + 10%.average of 10%, $5 // A percentage (10%) and an amount in USD cannot be averaged together: a percentage is a share of an amount, not an amount of its own. Write every value as a percentage, or write 10% as the number 0.1.The boundary: a variance of percentages would be in percent squared, which is
not a percentage, so it is refused, and the standard deviation gives the same
spread as a percentage. count of counts percentages like anything else. A
table column is the exception to gathering them: its summaries refuse a
percentage cell (see table columns). A product of
percentages is still worked out with *, so product of 10%, 20% is the
fraction 0.02. These aggregates used to read each percentage as its fraction,
so sum(10%, 20%) was 0.30 and sum(10%, 100) was 100.10, and a line range
or total above over percentages was refused.
Spread and shape
Section titled “Spread and shape”average and median find a list’s centre; these four say how spread out it
is. Each reads a bare list, or a table column (see
table columns).
standard deviation of 2, 4, 4, 4, 5, 5, 7, 9 // 2variance of 2, 4, 4, 4, 5, 5, 7, 9 // 4spread of 3, 7, 2, 9 // 7mode of 4, 2, 4, 3, 4, 2 // 4Standard deviation and variance take the population form by default, which is what a note over a fixed column of readings usually is: the whole set, not a draw from a larger one. The sample form (dividing by one less) is a named variant, for both.
sample standard deviation of 2, 4, 4, 4, 5, 5, 7, 9 // 2.14sample variance of 2, 4, 4, 4, 5, 5, 7, 9 // 4.57spread is the largest minus the smallest, spelled that way because range
already means a start:end interval elsewhere. A tie for mode goes to the
value that reached the top count first, so the same list always gives the same
answer.
The spread of quantities
Section titled “The spread of quantities”The spread forms follow the rule for a list that carries units: the values are read in the first one’s unit, and the answer is in it. So 1 kg and 1000 g, which are the same mass, have no spread at all, and the most common of 1 kg, 1000 g and 2 kg is 1 kg.
standard deviation of 1 kg, 1000 g // 0.00 kgstandard deviation of $10, $20, $30 // $8.16mode of 1 kg, 1000 g, 2 kg // 1.00 kgspread of 1 kg, 1000 g // 0.00 kgA variance is the average of the squared distances from the mean, so it is in
the square of the data’s unit. The engine has a unit for that only when the data
are lengths: the square of a length is an area. For anything else (kilograms
squared, dollars squared, degrees squared) there is no unit to give the answer,
so a variance of those is refused by name, the same way (2 kg)^2 is. The
standard deviation is the same spread, in the data’s own unit, and is the one to
ask for.
variance of 2 m, 4 m // 1.00 m²variance of 1 kg, 1000 g // A variance of quantities in kg would be in kg squared, which has no unit: only a length squared has one, an area. The standard deviation is the same spread in kg.Two measures in one list are refused as they are everywhere on this page, and a list with an infinity in it has no standard deviation or variance, since an infinite value is no finite distance from the mean.
standard deviation of 1 kg, 2 m // mass and length cannot be used in a standard deviationstandard deviation of 1/0, 1000 // A standard deviation of a list with an infinity in it has no value: an infinite value has no finite distance from the mean.Weighted average
Section titled “Weighted average”A plain average treats every value as equal; a weighted average pairs each value
with its own weight through at. The weights are normalised by their own total,
so they need not sum to 1 or to 100%.
weighted average of 72 at 30%, 88 at 70% // 83.20weighted average of 4.0 at 3 credits, 3.0 at 1 credit // 3.75weighted average of 10 at 2, 20 at 3 // 16weighted mean of 10 at 2, 20 at 3 // 16weighted mean of is the same form under the other name for an average.
The grade-point case divides by the four credits, giving 3.75; percentages that
already sum to 100 come out unchanged. A trailing label on a weight (3 credits)
is read for the number and the word ignored. A value written with no at clause
is reported as an error rather than given a silent weight of 1, since guessing
one would quietly change the answer of a list that was simply mistyped.
Relationships between two lists
Section titled “Relationships between two lists”The stats above describe one list. These describe how two lists move together:
whether taller people also tend to be heavier, and by how much. Give the two
lists as [bracketed] sets of the same length.
Correlation is a single number from -1 to 1: 1 means the two rise together in perfect step, -1 means one rises exactly as the other falls, and 0 means no straight-line relationship at all.
correlation of [1, 2, 3, 4] and [2, 4, 5, 8] // 0.98The line of best fit is the straight line that sits closest to the points.
slope is how steep it is (how much the second list changes per step of the
first), and intercept is where it crosses zero.
slope of [1, 2, 3, 4] and [2, 4, 5, 8] // 1.90intercept of [1, 2, 3, 4] and [2, 4, 5, 8] // 0r squared is the share of the variation the line explains, from 0 to 1; it is the correlation squared, and is written as a function.
rsquared([1, 2, 3, 4], [2, 4, 5, 8]) // 0.96Each two-list form also has a call spelling, correlation([a], [b]) and so on.
Two lists of different lengths, or fewer than two points, are reported as an
error rather than answered.
Position in a list
Section titled “Position in a list”A percentile is the value a given share of a list sits below: the 90th percentile is the value nine tenths of the data fall under. The share is a number from 0 to 100, and the 50th percentile is the median.
percentile([1, 2, 3, 4, 5, 6, 7, 8, 9, 10], 90) // 9.10A z-score says how far one value sits from the average, measured in standard deviations: a z-score of 2 is two standard deviations above the mean.
zscore(9, [2, 4, 4, 4, 5, 5, 7, 9]) // 2A z-score is also where the bell curve comes in: normalcdf turns one into the
share of a normal population below it. The normal, binomial, Poisson and t
distributions have their own page, probability distributions.
Every statistics call refuses an argument it does not read, rather than quietly leaving it out:
percentile([1, 2, 3], 50, 9) // ERROR: percentile takes 2 arguments, but was given 3, as in percentile([1, 2, 3], 90)zscore(1, [1, 2, 3], 5) // ERROR: zscore takes 2 arguments, but was given 3, as in zscore(5, [1, 2, 3])Comparisons and fractions
Section titled “Comparisons and fractions”Picking the biggest or the smallest of a few values, or the point halfway
between two, is a question in its own words too. larger of gives the largest
of the values joined by and, and smaller of the smallest. Each takes as many
values as are listed, so a third and adds a value to compare, never to sum.
greater of and lesser of are the same phrases, and gcd of and lcm of read
a list the same way. half of halves one value, and midpoint between gives the
value exactly halfway from one to the other.
larger of 10 and 4 // 10greater of 10 and 4 // 10larger of 10 and 4 and 12 // 12smaller of 10 and 4 and 12 // 4lesser of 10 and 4 // 4gcd of 12 and 18 and 8 // 2half of 50 // 25midpoint between 10 and 20 // 15Each value may be a sum, so larger of 1 + 1 and 3 compares 2 with 3. The
values are joined by and, not commas: for a list with commas, write
max(10, 4, 12) or min(10, 4, 12).
Ranges and clamping
Section titled “Ranges and clamping”Clamping a value holds it inside a range: a value inside is left as it is, and
one outside is moved to the nearer end. It is how a score is capped at the
maximum, or a setting kept within its limits. Write clamp, the value, and the
two ends after between.
clamp 15 between 1 and 10 // 10Proportions
Section titled “Proportions”A proportion says that two ratios are equal: 2 is to 4 as 5 is to 10, since each
second number is twice the first. Given three of the four numbers, the engine
finds the missing one, written what, which is the everyday way to scale a
recipe or a map distance.
2 is to 4 as 5 is to what // 10random number between 1 and 10 gives a fresh value in that range each time, so
it is shown here rather than pinned to one result.