The impossibility

How wrong a flat picture has to be

The rung below proves no flat picture of four places is exact and leaves the size of the failure to a determinant nobody can read. Measured directly, the least error falls as the square of how much sphere the places span — fitted exponent 2.0088 — and the same exponent comes back from five different arrangements while the constant in front of it moves by a factor of seventeen.

Assumes Four cities that cannot be drawn to scale.

Four cities that cannot be drawn to scale establishes that six distances taken off a sphere are not the six distances of four points on a sheet of paper, and decides it exactly, with a determinant. Then it hands back a number nobody can interpret: −3.344471 × 10²³, in square kilometres to the sixth power.

A determinant is a fine instrument for a yes or no. It is a poor one for a how much, because its units are the units of a volume that does not exist and its size scales as the sixth power of the set. Normalising by the diameter, which that rung does, removes the size and leaves a shape reading — which is better and is still not a quantity anybody has an opinion about.

The quantity anybody has an opinion about is the error. So this rung measures it, at every span from a county to a hemisphere, and finds a law.

The least error a flat picture can have, against how much sphere it spans. The same configuration of places, shrunk about its own centroid so that every bearing is kept and only the span changes, with the least worst-case relative error of the best flat picture at each size. Both axes are logarithmic. The fitted slope over the rows below ninety degrees is 2.0246: the error falls as the SQUARE of the diameter. That is Gauss's theorem arriving as a number for a finite set — curvature is a second derivative, so its first effect on a distance is quadratic in the separation — and it is why a county fits on a sheet and a hemisphere does not.
Fig. 1 The least worst-case relative error of a flat picture of London, New York, Tokyo and Sydney, against how much of the sphere they span. The configuration is the same at every point — the four are shrunk about their own centroid so that every bearing from the centre is unchanged and only the span varies — so the curve is a function of span and of nothing else. Both axes are logarithmic. Below ninety degrees the fitted slope is 2.0246; below forty-five it is 2.0088.

A slope of two on log axes is a square law. Halve the span and the least possible error falls by a factor of four; quarter it and the error falls by sixteen.

The number on that axis has to be defined before it can be believed, and two choices in it are not obvious.

The error is a ratio, and the picture is allowed a scale. A drawing has no natural size, so the only meaningful statement is about the proportions between its separations. Every arrangement below is therefore scored after being given the one free multiplier that a scale bar represents, and the multiplier chosen is the one that makes the score best — the geometric mean of the extreme ratios, because a minimax of a ratio is symmetric in the logarithm and not in the ratio itself. Centring on the arithmetic mean instead reports an error up to twice as large on these sets, which is a fact about the scoring rather than about the places.

The score is the worst edge, not the average. A picture in which five of six distances are perfect and the sixth is out by ten per cent is a bad picture, and an average would call it a good one. The worst edge is the claim a scale bar actually makes.

Minimising a worst case is not what the standard method does. Classical multidimensional scaling — double-centre the squared distances, take the top two eigenvectors — minimises a least-squares criterion in the squared distances, which is a different objective and lands somewhere else. Every point on the curve above is therefore a pattern search: start at the scaling solution, move each dot in turn by a step, halve the step when no move helps, and stop when the step is below a millionth of the diameter. A minimax objective is made of kinks, so no derivative method applies, and a pattern search is the crude thing that works.

Three ways of scoring the same picture, and one exponent. The same ladder of spans scored three ways: the worst edge of the best possible arrangement, the worst edge of the arrangement classical multidimensional scaling produces, and the root-mean-square error over all the edges. The three curves are parallel and are not the same curve — classical scaling sits a factor of 1.31 above the optimum at every span, and the root-mean-square sits below both because averaging hides the edge the minimax is about. The fitted exponents are 2.0246, 2.0223, 2.0111: the constant in front is a property of the estimator and the exponent is a property of the sphere.
Fig. 2 The same ladder of spans scored three ways: the worst edge of the best arrangement, the worst edge of the arrangement classical scaling produces, and the root-mean-square over all six edges. The three curves are parallel and are not the same curve. Classical scaling sits a fixed factor of 1.31 above the optimum at every span — it is not converging to it — and the root-mean-square sits below both, because averaging over six edges hides the one edge the minimax is about. All three fitted slopes are within a hundredth of two.

That figure separates the two things a measurement of this kind can be about. The exponent survives every change of scoring rule, which means it belongs to the geometry. The constant in front does not, which means it belongs to the instrument — and reporting a constant without naming the rule that produced it would be exactly the failure the ranking depends on the region records for the classical distortion indices.

Why the exponent is two

Gauss’s theorem says the sphere and the plane differ in a second derivative of the metric. A second derivative is the first term in an expansion that vanishes for the first two orders, so the first place it can appear in a distance between two places a small angle θ apart is at θ³ in the length, which is θ² in the relative length.

That is the whole derivation, and it is why the exponent is not 1 and not 3. A first-order difference between the two surfaces would mean they had different lengths at every scale, which they do not: shrink any configuration far enough and a flat picture becomes exact to any stated precision. A third-order difference would mean flat pictures were far better than they are.

The same arrangement at three sizes. The best flat picture of London, New York, Tokyo, Sydney, with the configuration shrunk about its own centroid so that every bearing from the centre is unchanged and only how much of the sphere it covers varies. At 152.8° across the worst edge is out by 0.629%; at 39.7° it is 0.0299%; at 9.9° it is 0.0018%. The pictures are indistinguishable to the eye and the numbers under them fall by a factor of sixteen for each factor of four, which is what a square law looks like when it is drawn rather than plotted.
Fig. 3 The best flat picture of the same four cities at three sizes, an octave and a half apart in span. The three drawings are indistinguishable — the same arrangement, shrunk — and the numbers under them fall by roughly sixteen for each factor of four in span, from 0.629 per cent through 0.0299 to 0.0018. There is nothing in the pictures to see, which is the point: the failure is not a visible deformation of the shape but a discrepancy in its proportions, and by the third panel it is three parts in a hundred thousand.

The same exponent is reached in this field from the other direction. How big a triangle it takes measures the spherical excess of a triangle, which is proportional to its area and therefore to the square of its size, and how small is flat enough asks at what extent a region stops being flat to a stated tolerance and finds the same quadratic. This rung is the finite-set version of both, and the agreement is worth something precisely because the three calculations share no code: one integrates a curvature over a cap, one sums three angles, and this one runs a pattern search over eight coordinates.

The exponent is the sphere’s and the constant is not

One configuration cannot separate a property of the surface from a property of the places. Five can.

The exponent belongs to the sphere and the constant to the arrangement. The square law fitted separately on five configurations of places — from five towns inside one country to sixteen cities on every continent — each shrunk through its own ladder of spans. The exponents agree: 2.025, 2.016, 1.999, 2.013, 1.978. The constants in front of them do not, and span a factor of 17. That separation is the point of running five sets rather than one: how badly a particular arrangement of places resists a flat sheet is a fact about the arrangement, and that the resistance falls as the square of the span is a fact about the surface they sit on.
Fig. 4 The square law fitted separately on five configurations — five towns inside Britain, four world cities, five, eight, and sixteen on every continent — each shrunk through its own ladder of spans and each fitted below ninety degrees. The exponents are 2.0246, 2.0163, 1.9986, 2.0132 and 1.9777, which agree to about one per cent. The constants in front span 1.80 × 10⁻⁷ to 2.99 × 10⁻⁶, a factor of seventeen. The sphere decides the exponent; the arrangement decides how much trouble it is in at any given span.

The constant is a real quantity and is worth understanding, because it is where all the design freedom in a real map of a real region lives. A set of places strung out along a line has a small constant, because a line embeds in a plane whatever its length; a set spread evenly over a cap has a large one. The sixteen-city set has the largest constant here and the four-city set the smallest, and the four-city set is not smaller because it has fewer places — the five-town Britain set has a constant seven times larger with one more place in it.

What decides it is how much of the two-dimensional extent of the sphere the places actually occupy, and that is a statement about arrangement rather than about count. Which is the same reason a country long in one direction and narrow in the other is easy to map, and is the design principle behind every choice of aspect this collection has priced: fitting the aspect to the region turns a wide region into a narrow one by rotating the axis, and the constant here is exactly what that rotation is reducing.

The two things a count does not decide

The figure above invites a reading it does not support, and it is worth blocking. The five sets have four, five, five, eight and sixteen places in them, and the constants are not ordered by that count.

Adding a place cannot lower the difficulty. A larger set contains the smaller one, so every distance the smaller set had to get right the larger one must also get right, plus more; the least error can only rise. The sixteen-city set does have the largest constant here and that inequality is why.

But the count is not what sets the size of it. Britain’s five towns carry a constant seven times the four world cities’, while spanning a twenty-seventh of the sphere. What separates them is that the four cities happen to sit close to a great circle — London, New York, Tokyo and Sydney are strung round the world rather than spread over a cap — and a set on a curve embeds in a plane far better than a set filling an area, because a curve is one-dimensional and the flat sheet’s difficulty is with the second dimension.

That is the same structure where a pseudocylindrical puts its error finds in a projection’s own freedom, arriving on a finite set: what a flat sheet has trouble with is extent in two directions at once, and a region that is long and thin is cheap however long it is.

Where the law stops

The fitted exponents above all exclude the largest spans, and the reason is visible in the hero figure: the top-right point sits above the line.

The law is a leading-order expansion in the span, so it must fail when the span stops being small — and on a sphere “small” has a definite meaning, because there is a largest possible separation. At 152.83 degrees the four cities span most of the available range, and the measured error of 0.629 per cent is above the 0.52 the fitted line predicts. Include that point and the fitted exponent rises from 2.0088 to 2.0703, which is the expansion’s next term making itself felt rather than the law being wrong.

The least error a flat picture can have, against how much sphere it spans. The same configuration of places, shrunk about its own centroid so that every bearing is kept and only the span changes, with the least worst-case relative error of the best flat picture at each size. Both axes are logarithmic. The fitted slope over the rows below ninety degrees is 1.9777: the error falls as the SQUARE of the diameter. That is Gauss's theorem arriving as a number for a finite set — curvature is a second derivative, so its first effect on a distance is quadratic in the separation — and it is why a county fits on a sheet and a hemisphere does not.
Fig. 5 The same measurement on sixteen cities on every continent, spanning 167 degrees at full size. The fitted slope below ninety degrees is 1.9777 and the four points that carry it lie on a line to the width of the marker. The two points above ninety do not: at 127 degrees the least error is 4.68 per cent against the 2.6 the fitted line predicts, and at 167 degrees it is 71.7 per cent against 4.9. The law has a domain, the domain has an edge, and the edge is where the set starts containing near-antipodal pairs.

For the sixteen-city set the departure is dramatic. At 63 degrees the least error is 1.10 per cent; at 127 it is 4.68, which the square law under-predicts by a factor of not quite two; at 167 it is 71.7 per cent, which it misses by more than an order of magnitude. Beyond about a hundred and thirty degrees the set contains pairs of places approaching antipodal, and near-antipodal pairs are where every flat picture of the sphere breaks down completely — the same neighbourhood the route with no shortest path finds the distance formula itself stops converging in.

The mechanism is worth naming rather than gesturing at. Two places nearly opposite each other are joined by a great circle whose length barely changes as either place moves, because the derivative of the distance with respect to position vanishes at the antipode; on a flat sheet the distance between two dots has no such stationary point anywhere. So a set containing a near-antipodal pair is asking the plane to reproduce a feature the plane does not have, and the failure stops being a small correction to a good picture and becomes a structural one. That is why the sixteen-city curve bends up rather than merely leaving the line.

So the honest statement of the law has a domain attached to it. Below about ninety degrees of span the least error of a flat picture is c·θ² with c between 10⁻⁷ and 3 × 10⁻⁶ per square degree, depending on the arrangement. Above that the expansion has run out and the number has to be computed.

The determinant and the error agree, which they had to

The rung below reads its exponent off the Cayley–Menger determinant and gets 2.06; this one reads it off the error of an optimised picture and gets 2.0088. Those are two entirely separate calculations — one is a five-by-five determinant in six squared distances with no optimisation in it at all, the other is a pattern search over eight coordinates minimising a worst case — and they are measuring the same thing through different instruments.

They had to agree and it is worth saying why they are still worth running both ways. The determinant is 288 times the squared volume of a tetrahedron that does not exist, so its square root is a length-cubed and its normalised sixth root is a length; a dimensional argument alone says the determinant should carry three times the exponent of a linear error, and the observed ratio of 2.06 to 2.01 is nothing of the kind. The agreement is therefore not dimensional bookkeeping: the determinant’s dependence on the span is dominated by the shape term, which is the same second-order departure the error measures, and the sixth-power normalisation removes the rest.

The reason to keep both is that a shared exponent from unshared code is the only kind of confirmation available here. There is no closed form for the least worst-case error of an arbitrary set of places, so nothing can be checked against an exact answer; what can be done is to measure the same underlying quantity two ways and require the exponents to match, which is the habit measuring curvature from inside applies to curvature itself.

What the law is for

A square law with a measured constant is a tolerance calculator, and it answers a question nobody in this collection had been able to answer.

A survey plan, a floor plan, an orienteering map and a national atlas are the same object at four sizes, and the question of when the flat sheet stops being adequate is a question every one of them answers implicitly and none states. With a constant of 10⁻⁶ per square degree, a flat picture of places within one degree of each other — a hundred and eleven kilometres — is wrong by a part in a million, or eleven centimetres in a hundred kilometres. Within ten degrees it is a part in ten thousand, or eleven metres. Within a hundred degrees it is a part in a hundred, and every distance on the sheet is wrong in the second digit.

Those three sentences describe, in order, a cadastral survey, a national grid and a world map, and the reason the three are treated as different problems by different professions is not conventional. It is the square law.

The threshold that matters in practice is where the error crosses the precision of the measurement rather than where it becomes visible. A tape-and-theodolite survey good to one part in a hundred thousand meets the flat sheet’s own error at about three degrees of extent, or three hundred and thirty kilometres; satellite positioning good to one part in ten million meets it at a third of a degree, thirty-seven kilometres. So improving the instrument by two orders of magnitude shrinks the region a flat sheet can carry by one, which is the square law read backwards and is the reason the plane-survey tolerance in how small is flat enough has kept shrinking through the history of the subject without anybody deciding it should.

It also says something about the shape of the trade-off that the differential account does not. The trade-off is two lines shows that a projection cannot be both conformal and equal-area and that the two properties bound each other; that is a statement about which errors a map can trade against which. The law here is about neither: it says that the total available error, whatever a map maker chooses to spend it on, is fixed by the span before any design decision is made. The design decides where the error goes and the span decides how much there is.

What a square law is not

Two readings of the curve are available and only one of them is right, and the wrong one is the more comfortable.

The wrong reading is that a flat picture becomes correct below some size. It does not. The determinant is non-zero at every span in every figure here, the negative eigenvalue is negative at every span, and the exact statement of the rung below holds at three degrees exactly as firmly as at a hundred and fifty. What falls quadratically is the magnitude of a failure that never becomes an absence.

The right reading is that the failure becomes smaller than something else. Below three degrees a flat picture of a set of places is wrong by less than a tape measure can detect; below a third of a degree it is wrong by less than a satellite receiver can detect; and there is no span at which it is wrong by nothing. What the instrument is, and therefore where the crossing sits, is a decision somebody makes and writes down — which is why every specification in surveying that permits plane methods states a maximum extent, and why the number in it differs between countries and between decades.

That distinction matters because the two readings prescribe different behaviour when the tolerance tightens. Under the first, a small region is settled and a tightening tolerance changes nothing about it. Under the second, tightening the tolerance by a factor of a hundred divides the region a flat sheet can carry by ten, and a specification written for one instrument is silently wrong for its successor. The subject has met that transition twice — once when triangulation replaced chaining and once when satellite positioning replaced triangulation — and on both occasions the extent limits in national specifications had to come down.

The next thing to compare it against

Every number in this rung is the error of the best flat arrangement of a specific set of places. Nothing in it is a map. The arrangement has no graticule, no inverse and nothing to say about a place that was not in the set, and the optimisation that produced it would move every dot if a seventeenth city arrived.

A projection gives all of that up in exchange for being a rule fixed in advance, and the price of the exchange has never been measured here. The comparison is direct — put the same set of places through every projection in the library, score them with the same worst-edge rule, and read the ratio — and the answer turns out to depend violently on how many places are in the set. On four cities the free picture is thirteen times better than the best projection. On sixteen it is better by eight per cent.

What this makes readable

Essays that name this one as a prerequisite.

Named alongside this one

Essays reaching for the same objects. Nobody chose these; they are what the concept index makes visible.

What links here

Every essay whose body links to this one.

The objects this essay names

Each one links to every other essay that touches it.

Cayley mengerClosed formConvergence rateDistance matrixEmbeddingEstimatorGaussian curvatureMinimaxOptimisationQuadratic lawScaleToleranceVerification