What is taught wrongly

A map with no graticule

Ten rungs are handed control points, and a great many maps have none. Handed an outline with no labels on it at all, the method still works — and works better: the correspondence between ink and ground is recoverable exactly, because a similarity preserves ratios of arc length, and the margin on clean observations is 1.6 × 10¹⁰ against a graticule's 9.9 × 10⁶. What breaks it is noise, at three parts in a thousand.

Assumes The sheet moved before it was measured.

Every rung of this anchor is handed control points: crossings of the graticule whose latitude and longitude are known, read off the sheet, so the correspondence between the ink and the ground is given and the only unknowns are the projection and the reproduction.

A great many maps have no graticule at all. A road atlas page, an estate plan, a bird’s-eye town view, a tourist sheet, most of what is in a local archive — none of them carries a single labelled crossing, and the method this anchor is built on has nothing to be handed.

What such a map does have is a shape.

The whole of what an unlabelled map gives you. The outline of Japan as drawn on Conformal conic, delivered as an ordered list of page positions with nothing attached to any of them. No latitude, no longitude, no scale, no north. The rung's question is whether a projection can be recovered from that, and it can: the correspondence between the ink and the ground is found by sweeping the starting point round the curve and both directions, and the true candidate comes back with a residual of 2.88e-14 against the runner-up's 4.57e-4.
Fig. 1 The outline of Japan as drawn on a conformal conic, delivered as an ordered list of page positions with nothing attached to any of them. No latitude, no longitude, no scale, no north. This is the whole of what an unlabelled map gives, and the rung’s question is whether a projection can be recovered from it.

It can, and the expectation going in was that it would cost something. It costs nothing.

The correspondence is recoverable exactly

The difficulty of an unlabelled outline is that nothing says which point of the drawn curve is which point of the ground curve. What is available is the order, which survives any projection, and the arc length, which does not.

So the method is: resample both curves at equally spaced points along their own lengths, sweep the starting point of one round all n positions and both directions, fit a similarity at each, and take the best. That is a search over one discrete unknown and it is exhaustive.

The expectation was a residual. A projection stretches a curve unevenly, so the point a third of the way round the drawn outline should not be the image of the point a third of the way round the ground one, and the best correspondence should be a compromise with something left over.

There is nothing left over. The residual for the true candidate is 8 × 10⁻¹⁴ of the map’s own width, which is the arithmetic’s floor.

The reason is one sentence and it should have been obvious. The reproduction is a similarity — a scale, a rotation and two shifts — and a similarity multiplies every arc length by the same factor. So the drawn curve and the true candidate’s own curve have the same arc-length parameterisation up to a constant, the sweep finds the offset that aligns them, and the fit is exact.

The graticule was never carrying information the outline does not have.

Every candidate against an unlabelled outline. The twelve best candidates fitted to the outline of Japan drawn on Conformal conic, with no labels and the correspondence swept. The bar is the log of the residual, so a longer bar is a worse fit. The truth is separated from the runner-up by a factor of 1.6e+10, which is the same shape of answer a labelled fit gives: the residual alone says nothing and the margin says everything.
Fig. 2 The twelve best candidates fitted to that outline, with the correspondence swept. The bar is the log of the residual, so a longer bar is a worse fit. The truth sits at 10⁻¹⁴ and the runner-up ten orders of magnitude above it. That is the same shape of answer a labelled fit gives, and when the answer is not in the library is the rung establishing what to do with it: the residual alone is a ceremony and the margin is the measurement.

The refusal, and why it had to be rewritten

This rung was written with a prediction in it and the prediction was wrong, which changed what the machinery asserts.

The assertion drafted first was that the outline’s residual must be far above the labelled fit’s, on the reasoning that an arc-length correspondence is a compromise. Run, it came back three orders of magnitude below it, and the correct response is not to hunt for the bug — it is to work out why the measurement is right, which took one sentence about similarities, and then to assert what was measured.

So the check the file now carries is that the outline’s residual is below 10⁻⁹, which is a strong claim and a falsifiable one: it fails the moment the reproduction stops being a similarity, and the rung below this one is about a reproduction that does exactly that.

The second half of the refusal is that something must break the method, or the rung has measured nothing. Digitising error does, and the assertion requires it: at three parts in a thousand the outline’s margin must have collapsed by a factor of a hundred against its clean value, which it has.

And on clean observations it does better than a graticule

Putting the two methods side by side on the same map gives a result that runs the wrong way for the essay’s own premise.

An outline’s margin over the second-best candidate is 1.6 × 10¹⁰. A graticule’s, on the same projection over a comparable region, is 9.9 × 10⁶ — a thousand times smaller.

The reason is that an outline is a lot of constraints. Two hundred and forty points on a closed curve, every one of them required to fall in the right place relative to every other, against a graticule’s hundred and twenty-one crossings on a lattice — and a lattice is a much more forgiving object to fit than a curve with corners in it, because a wrong projection can distort a lattice into another lattice and cannot distort a coastline into another coastline.

So the honest statement of what a graticule is for is not that it makes identification possible. It is that it makes identification easy to explain: the correspondence is given, the fit is one linear solve, and nothing has to be swept.

What actually breaks it

Everything above is on clean observations, and the difficulty of an unlabelled map in a real archive is not the missing labels.

What an outline survives, and what a graticule survives. The margin over the second-best candidate, against how accurately the ink was measured, for an outline with no labels and for a graticule with them. On clean observations the outline wins — its margin is 1.6 × 10¹⁰ against the graticule's 9.9 × 10⁶ — because a similarity preserves ratios of arc length, so the correspondence is recovered exactly and there is nothing left over. Under noise the outline loses about two and a half times faster, and at 0.003 it names the wrong projection while the graticule is still right. A margin of one is where the wrong answer wins.
Fig. 3 The margin over the second-best candidate against how accurately the ink was measured, for both methods. On clean observations the outline wins by three orders of magnitude. Under noise it loses about two and a half times faster: at a hundredth of a per cent both are still right, at three hundredths the outline’s margin is 1.59 against the graticule’s 4.31, and at three tenths of a per cent the outline names an equal-area conic while the graticule is still right.

The mechanism is the sweep itself. Choosing the best of 2n correspondences is choosing the best of 2n chances to fit the noise, so the search that recovers the correspondence exactly on clean data overfits on noisy data — and the overfitting helps the wrong candidates more than the right one, because a wrong candidate has more to gain from a favourable alignment.

That is a general property of a search over a discrete nuisance parameter and it has a general name in statistics. What is worth having here is the size: the outline method’s usable noise floor is about a tenth of the graticule method’s, so a map measured to a part in a thousand can be identified from its graticule and not from its coast.

Three parts in a thousand of a half-metre sheet is a millimetre and a half. That is not a demanding standard — it is about what a careful hand with a rule achieves — so the outline method is usable, and it is usable with less margin for error than the anchor’s other nine rungs assume.

Which shape, and how much that matters

The outline used throughout is Japan’s, which in this collection’s region library is an ellipse on the sphere with its long axis running north-east. That is a shape with two useful properties and one limitation, and all three are worth stating.

It has a direction, so a wrong candidate cannot align with it by symmetry. A circular region would be nearly invariant under rotation, which is one of the freedoms the fit removes, so a round outline would carry far less information than an elongated one and every margin here would be smaller.

It has extent in two dimensions, so it samples the projection’s behaviour in latitude and in longitude at once. A boundary that ran along a single parallel would be the outline version of the degeneracy where the control points are measures for points, where an equirectangular’s standard parallel is not merely hard to find but exactly invisible.

And it is smooth, which is the limitation. A real coastline is not: it has headlands and bays at every scale, and the score is not stable at any scale shows its measured length has no limit at all. That roughness is more information rather than less — a rough curve is far harder for a wrong projection to imitate — so the margins here are a floor for what a real coast would give, and the noise sensitivity is a ceiling, because a rough curve’s shape survives smoothing better than a smooth one’s does.

How much of the coast has to be drawn

A closed island is a favourable case. A real map shows a stretch of coast, so the drawn curve is an arc of the ground curve and neither where it starts nor how much of the whole it covers is known.

How much of the coast has to be drawn. A real map shows a stretch of coast rather than a closed island, so the drawn curve is an ARC of the ground curve and neither where it starts nor how much of the whole it covers is known. Both are swept. The whole outline identifies the projection with a margin of 2.59; 80 per cent of it with 1.00; and by 80 per cent the method names Sinusoidal and puts the truth 4th. What runs out is not the correspondence but the shape: a short arc of any smooth curve looks like a short arc of any other.
Fig. 4 The same identification with only part of the outline drawn, sweeping the start and the extent as well as the correspondence. The whole outline gives a margin of 2.59; four fifths of it 2.14; and by three tenths the method names an equal-area conic and puts the truth third. What runs out is not the correspondence but the shape.

A short arc of any smooth curve looks like a short arc of any other, which is the same thing two projections that cannot be told apart measures for a graticule: over a small enough region every projection is the same picture, and the question is how small. For a graticule the answer is about twelve degrees of extent between two conformal projections. For an outline it is about half the curve, and the two are not the same quantity — one is an extent on the ground and the other is a fraction of a feature.

The residual also stops being exact once the arc is partial, at 2.4 × 10⁻⁴ rather than 10⁻¹⁴, and that is the search’s own coarseness rather than a geometric fact: the start and the extent are swept on a grid, so the best correspondence found is near the right one rather than at it.

What the search costs

The method is exhaustive and the arithmetic is worth stating, because the cost is what decides whether it is usable on a shelf of maps rather than on one.

For a closed outline: two directions times n starting positions times a four-by-four linear solve over n points, per candidate. At n = 120 and twenty candidates that is 4,800 similarity fits of 120 points, which runs in about sixty milliseconds.

For a partial arc it is worse by the number of extents swept — seven here — and by the finer start grid a partial curve needs, which brings it to about seven seconds for the same twenty candidates. Still nothing.

What would make it expensive is doing the correspondence properly. The sweep here assumes the drawn curve and the ground curve are traversed at proportional speeds, which a similarity guarantees; when they are not — under a shrunk sheet, or a generalised outline — the correspondence is a monotone reparameterisation rather than an offset, and finding the best one is a dynamic program over pairs of points rather than a sweep over one number. That is the standard tool for the job and it is a hundred times the work, and nothing here needs it because nothing here has a reason to reparameterise.

What this rung does not model

Three things stand between the measurement here and an archive, and all three make the problem harder.

The ink is generalised. A drawn coastline is a simplified version of the ground curve, drawn at a scale that decides how much detail survives — and a boundary that two features share is the essay about what a simplification tolerance does to an outline. That is not noise: it is a systematic smoothing that removes exactly the high-frequency detail an arc-length correspondence relies on.

The ground curve is a choice. The outline here is one of this collection’s own region shapes, whose ground position is exact by construction. A real coastline’s ground position depends on the tide, the date, the survey and the generalisation of whatever dataset supplies it, and two datasets disagree by more than the noise floor measured above.

And the reproduction may not be a similarity. The rung below this one shows that a paper sheet’s shrinkage is an anisotropic stretch, which does not preserve ratios of arc length — so the exactness that makes this whole method work is exactly what a shrunk sheet destroys. An outline read off an old sheet needs an affine fit at every one of the 2n correspondences, and an affine fit is the one that cannot separate a family.

That last one is the most interesting because it joins the two rungs. The exactness here rests on the reproduction being a similarity, and the rung below establishes that for a paper map it is not.

What an unlabelled map still cannot give

Two things a graticule supplies that an outline does not, and neither is the projection.

The scale. A similarity fit recovers a scale factor, and that factor is the map’s scale times whatever the reproduction did — which is exactly the ambiguity the rung below this one is about, arriving from a different direction. A graticule does not fix it either, so this is a tie rather than a loss.

The orientation on the ground. Here the two methods genuinely differ. A labelled crossing says where north is; an outline says only that the shape matches, and a shape matched to its mirror image is a different map with the same residual. The sweep searches both directions for that reason, and it returns which one won — so the method reports the handedness rather than assuming it, and a map printed in reverse, which archives contain, is identified correctly and flagged.

The whole of what an unlabelled map gives you. The outline of Chile as drawn on Sinusoidal, delivered as an ordered list of page positions with nothing attached to any of them. No latitude, no longitude, no scale, no north. The rung's question is whether a projection can be recovered from that, and it can: the correspondence between the ink and the ground is found by sweeping the starting point round the curve and both directions, and the true candidate comes back with a residual of 8.02e-15 against the runner-up's 1.00e-2.
Fig. 5 The same measurement on a different shape and a different projection: Chile’s outline drawn on a sinusoidal. Chile is long and narrow, thirty-nine degrees of latitude and ten of longitude, which is the case a north-south projection is chosen for and the case an outline carries most information about — a curve with a strong direction pins a rotation, and a curve spanning many latitudes pins whatever the projection does with latitude.

Where this leaves the anchor

Eleven rungs take identification from a fit to a set to a shape. What the last two together say about a real map in a real archive is a shorter list than the anchor’s optimism suggests.

If the map has a graticule and a modern reproduction, the method works and the margin is the quantity to report, at the thresholds when the answer is not in the library calibrates.

If it has a graticule and is on paper, the recovered parameters carry a bias that no fit can remove, and the answer is a curve in a two-dimensional space rather than a number.

If it has no graticule and a modern reproduction, the outline is enough, with a noise floor about a tenth as forgiving.

And if it has no graticule and is on paper, the shrinkage breaks the arc-length exactness that the outline method rests on, and nothing in this anchor covers it.

That last case is the commonest one in any archive, which is an uncomfortable place for an anchor to have arrived after eleven rungs and is where it honestly is.

What a graticule is actually for

Reading the rung back, the conclusion inverts a premise the anchor has carried since its first essay.

A map does not say what it is opens the anchor with the observation that a map’s projection is a sentence in a corner or nothing at all, and that the graticule is enough to recover it. The second half of that is true and the emphasis is wrong: the graticule is enough and it is not necessary, and on clean observations it is not even the best available evidence.

What a graticule is for is something else, and it is worth naming because it is what the anchor has been using it for without saying so.

It makes the correspondence explicit, so the fit is one linear solve rather than a search. That is a computational convenience and, on a shelf of maps, a real one.

It is checkable by eye. A person can look at a crossing labelled 45° north and see whether the fit put it there. A correspondence found by sweeping an outline cannot be inspected that way, and a wrong offset produces a plausible-looking answer with no visible symptom.

And it is what the cartographer intended to be read. The ink of a coastline is a depiction; the ink of a graticule is a claim about coordinates, made deliberately. Fitting to it is fitting to what the map asserts, and fitting to the coast is fitting to what the map depicts, and those are different kinds of evidence even when they carry the same information.

That third one is why the graticule remains the right thing to use when there is one, and why the outline method is a fallback rather than an improvement despite winning on margin.

What comes next: the map that was compiled rather than projected

Every rung here assumes the map was drawn in one projection, correctly, once — that somewhere behind the ink there is a single pair of formulae and a single set of parameters that the fit is trying to recover.

A great many historical maps were not made that way. They were compiled from other maps of different projections, joined along seams nobody recorded, adjusted by eye where they disagreed, and redrawn by hands that were copying rather than projecting. For such a map no single projection fits any of it, and the residual left by the best candidate is a mixture of construction error and compilation error with no way to separate the two.

Telling those apart is the thing this anchor still cannot do — and it now has one more instrument towards it than it had. A compiled map’s seams are places where the shape stops matching while the graticule, if there is one, goes on looking regular; so the residual field of an outline fit, sampled along the curve, should have structure at the seams that a projection error does not. A residual has more than one explanation is the rung about reading a residual for what else it could be, and a per-point residual along a coastline is a richer object than the single number a graticule fit returns.

That is a lead rather than a measurement, and it is the one the anchor should take next.

Named alongside this one

Essays reaching for the same objects. Nobody chose these; they are what the concept index makes visible.

The objects this essay names

Each one links to every other essay that touches it.

AffineControl pointsDegeneracyEstimatorGeneralisationIdentificationMarginPurposeReproductionResidualShape matchingSimilarity transformationToleranceVerification