A residual has more than one explanation
The base of this ladder takes a map that carries no statement of its projection, fits every candidate in a library to its graticule crossings, and names the one that fits best. It works: the winner wins by a factor of 3.5 × 10⁷ on a clean example, and the fourth rung taught it to refuse when the truth is held out.
What none of the four rungs has asked is what a residual is. The method treats it as the projection’s misfit, and a residual is not labelled. A map made from coordinates that were never converted between datums has the same symptom — a small, smooth, systematic departure from every candidate — and until this rung the method had never been given one to see what it would do.
The datum shift is invisible, and that is the good news
The first result is the reassuring one and it deserves to be stated first.
The identification is not fooled by a datum shift. At every region size from half a degree to sixteen, with control coordinates on OSGB36 read as though they were on WGS84 — a displacement of about a hundred metres on the ground — the method still names the right projection, and the residual it leaves is a flat 1.1 × 10⁻⁶ of the map’s own width. That is arithmetic noise.
The mechanism is the fit. Identifying a projection requires removing the reproduction — the unknown scale, rotation and offset with which somebody drew or scanned the map — and the fit that removes it is a similarity in the plane. A datum shift over a region is, to a very good approximation, a translation with a small stretch: it is a rigid displacement of the reference ellipsoid plus a scale term, seen through the projection. The similarity fit removes almost all of it, and what it does not remove is below the noise.
So a map identified from control points does not need its datum to be right. That is worth knowing, because a historical map’s datum is usually less well known than its projection, and a careless copy hides its own distortions on top of both, and it would be awkward if the second could not be recovered without the first.
And that is also the bad news
The other side of the same fact is that the method cannot detect a datum shift either. A residual of 1.1 × 10⁻⁶ is what a correct projection on the correct datum gives, and it is what a correct projection on a hundred-metre-wrong datum gives. Nothing in the fit distinguishes them.
This is the shape of every confounded measurement: an instrument that is insensitive to a nuisance variable is also blind to it, and blindness is only a virtue when nobody was going to ask.
Somebody does ask. A historical map’s datum is often the interesting question — which meridian was prime, which figure of the Earth was assumed, whose triangulation the coordinates came from — and an identification that reports a projection and a residual has said nothing about any of it. Worse, it has reported a residual small enough to look like a confirmation.
How the two explanations separate
They separate, but not by their size at any one region. They separate by how they scale with it.
| half-extent | datum shift | wrong projection | ratio |
|---|---|---|---|
| ±0.5° | 1.25 × 10⁻⁶ | 4.42 × 10⁻⁶ | 3.5 |
| ±1° | 1.14 × 10⁻⁶ | 1.59 × 10⁻⁵ | 13.9 |
| ±2° | 1.11 × 10⁻⁶ | 5.33 × 10⁻⁵ | 48.0 |
| ±4° | 1.11 × 10⁻⁶ | 2.01 × 10⁻⁴ | 180.8 |
| ±8° | 1.15 × 10⁻⁶ | 7.92 × 10⁻⁴ | 687.4 |
| ±16° | 1.36 × 10⁻⁶ | 3.16 × 10⁻³ | 2,326 |
A wrong projection’s residual grows because the two projections’ shapes diverge with extent — which is exactly what rung two of this ladder measured when it asked how large a region has to be before two candidates can be separated. A datum shift’s residual does not grow at all, because the similarity absorbs it at every scale.
So the diagnostic is a ladder rather than a number: fit the same map at several extents and look at the trend. A residual that is flat in the region size is a nuisance the fit is removing; one that grows is a shape difference. This is the same reasoning as reading a rate of convergence rather than an error, which this collection uses when a raster edge has an order — and it needs the same thing, which is measurements at more than one scale.
What the trend can be used for
If a flat residual means a nuisance the fit absorbed, then the flatness is itself a measurement, and it can be turned round.
Suppose a map’s projection is known — a national series says so on the sheet — and the fit still leaves a residual. Run the fit at several extents. If the residual is flat, the projection is right and the discrepancy is a datum: coordinates from an old triangulation, read as though they were modern. And the fit’s own translation is then an estimate of that datum shift, over that region, recovered from a printed map.
That is a measurement nobody appears to make, and it is available from material that survives in quantity. A nineteenth-century sheet with a graticule is a set of control points; its projection is documented; and the difference between where its graticule crossings sit and where a modern conversion puts them is the difference between two realisations of the ground.
The obstacle is the one the last section names: at the extent of a single sheet, half a degree or less, every explanation leaves the same residual. What makes it tractable is the same thing that makes any of it tractable — a series of sheets covering several degrees is one large region, and the trend across it is readable where the trend within one sheet is not.
The diagnostic has a rate, and the rate is quadratic
The table’s third column is quoted as a ratio at six sizes, and it is a law. Reading down it — 3.5, 13.9, 48.0, 180.8, 687.4, 2,326 — each entry is close to four times the one above, so the ratio grows as the square of the region’s half-extent.
That is not a coincidence of these two projections. Two maps that agree at a point differ at second order in the distance from it, which is the same r² law a local model has an order fits for a polynomial approximation and how small is flat enough fits for a tangent plane. A wrong candidate’s residual is that second-order difference; a datum shift’s is a constant the similarity cannot quite remove. Quadratic against flat gives quadratic.
Which turns below a degree nothing means anything into a graded rule with a number in it. Taking the measured 3.5 at half a degree and scaling as the square, the region needed for a stated discrimination ratio D is
A factor of ten needs about 0.85°. A factor of a hundred needs 2.7°. A factor of a thousand needs 8.5°. So a single sheet at 1:50,000 — a quarter of a degree or so — cannot reach even a factor of two, a county-sized region reaches ten, and only a region the size of a small country reaches a hundred.
The quadratic law also says something about how to spend effort. Reading the control points more carefully lowers the noise floor and does nothing to the ratio, because both explanations are measured against the same map width; extending the region raises the ratio as the square and is the only lever that moves the discrimination at all. A study with twice the sheets over twice the ground is four times better placed to separate the two explanations than one with four times the care over the same ground.
That is an unusual shape for a piece of advice about measurement, and it follows from the arithmetic rather than from any judgement about instruments: the quantity being separated is a shape difference, shape differences are second order, and second order rewards extent and is indifferent to precision.
Below a degree, nothing means anything
The bottom row of the table is the one to take away.
At ±0.5° the wrong projection’s residual is 4.4 × 10⁻⁶ and the datum shift’s is 1.3 × 10⁻⁶: a factor of 3.5, which no measurement precision available from a paper map can resolve. At that size the two explanations are the same number, and a third explanation — the map was drawn carelessly, or scanned with a slight nonlinearity, or the control points were read to the nearest half-millimetre — is the same number again.
This joins the two limits the ladder already carries. A small region cannot be identified at all, because the candidates have not yet diverged. A truth outside the library produces a confident wrong answer unless a rejection threshold is set. And now: a small residual on a small region has at least three explanations and the method distinguishes none of them.
What was computed, and how
The shifted observations are what an analyst actually holds. A map was drawn from the local datum’s coordinates, so the paper position of a point is the projection of its shifted latitude and longitude; the analyst then labels each position with the coordinate they believe it has, which is the unshifted one. The fit sees a set of paper positions and a set of geographic labels, and the labels are wrong by a datum shift.
Both experiments use the same graticule step, explicitly. The two generating functions defaulted their step differently — one at a sixth of the half-extent and one at a fifth — so leaving it out builds two different graticules, and the pairing of shifted against unshifted points silently matched only three of a hundred and sixty-nine. That is the kind of mistake that produces a figure rather than an error.
The “wrong projection” comparison is the conformal conic’s map fitted with an equal-area conic, which is the runner-up the identification actually offers on this region: 1.95 × 10⁻⁴ against the winner’s 1.11 × 10⁻⁶, a margin of 175.
Where the model stops
One datum and one region. OSGB36 over Britain is a shift of about a hundred metres that is very nearly a translation across a region that size. A datum whose difference from the reference is dominated by rotation — which is what a badly oriented old network gives, and what the seven parameters’ own rotations are for — would leave a residual the similarity absorbs less completely, and its trend with region size would not be flat. That case is not measured here and it is the one where the method might see something.
A similarity fit, not an affine one. The ladder’s own third rung showed that an affine fit cannot separate the cylindrical equal-area family, because the family differs by an affine map. An affine fit would absorb even more of a datum shift than a similarity does, which pushes the flat line lower and makes the confounding worse.
And the residual’s shape is not used. A datum shift’s residual is almost entirely first-order in the coordinates — a plane, explaining 95 to 99 per cent of its own variance at the smaller extents — and a wrong projection’s is not, at 63 per cent. That looks like a discriminant and it is not a reliable one: at ±16° the datum’s residual is 62 per cent planar, the projection’s 61, and the two are indistinguishable by that test exactly where they are most distinguishable by size.
Four limits, and the ladder they make
With this rung the identify ladder has four statements about a method that works, and they are worth listing together because they compose.
The region has to be large enough. Below a few degrees, two candidate projections have not diverged enough to be separated by any residual, however precisely the control points are read.
The truth has to be in the library, or a threshold has to refuse. Otherwise the method returns the nearest thing with a margin that looks like confidence.
A careless copy adds its own distortions on top of the projection’s, and the fit absorbs some of them and not others.
And a small residual has more than one explanation — a datum shift, a copying error, a measurement error and a correct answer all produce the same number below about a degree.
The four are not independent failures of four different kinds. They are the same failure seen four ways: the method measures a discrepancy and reports a name, and everything between the discrepancy and the name is an assumption. A ladder that has spent five rungs on one method has spent them finding out what those assumptions are, which is more useful than five rungs of improvements would have been.
The generalisation
A goodness-of-fit statistic answers the question “does this model fit”, and is routinely read as answering “is this model right”. They are different questions, and the gap between them is filled by whatever else could have produced the same residual.
The habit that closes the gap is not a better statistic. It is varying something the two explanations respond to differently — here the region size, which the projection difference grows with and the datum shift does not. That is a designed comparison rather than a computed one, and no amount of care with the fit produces it.
This is the third distinct limit this ladder has found in a method that works, and the three of them together are the ladder’s real subject. A method can be right, be confident, be checked against a held-out truth, and still be answering a question whose answer was decided by something it never looked at. The remedy in each case has been the same: find a knob the alternatives disagree about, and turn it.
One caution about the scaling law, since it is being used as a design rule. The quadratic is the leading behaviour of a difference between two smooth maps, and the measured ratios climb from 3.5 to 3.99 per doubling as the region grows — approaching four from below rather than sitting on it. So the rule understates the discrimination at large extents and overstates it at small ones, which is the safe direction for a design and the wrong one for a claim about a region already measured.
Who found it, and when
Identifying the projection of an undocumented map is a small and practical literature, mostly from the digital-humanities and historical-cartography side, and its standard method is the one this ladder implements: fit candidates, compare residuals. The best of it is careful about the reproduction transformation and about control point quality.
The confounding with the datum is, as far as this collection has found, not discussed in that literature, and there is a straightforward reason. Historical maps predate datums in the modern sense: a nineteenth-century map’s coordinates come from a national triangulation whose relationship to any modern frame is exactly what a datum transformation describes, and the practitioners fitting projections to such maps generally treat the whole discrepancy as one bundle and do not try to split it.
Splitting it is what would make the method say something new. A map whose projection is known and whose residual is flat in the region size is a map whose coordinates carry a datum shift, and the shift is recoverable from the fit’s own translation — which is a measurement of a historical triangulation, obtained from a printed map, and it is a use of this machinery that nobody appears to have made.
The same caution applies to the flat line, from the other direction. The datum’s residual is quoted as flat and it is not exactly flat — it dips to 1.11 × 10⁻⁶ in the middle of the range and rises at both ends, which is the similarity fit doing best where the region best matches a translation. That variation is a tenth of a per cent of the projection term at the largest extent and is the whole of the datum term at the smallest, which is one more reason the diagnostic needs the trend rather than any single reading.
Where the ladder goes next
This rung finds that the method’s residual is not about what it is read as being about. What the ladder has still never done is turn its own apparatus on a real map rather than a synthetic one — and the reason is that a real map’s control points come with a measurement error that every one of the four limits interacts with.
Named alongside this one
Essays reaching for the same objects. Nobody chose these; they are what the concept index makes visible.
- The sheet moved before it was measured control points · identification · residual · similarity transformation
- Where a fit leaves residuals datum · residual · similarity transformation
- A coordinate is the output of a solve datum · residual
- A local model has an order residual · similarity transformation
- The height a coordinate does not carry datum · similarity transformation
- The seven parameters, and what each one does datum · similarity transformation
What links here
Every essay whose body links to this one.
The objects this essay names
Each one links to every other essay that touches it.
ConfoundingControl pointsDatumDiagnosisEvidenceFitIdentificationProjectionResidualRobustnessScaleSimilarity transformation