What is taught wrongly

A map does not say what it is

Every map on this site is one the site drew, from a projection it chose. Every map a reader has ever used is the other kind — a picture whose projection is a sentence in a corner, a legend, or nothing at all. The graticule is enough to recover it: fit every candidate to the crossings and rank what is left over.

Every map on this site is one the site drew. The projection was chosen, its parameters were passed as arguments, and the distortion machinery was pointed at a function whose formula was in hand. That is the right arrangement for asking what a projection does, and it is not the situation any reader is ever in.

Every map anybody has actually used is the other kind: a picture, with the projection stated in a corner, in a legend, in a metadata field nobody filled in, or nowhere at all. A scanned sheet, a figure in a paper, a screenshot, a historical atlas plate — the geometry is right there in the ink and the statement about it is missing.

The statement is recoverable, and it needs nothing but the graticule.

Every candidate fitted to one map's graticule, ranked by what is left over. The map is drawn in Conformal conic over a region 40° tall centred at 45° north, and the projection is not told to the fit. Each candidate is evaluated at the same 121 graticule crossings, its own parameters are searched, the best plane similarity between its output and the picture is removed, and what remains is drawn as a proportion of the map's own width on a logarithmic scale. Conformal conic fits to 1.3e-10, which is the arithmetic's floor; the next candidate is 3.7e+7 times worse.
Fig. 1 Every candidate in the library fitted to one map’s graticule crossings and ranked by what is left over. The map is a conformal conic and the fit is not told so; the right candidate lands at 10⁻¹⁰ of the map’s own width, which is the arithmetic’s floor, and the next is thirty-seven million times worse.

What a drawn graticule is

A map that draws its own meridians and parallels is publishing a table of correspondences: this longitude and this latitude went to this place on the paper. A projection is exactly a rule for producing such a table.

So the inverse problem has an obvious shape. Run every candidate rule over the same longitudes and latitudes, ask which one lands where the picture does, and rank the answers. What makes it work rather than merely sound plausible is that the ranking is not close: the right candidate agrees to the arithmetic’s floor and the wrong ones do not agree at all.

The reproduction has to be removed first

A picture has been scaled to fit a page, possibly rotated, and its origin is wherever the crop put it. None of that is a property of the projection, so a candidate has to be scored after the best plane transformation between its own output and the observed points has been taken out.

That is a linear least-squares problem with no starting guess and no iteration. Writing the rotation and scale as a = s cos θ and b = s sin θ makes the four unknowns enter linearly:

x=aXbY+tx,y=bX+aY+tyx = aX - bY + t_x, \qquad y = bX + aY + t_y

so the normal equations are 4 × 4 and solve once. The residual left over is the shape difference, and dividing it by the map’s own width makes it a proportion rather than a length — because the physical size of the sheet is a fact about the paper.

Which plane transformations are allowed is the load-bearing decision in the whole exercise, and it is the subject of the third rung of this ladder. A similarity — uniform scale, rotation, translation — is what a photocopier and a crop do. Anything more absorbs part of the projection.

The parameters are searched, the plane part is solved

A candidate is not a projection but a projection with parameters: a conic has a cone constant, an azimuthal has a centre, an equirectangular has a standard parallel. Those enter the formulae nonlinearly, so they are searched — a coarse grid, then golden-section refinement on each in turn — while the plane transformation is solved exactly inside every evaluation.

The recovered parameters are part of the answer and they are worth checking. Fitted to a map drawn as a conformal conic centred at 45° north, the search returns a cone constant of 0.70711, and sin 45° = 0.70711: the cone constant of a tangent conic at its own standard parallel. Nothing in the fit was told the standard parallel or the latitude of the region.

The published graticule, and the best that Equal-area conic can do with it. The map's own meridians and parallels at 4° spacing, drawn from a Conformal conic projection over 40° of latitude centred at 45°, with Equal-area conic fitted to the same crossings and drawn over them. The fit has removed the best scale, rotation and offset, so what is visible is the shape difference alone: 0.498 per cent of the map's own width, root-mean-square. Drawn in the picture's own coordinates rather than in any projection of this site's, because a scanned map has no projection until one has been identified.
Fig. 2 The published graticule with the runner-up fitted to it and drawn over it. The equal-area conic has had the best scale, rotation and offset removed, so what is visible is the shape difference alone — half a per cent of the map’s own width, which on a printed sheet is a millimetre and a half and is exactly the kind of thing an eye does not catch.

What the residual looks like when it is wrong

What is left over, at every crossing, exaggerated 24×. The residual of the best Equal-area conic fit to a Conformal conic graticule, drawn as a displacement at each of the 121 crossings and multiplied by 24 to be visible at all. The largest is 0.989 per cent of the map's width. What matters is that the residual is not noise: it is a smooth pattern, which is what a shape difference looks like and what distinguishes the wrong projection from the right one measured badly.
Fig. 3 The residual at each crossing, drawn as a displacement and multiplied to be visible at all. It is not noise: it is a smooth pattern, which is what a shape difference looks like — and it is the check that separates a wrong projection from a right one measured badly.

The distinction between a wrong candidate and a right candidate measured badly is the one thing a ranking alone cannot make, and the residual’s own shape makes it.

A wrong projection leaves a residual that is smooth and structured — it varies slowly across the map, because it is the difference between two smooth functions. Measurement error leaves a residual that is not: it varies from crossing to crossing. So the picture above is not decoration. A structured residual says the candidate is wrong; a scattered one says the reading of the map is.

Nothing on this site has an error model in it — the crossings here are computed exactly, which is the boundary this collection agreed to when it opened its field about surveys — so every residual in these figures is pure shape difference, and the structure is total.

The margin is the answer’s strength

The ranking’s first place is not the whole result. The ratio between the first and second residuals is what says whether the map has been identified:

  • a conformal conic map: right candidate at 1.3 × 10⁻¹⁰, next at 5.0 × 10⁻³ — a margin of 3.7 × 10⁷;
  • a Mollweide map: right candidate at 1.2 × 10⁻¹⁵, next at 1.3 × 10⁻², a margin of 10¹³.

A margin of ten million is an identification, in the sense measuring instead of naming uses throughout: the number decides, and the name follows from it. A margin of two is a coincidence with two candidates in it, and it happens: over a small enough region, or between two projections that are close relatives, the fit cannot separate them at all. That is the second rung of this ladder, and it is not a defect of the method but a property of the pair.

Drawing the graticule more finely does not change the answer. The residual of a Stereographic fit to a Mercator map of half-extent 20°, computed at five graticule spacings from 10° down to 0.625° — from 25 crossings to 4225. Each halving of the step closes about half of what remains, which is first-order convergence to 4.454e-3: a property of the region rather than of how the map was drawn. The dashed line is that limit, obtained from the two finest spacings.
Fig. 4 The residual of one wrong candidate at five graticule spacings, from 25 crossings to 4,225. Each halving of the step closes about half of what remains, converging to a limit that is a property of the region rather than of how finely the map was drawn.

What was computed, and how

Twenty candidates, each with up to one free parameter, fitted to between 25 and 4,225 crossings.

The observations are generated by drawing a map with a stated projection and a stated reproduction — a scale of 137.3, a rotation of 11° and an offset — so that the fit’s recovery of those three can be checked as well as its ranking. That is deliberate: a test in which the reproduction is the identity would pass with a fit that ignored the reproduction entirely.

The search over each candidate’s own parameter is a grid of thirteen points followed by two passes of golden section, because the residual surfaces are not unimodal. A conic’s residual against a cylindrical map falls all the way to the n → 0 end of its range, and a refinement started from the wrong basin would report a conic that fits a Mercator map badly rather than one that fits it as well as the family permits.

The one candidate class that needs care is the azimuthals, whose centre is a parameter of the projection rather than a rotation of the sphere. Building them through the general aspect machinery — rotating the sphere so that the projection’s pole lands on the region — puts the map somewhere else entirely, and this collection has made that mistake once already in a different figure and found it by getting an answer that was too dramatic to believe.

What the fit recovers besides the name

The published graticule, and the best that Eckert IV can do with it. The map's own meridians and parallels at 12° spacing, drawn from a Mollweide projection over 120° of latitude centred at 0°, with Eckert IV fitted to the same crossings and drawn over them. The fit has removed the best scale, rotation and offset, so what is visible is the shape difference alone: 1.292 per cent of the map's own width, root-mean-square. Drawn in the picture's own coordinates rather than in any projection of this site's, because a scanned map has no projection until one has been identified.
Fig. 5 A Mollweide graticule with Eckert IV fitted to it — two pseudocylindrical equal-area projections that differ in what they do at the pole. The fit has removed the best scale, rotation and offset, so the visible gap is the shape difference: one and a half per cent of the map’s own width, which is unmissable on a world map and would be a millimetre on a page.

A ranking answers which projection. The fit answers three more questions at the same time, and each has a use.

The projection’s own parameters. The cone constant, the standard parallel, the centre of an azimuthal — recovered as the values that minimise the residual, and interpretable directly: 0.70711 is the tangent conic at 45°, and a recovered standard parallel of 30° on a cylindrical equal-area map identifies it as Behrmann’s rather than Lambert’s.

The reproduction. Scale, rotation and offset come out of the same solve, which is what a georeferencing tool needs in order to place the sheet: the scale gives the map’s representative fraction once the sheet’s own size is known, and the rotation says how far the print is off square.

And the residual itself, which is the only one of the four that says whether to believe the other three.

The question a machine asks of every coordinate

54.5° north, 2° west: four places one pair of numbers could be. The same two numbers, read as three datums and as the pair in the other order. The centre is the WGS84 reading; the marks are where the ground is if the numbers were measured against OSGB36, ED50, NAD27 instead, at 99 m, 134 m, 195 m. Every one of those is a plausible position — near enough to look like a survey difference and far enough to be a different field. Reversing the two numbers instead moves the point 8111 kilometres, and is the mistake nobody worries about because it is usually obvious.
Fig. 6 One latitude and longitude read on three datums: the same pair of numbers, three places on the ground, hundreds of metres apart. Identifying a projection from a picture answers one of the five declarations a coordinate needs, and this figure is about a different one.

Recovering a projection is one answer to a general difficulty this collection states in several places. A coordinate without its system is not a location; a projected coordinate declares five things and four of them are invisible in the numbers; the units are part of the coordinate and are not in it either.

A picture is the same problem with a different missing declaration. The projection is one of the five, and it is the only one the geometry itself can supply — the datum, the units and the axis order leave no trace in the shape of a graticule, because they are conventions about numbers rather than about pictures.

So an identification from a map is exactly as partial as the geometry allows. The projection can be recovered to ten decimal places and the sheet still cannot be georeferenced, because the datum it is on moved the ground under it by hundreds of metres before the projection was applied.

That is the honest boundary of this ladder, and it is worth stating at its first rung rather than its last. What is being recovered here is the rule that drew the picture, and the rule is not the whole of what a map means — which is the same distinction measuring instead of naming makes about a projection’s properties, one level down.

Where the model stops

The graticule has to be drawn and labelled. A map with no graticule publishes no correspondences, and the same method then needs identifiable features instead — coastline vertices, city positions, grid ticks — which brings back everything this site’s founding decision avoids about datasets and their simplification levels.

The candidate has to be in the library. A fit ranks the twenty candidates it knows; a map drawn in the twenty-first returns a best fit that is merely the least bad. The margin is what protects against this: a genuine identification has a margin of millions, and a map in an unlisted projection produces a ranking whose top few are within a factor of two of one another.

A composite is not a candidate. An interrupted projection, a map assembled from several sheets, or a projection with a locally applied adjustment is not any single rule, and the residual will say so without saying what to do about it.

And the crossings are exact. A real reading of a printed sheet has a position error of a fraction of a millimetre and a paper distortion of its own, which sets a floor under every residual. Nothing here estimates that floor, and estimating it would be the first thing a working implementation needed.

What sets the floor, and it is not the reading

The measurement’s boundary is stated above as the crossings are exact and the working floor is left unestimated. It is worth estimating, because the answer changes which of the two obvious error sources matters and it explains why the third rung of this ladder is about the plane transformation.

Reading error averages away. A crossing read off a printed sheet is wrong by a fraction of a millimetre — say 0.2 mm on a 300-millimetre sheet, which is 0.07 per cent of the map’s width. That error is independent from crossing to crossing, so its contribution to a fitted residual falls as one over the square root of the number of crossings: at 120 crossings it is under 0.01 per cent. Against the 0.5 per cent shape difference between a conformal and an equal-area conic, reading error is nowhere near the limit and adding crossings makes it less so.

Paper distortion does not. A printed sheet shrinks with humidity, and it shrinks anisotropically — more across the grain than along it, by a few tenths of a per cent on ordinary map paper. That is a differential scale between two directions, it is smooth across the sheet, and it is therefore indistinguishable in kind from what a wrong projection leaves behind. Averaging over crossings does nothing to it, because it is not noise.

The sizes are the uncomfortable part. A differential shrinkage of 0.2 per cent leaves a residual of the same order as the 0.5 per cent that separates two genuinely different conics, and half the 1.5 per cent that separates Mollweide from Eckert IV. So the floor under a real identification is set by the paper, at a level comparable with the signal being measured, and it is a floor that no amount of care in reading the sheet lowers.

That is exactly why the choice of plane transformation is load-bearing. A similarity has one scale, so it cannot absorb differential shrinkage and the shrinkage stays in the residual. An affine transformation has two scales and a shear, so it absorbs the shrinkage completely — and it also absorbs part of every projection’s own anisotropy, which is a large part of what distinguishes an equal-area conic from a conformal one.

The two failures are therefore a genuine dilemma rather than a choice with a right answer: fit a similarity and the paper is in the residual, fit an affine and the projection is in the transformation. Neither the exact crossings of this rung nor the ranking margin it reports survives contact with a real sheet, and the margin of thirty-seven million is a measurement of how much room there is to lose.

Why the wrong candidates fail so badly

It is worth asking why the ranking is so decisive, because a margin of thirty-seven million is not what a fit usually gives.

The reason is that a projection is a rigid rule with almost no freedom in it. A candidate has at most one parameter to move; the plane transformation has four; and everything else about the map — every one of the hundred and twenty crossings — is then determined. Fitting five numbers to two hundred and forty is a heavily over-determined problem, and an over-determined fit with the wrong model has nowhere to hide.

Compare the usual situation in curve fitting, where a model with several free parameters can absorb a good deal of the wrong shape before its residual rises. Here the model cannot bend. Every projection is a formula with two or three constants and no local adjustment, which is exactly the property that makes the site’s own machinery able to assert things about them — and it is the same property that makes them identifiable.

Rigidity is why this works, and it is also why the method fails in the one case the next rung is about: two projections that are rigid in the same way, differing only at an order the fit’s own freedom conceals.

The generalisation

The structure is the standard one for inverse problems with a nuisance parameter, and this instance is unusually clean: the parameters that describe the reproduction enter linearly and are solved exactly, while the parameters that describe the projection enter nonlinearly and are searched. Separating the two is what makes the search small enough to do exhaustively, and the separation is available because a reproduction is an affine map and a projection is not.

The habit worth carrying is the margin. A ranking of candidates by a fitted residual always has a first place, and a first place means nothing until the second is beside it. That is true of model selection everywhere, and it is unusually easy to check here because the right answer, when it is present, wins by seven orders of magnitude rather than by a comfortable margin.

Who found it, and when

Identifying the projection of an undocumented map is a working problem in cartographic history and in georeferencing, and there is a small literature on it — most of it fitting a handful of candidates to control points and reporting a root-mean-square error, which is this method with a shorter list. The systematic version, with every candidate’s parameters searched and the plane transformation solved rather than assumed, belongs to the software era.

What is old is the difficulty. A nineteenth-century atlas plate typically names its projection; a sixteenth-century one does not, and the question of what Mercator’s contemporaries were actually drawing has been argued from the graticule’s own geometry since at least the 1930s.

Where the ladder goes next

The margin in this rung is thirty-seven million, and that is a fact about the region as much as about the projections: it was measured over forty degrees of latitude. Shrink the region and the margin shrinks with it, until two candidates are the same picture and no measurement can separate them. Two projections that cannot be told apart computes the extent at which each pair crosses that line, and finds that what decides it is the order at which the two projections first differ.

What this makes readable

Essays that name this one as a prerequisite.

Named alongside this one

Essays reaching for the same objects. Nobody chose these; they are what the concept index makes visible.

What links here

The 8 essays that link to this one and share the most of its objects, of 13 that link here.

The objects this essay names

Each one links to every other essay that touches it.

CandidateCone constantConformalityGraticuleInverse problemLeast-squaresMap metadataParameter searchProjection identificationRankingReproductionResidualSimilarity transformation