A map does not say what it is
Every map on this site is one the site drew. The projection was chosen, its parameters were passed as arguments, and the distortion machinery was pointed at a function whose formula was in hand. That is the right arrangement for asking what a projection does, and it is not the situation any reader is ever in.
Every map anybody has actually used is the other kind: a picture, with the projection stated in a corner, in a legend, in a metadata field nobody filled in, or nowhere at all. A scanned sheet, a figure in a paper, a screenshot, a historical atlas plate — the geometry is right there in the ink and the statement about it is missing.
The statement is recoverable, and it needs nothing but the graticule.
What a drawn graticule is
A map that draws its own meridians and parallels is publishing a table of correspondences: this longitude and this latitude went to this place on the paper. A projection is exactly a rule for producing such a table.
So the inverse problem has an obvious shape. Run every candidate rule over the same longitudes and latitudes, ask which one lands where the picture does, and rank the answers. What makes it work rather than merely sound plausible is that the ranking is not close: the right candidate agrees to the arithmetic’s floor and the wrong ones do not agree at all.
The reproduction has to be removed first
A picture has been scaled to fit a page, possibly rotated, and its origin is wherever the crop put it. None of that is a property of the projection, so a candidate has to be scored after the best plane transformation between its own output and the observed points has been taken out.
That is a linear least-squares problem with no starting guess and no iteration. Writing the rotation and scale as a = s cos θ and b = s sin θ makes the four unknowns enter linearly:
so the normal equations are 4 × 4 and solve once. The residual left over is the shape difference, and dividing it by the map’s own width makes it a proportion rather than a length — because the physical size of the sheet is a fact about the paper.
Which plane transformations are allowed is the load-bearing decision in the whole exercise, and it is the subject of the third rung of this ladder. A similarity — uniform scale, rotation, translation — is what a photocopier and a crop do. Anything more absorbs part of the projection.
The parameters are searched, the plane part is solved
A candidate is not a projection but a projection with parameters: a conic has a cone constant, an azimuthal has a centre, an equirectangular has a standard parallel. Those enter the formulae nonlinearly, so they are searched — a coarse grid, then golden-section refinement on each in turn — while the plane transformation is solved exactly inside every evaluation.
The recovered parameters are part of the answer and they are worth checking. Fitted to a map drawn as a conformal conic centred at 45° north, the search returns a cone constant of 0.70711, and sin 45° = 0.70711: the cone constant of a tangent conic at its own standard parallel. Nothing in the fit was told the standard parallel or the latitude of the region.
What the residual looks like when it is wrong
The distinction between a wrong candidate and a right candidate measured badly is the one thing a ranking alone cannot make, and the residual’s own shape makes it.
A wrong projection leaves a residual that is smooth and structured — it varies slowly across the map, because it is the difference between two smooth functions. Measurement error leaves a residual that is not: it varies from crossing to crossing. So the picture above is not decoration. A structured residual says the candidate is wrong; a scattered one says the reading of the map is.
Nothing on this site has an error model in it — the crossings here are computed exactly, which is the boundary this collection agreed to when it opened its field about surveys — so every residual in these figures is pure shape difference, and the structure is total.
The margin is the answer’s strength
The ranking’s first place is not the whole result. The ratio between the first and second residuals is what says whether the map has been identified:
- a conformal conic map: right candidate at 1.3 × 10⁻¹⁰, next at 5.0 × 10⁻³ — a margin of 3.7 × 10⁷;
- a Mollweide map: right candidate at 1.2 × 10⁻¹⁵, next at 1.3 × 10⁻², a margin of 10¹³.
A margin of ten million is an identification, in the sense measuring instead of naming uses throughout: the number decides, and the name follows from it. A margin of two is a coincidence with two candidates in it, and it happens: over a small enough region, or between two projections that are close relatives, the fit cannot separate them at all. That is the second rung of this ladder, and it is not a defect of the method but a property of the pair.
What was computed, and how
Twenty candidates, each with up to one free parameter, fitted to between 25 and 4,225 crossings.
The observations are generated by drawing a map with a stated projection and a stated reproduction — a scale of 137.3, a rotation of 11° and an offset — so that the fit’s recovery of those three can be checked as well as its ranking. That is deliberate: a test in which the reproduction is the identity would pass with a fit that ignored the reproduction entirely.
The search over each candidate’s own parameter is a grid of thirteen points followed by two passes of golden section, because the residual surfaces are not unimodal. A conic’s residual against a cylindrical map falls all the way to the n → 0 end of its range, and a refinement started from the wrong basin would report a conic that fits a Mercator map badly rather than one that fits it as well as the family permits.
The one candidate class that needs care is the azimuthals, whose centre is a parameter of the projection rather than a rotation of the sphere. Building them through the general aspect machinery — rotating the sphere so that the projection’s pole lands on the region — puts the map somewhere else entirely, and this collection has made that mistake once already in a different figure and found it by getting an answer that was too dramatic to believe.
What the fit recovers besides the name
A ranking answers which projection. The fit answers three more questions at the same time, and each has a use.
The projection’s own parameters. The cone constant, the standard parallel, the centre of an azimuthal — recovered as the values that minimise the residual, and interpretable directly: 0.70711 is the tangent conic at 45°, and a recovered standard parallel of 30° on a cylindrical equal-area map identifies it as Behrmann’s rather than Lambert’s.
The reproduction. Scale, rotation and offset come out of the same solve, which is what a georeferencing tool needs in order to place the sheet: the scale gives the map’s representative fraction once the sheet’s own size is known, and the rotation says how far the print is off square.
And the residual itself, which is the only one of the four that says whether to believe the other three.
The question a machine asks of every coordinate
Recovering a projection is one answer to a general difficulty this collection states in several places. A coordinate without its system is not a location; a projected coordinate declares five things and four of them are invisible in the numbers; the units are part of the coordinate and are not in it either.
A picture is the same problem with a different missing declaration. The projection is one of the five, and it is the only one the geometry itself can supply — the datum, the units and the axis order leave no trace in the shape of a graticule, because they are conventions about numbers rather than about pictures.
So an identification from a map is exactly as partial as the geometry allows. The projection can be recovered to ten decimal places and the sheet still cannot be georeferenced, because the datum it is on moved the ground under it by hundreds of metres before the projection was applied.
That is the honest boundary of this ladder, and it is worth stating at its first rung rather than its last. What is being recovered here is the rule that drew the picture, and the rule is not the whole of what a map means — which is the same distinction measuring instead of naming makes about a projection’s properties, one level down.
Where the model stops
The graticule has to be drawn and labelled. A map with no graticule publishes no correspondences, and the same method then needs identifiable features instead — coastline vertices, city positions, grid ticks — which brings back everything this site’s founding decision avoids about datasets and their simplification levels.
The candidate has to be in the library. A fit ranks the twenty candidates it knows; a map drawn in the twenty-first returns a best fit that is merely the least bad. The margin is what protects against this: a genuine identification has a margin of millions, and a map in an unlisted projection produces a ranking whose top few are within a factor of two of one another.
A composite is not a candidate. An interrupted projection, a map assembled from several sheets, or a projection with a locally applied adjustment is not any single rule, and the residual will say so without saying what to do about it.
And the crossings are exact. A real reading of a printed sheet has a position error of a fraction of a millimetre and a paper distortion of its own, which sets a floor under every residual. Nothing here estimates that floor, and estimating it would be the first thing a working implementation needed.
What sets the floor, and it is not the reading
The measurement’s boundary is stated above as the crossings are exact and the working floor is left unestimated. It is worth estimating, because the answer changes which of the two obvious error sources matters and it explains why the third rung of this ladder is about the plane transformation.
Reading error averages away. A crossing read off a printed sheet is wrong by a fraction of a millimetre — say 0.2 mm on a 300-millimetre sheet, which is 0.07 per cent of the map’s width. That error is independent from crossing to crossing, so its contribution to a fitted residual falls as one over the square root of the number of crossings: at 120 crossings it is under 0.01 per cent. Against the 0.5 per cent shape difference between a conformal and an equal-area conic, reading error is nowhere near the limit and adding crossings makes it less so.
Paper distortion does not. A printed sheet shrinks with humidity, and it shrinks anisotropically — more across the grain than along it, by a few tenths of a per cent on ordinary map paper. That is a differential scale between two directions, it is smooth across the sheet, and it is therefore indistinguishable in kind from what a wrong projection leaves behind. Averaging over crossings does nothing to it, because it is not noise.
The sizes are the uncomfortable part. A differential shrinkage of 0.2 per cent leaves a residual of the same order as the 0.5 per cent that separates two genuinely different conics, and half the 1.5 per cent that separates Mollweide from Eckert IV. So the floor under a real identification is set by the paper, at a level comparable with the signal being measured, and it is a floor that no amount of care in reading the sheet lowers.
That is exactly why the choice of plane transformation is load-bearing. A similarity has one scale, so it cannot absorb differential shrinkage and the shrinkage stays in the residual. An affine transformation has two scales and a shear, so it absorbs the shrinkage completely — and it also absorbs part of every projection’s own anisotropy, which is a large part of what distinguishes an equal-area conic from a conformal one.
The two failures are therefore a genuine dilemma rather than a choice with a right answer: fit a similarity and the paper is in the residual, fit an affine and the projection is in the transformation. Neither the exact crossings of this rung nor the ranking margin it reports survives contact with a real sheet, and the margin of thirty-seven million is a measurement of how much room there is to lose.
Why the wrong candidates fail so badly
It is worth asking why the ranking is so decisive, because a margin of thirty-seven million is not what a fit usually gives.
The reason is that a projection is a rigid rule with almost no freedom in it. A candidate has at most one parameter to move; the plane transformation has four; and everything else about the map — every one of the hundred and twenty crossings — is then determined. Fitting five numbers to two hundred and forty is a heavily over-determined problem, and an over-determined fit with the wrong model has nowhere to hide.
Compare the usual situation in curve fitting, where a model with several free parameters can absorb a good deal of the wrong shape before its residual rises. Here the model cannot bend. Every projection is a formula with two or three constants and no local adjustment, which is exactly the property that makes the site’s own machinery able to assert things about them — and it is the same property that makes them identifiable.
Rigidity is why this works, and it is also why the method fails in the one case the next rung is about: two projections that are rigid in the same way, differing only at an order the fit’s own freedom conceals.
The generalisation
The structure is the standard one for inverse problems with a nuisance parameter, and this instance is unusually clean: the parameters that describe the reproduction enter linearly and are solved exactly, while the parameters that describe the projection enter nonlinearly and are searched. Separating the two is what makes the search small enough to do exhaustively, and the separation is available because a reproduction is an affine map and a projection is not.
The habit worth carrying is the margin. A ranking of candidates by a fitted residual always has a first place, and a first place means nothing until the second is beside it. That is true of model selection everywhere, and it is unusually easy to check here because the right answer, when it is present, wins by seven orders of magnitude rather than by a comfortable margin.
Who found it, and when
Identifying the projection of an undocumented map is a working problem in cartographic history and in georeferencing, and there is a small literature on it — most of it fitting a handful of candidates to control points and reporting a root-mean-square error, which is this method with a shorter list. The systematic version, with every candidate’s parameters searched and the plane transformation solved rather than assumed, belongs to the software era.
What is old is the difficulty. A nineteenth-century atlas plate typically names its projection; a sixteenth-century one does not, and the question of what Mercator’s contemporaries were actually drawing has been argued from the graticule’s own geometry since at least the 1930s.
Where the ladder goes next
The margin in this rung is thirty-seven million, and that is a fact about the region as much as about the projections: it was measured over forty degrees of latitude. Shrink the region and the margin shrinks with it, until two candidates are the same picture and no measurement can separate them. Two projections that cannot be told apart computes the extent at which each pair crosses that line, and finds that what decides it is the order at which the two projections first differ.
What this makes readable
Essays that name this one as a prerequisite.
Named alongside this one
Essays reaching for the same objects. Nobody chose these; they are what the concept index makes visible.
- What a careless copy hides least-squares · projection identification · residual · similarity transformation
- A condition imposed at points is not a condition conformality · least-squares · residual
- A local model has an order least-squares · residual · similarity transformation
- The answer is a set inverse problem · least-squares · residual
- The nearest map to an impossible request conformality · least-squares · residual
- The nodes were evenly spaced conformality · least-squares · residual
What links here
The 8 essays that link to this one and share the most of its objects, of 13 that link here.
The objects this essay names
Each one links to every other essay that touches it.
CandidateCone constantConformalityGraticuleInverse problemLeast-squaresMap metadataParameter searchProjection identificationRankingReproductionResidualSimilarity transformation