The rule scored out of sample
The taught rule for choosing a projection family — cylindrical near the equator, conic in the middle latitudes, azimuthal at the poles — was scored here against thirty regions and found right on nineteen of them. Every one of its eleven failures had the same shape: a compact region, where the azimuthal wins at every latitude from the equator to 80°.
A replacement fell out of that: azimuthal for a compact region, conic for a wide one, transverse cylindrical for a tall one. It scores better on those thirty, and this collection wrote down, at the time, that scoring a rule on the population it was read from is not evidence about any other population. The same objection applies to every claim in this field that begins in practice — including, as conformal does not mean the angles are right found, the ones stated as properties.
Paying that debt turns out to produce two findings, and only one of them is the one the debt was for.
The test that was asked for, and passes
The held-out population is a different grid: latitudes at 10°, 30°, 50°, 70° and 85° against the training set’s 0°, 20°, 40°, 60° and 80°; heights of 4°, 18° and 45° against 10° and 30°; and aspect ratios of 0.15, 0.7 and 6 against 0.4, 1 and 2.5. Forty-five regions, none of them in the training set, several of them at shapes more extreme than any it contained.
The shape rule scores 41 of 45, which is 91 per cent, against the 25 of 30 — 83 per cent — it managed on its own population.
It went up. The expectation on record was that a rule fitted to a sample degrades on a fresh one, and here it does not, which is worth being precise about: nothing has been fitted in the statistical sense. The rule has two thresholds and three outcomes, it was read off a pattern rather than optimised against a loss, and a rule that simple has very little capacity to memorise thirty examples. Its failure mode is not overfitting.
The taught rule, on the same held-out set, falls from 63 per cent to 51.
The failures that remain, and what they are
Four of the forty-five. All four are at 85° north, and they are:
| region | aspect | rule said | measured winner | cost |
|---|---|---|---|---|
| 4° tall at 85° | 0.70 | cylindrical | azimuthal | 1.63× |
| 4° tall at 85° | 6.00 | conic | azimuthal | 7.04× |
| 18° tall at 85° | 1.68 | conic | azimuthal | 9.21× |
| 18° tall at 85° | 14.37 | conic | azimuthal | 2.53× |
Every one of them wants the azimuthal, and the rule declines to name it because the region is not compact.
That is the taught rule’s own clause. Azimuthal at the poles is the one part of the classical advice the shape rule threw away, and it is the one part that is right for a reason the shape rule cannot see: near a pole every region is compact in the geometry that matters, whatever its bounding box says, because the meridians converge. A box 14 times wider than it is tall at 85° north is a wedge, and a wedge with its apex on the pole is exactly what an azimuthal projection is built for.
So the obvious repair is to put the discarded clause back: azimuthal above 80°, whatever the shape; below that, azimuthal if compact, conic if wide, transverse cylindrical if tall. Measured rather than assumed, that clause repairs all four failures and breaks two regions that were right, for a net 43 of 45.
The two it breaks are the narrow slivers — a region 0.15 as wide as it is tall at 85°, and one at 0.36 — which are won by the transverse cylindrical, and where following the polar clause costs 10.11 times. A tall narrow region at 85° runs down away from the pole, so it is not a wedge at all, and the classical clause is wrong about it in exactly the way the shape clause is wrong about a wide one.
Neither clause is right on its own and their combination is not right either. What that says is not that a third threshold is needed; it is that both rules are reading the wrong number.
The test that was not asked for, and fails
Two of seven, for both rules, and the costs are not marginal:
- Europe — the case every textbook uses to illustrate conic in the middle latitudes — is won by the azimuthal, and the taught rule’s advice costs 3.12 times. The shape rule gets this one right, because Europe’s bounding box is 0.87 wide against tall.
- New Zealand is won by the azimuthal and both rules say conic, at a cost of 92.94 times.
- Japan is won by the azimuthal and both rules say conic, at 3.25 times.
- Britain is won by the azimuthal; the taught rule says conic at 7.44 times and the shape rule says cylindrical at 1.15.
Five of the seven are won by the azimuthal. Neither rule names it more than twice.
Why the real regions break it
The two worst failures are Japan and New Zealand, and they have something in common that no region in either synthetic population has: their long axis is oblique. Japan’s runs 42° from north; New Zealand’s runs north-east. Both are stored here as ellipses at an azimuth rather than as boxes.
Both rules read a region’s shape from its bounding box in longitude and latitude. For a region whose long axis is diagonal, that box is very much larger than the region and has a completely different aspect ratio — New Zealand’s box is 2.18 wide against tall while the region inside it is nearly four times as long as it is broad, lying across the box’s diagonal.
This is not a new fact about projections. It is the third parameter of an aspect arriving from the other direction: a region whose extent is oblique needs a projection turned to match it, the turn is worth up to 2.11 times when it is searched for, and a family rule that never mentions orientation cannot express the answer at all.
The other five named regions are boxes, and the rules do better on them — three of five for the shape rule against the taught rule’s two. The tropics is the one genuine disagreement between the rules on a box, and the taught rule wins it: a band 7.66 times wider than it is tall at the equator is exactly the cylindrical case, and the shape rule’s conic for a wide region clause is wrong there for the same reason its polar failures are wrong. Wide near the equator and wide near the pole are different situations.
What a rule can be, after this
Three conclusions, in increasing order of how much they cost to accept. Every projection minimises something, and a family rule is an attempt to guess which minimisation will win without performing any of them.
A two-clause rule keyed on one number is nearly good enough for boxes. Patched with the polar clause it gets 43 of 45 and 25 of 30 on built regions. If the region really is a box and really is at a modest latitude, the shape rule is sound advice and the taught rule is not.
No rule keyed on a bounding box can handle an oblique region, and real regions are frequently oblique. The failure is not in the thresholds; it is in the summary statistic, and no adjustment to the thresholds fixes a 93.
And the cost of the advice is not what the hit rate suggests. On the named regions the shape rule names the winner twice in seven and its mean cost is 14.6 — but five of those seven cost less than 1.6 times, and one costs 93. A rule’s average is dominated by the case it cannot see, which is the general shape of what happens when a heuristic meets a population it was not read from.
The alternative is not a better rule. It is running the measurement, which takes a second per region and answers the question the rule is a substitute for. That is the same conclusion which projection is best reached from the other end, and designing a grid for one region reached again for a country that had to commit to one answer for a century.
What was computed, and how
Each region is scored by giving every family its best parameters before comparing them: the conic its cone constant searched across the range, the cylindrical a standard parallel and the choice between a normal and a transverse axis, the azimuthal its centre. Scoring a family with parameters chosen for somewhere else measures the parameters rather than the family, and is the way this comparison is usually got wrong.
The criterion is Kavrayskiy’s — the root mean square of the logarithms of the two principal scale factors over the region — with each candidate normalised for overall scale first, because the scale is free and is not what is being compared.
The three populations are generated rather than curated. The trained set is the five latitudes, two heights and three shapes the earlier rung used; the held-out set is five different latitudes, three different heights and three different shapes; the named set is every region in the library except world and hemisphere, which are degenerate cases rather than places.
The assertion that was written first is the one that failed. It required the rule to score lower out of sample, on the reasonable grounds that a rule read off a sample usually does, and the measurement refused it at 91 against 83. The assertion now standing says four things: the shape rule beats the taught rule in sample; it does not degrade on a second population of the same kind; both collapse below half on the named regions; and every held-out failure is polar. The fourth is the one that turned a null result into a finding.
Why this kind of test is worth running at all
The reason to score a rule of thumb is not to embarrass it. A rule that is right two thirds of the time on built regions is a useful thing to carry in one’s head, and nothing here suggests teaching something else instead.
The reason is that a rule with no measured hit rate cannot be improved. Until it is scored there is nothing to compare a replacement against, no way to tell whether a proposed clause helps, and no way to notice that its failures all have the same shape — which was the finding that produced the shape rule in the first place, and which fell straight out of the scoring rather than out of any insight about projections.
The same applies to the shape rule now. It has a measured rate on three populations, a known failure mode, and a patch that helps on two of the four cases it was meant for and hurts on two others. That is a great deal more than the taught rule had after a century.
Where the model stops
Seven named regions is a small population and its smallness is a limit rather than a defect — seven is what the library holds. Two of the seven are ellipses at an azimuth and two of the seven are the regions the failure is about, so the 29 per cent is a rate over a handful and should be read as both rules fail on the oblique ones rather than as a percentage.
The synthetic populations are all boxes centred on their own meridian. That is what makes them easy, and it is the reason the held-out test passed: a box at a grid latitude is the case both rules were designed around, and generating forty-five more of them tests the thresholds without testing the summary statistic. A held-out sample from the same generator tests the parameters of a rule and not its shape, which is the methodological point this rung is really about.
And the whole exercise assumes the family is the choice being made. In practice the choice is a named projection with named parameters, and the family is a way of talking about it. A rule that names the right family and the wrong member of it has not helped anybody.
Who found it, and when
The rule is older than any citation for it and appears in essentially the same words in Deetz and Adams’ Elements of Map Projection of 1921, in Steers, in Maling, and in every introductory course since. Its origin is the physical construction — a cylinder, a cone and a plane wrapped round a globe — and the latitudes in it are the latitudes at which each surface touches.
Its persistence is the interesting part. A rule that is right 63 per cent of the time on built regions and 29 per cent on real ones has survived a century of teaching because it is memorable, because it comes with a picture, and because until a projection library and a scoring criterion could be run over a population nobody could say what its hit rate was.
What a hit rate of 29 per cent is and is not
The out-of-sample figure is the rung’s headline and it invites a stronger conclusion than it supports, so it is worth being exact about what a hit rate measures here.
It counts prescriptions, not damage. The rule is scored right or wrong on whether it names the family that wins, and a rule that names the runner-up on a region where the top two are within a per cent has been counted wrong for a difference no reader could see. So the hit rate is a harsh statistic by construction.
The complementary number is the one a user needs, and it is a different quantity: how much worse the map the rule prescribes is than the best available. A rule that is wrong most of the time and cheap when wrong is a perfectly serviceable rule; one that is wrong a third of the time and catastrophic when wrong is not.
Both are computed on the way to the hit rate, since the scoring produces every family’s score on every region. Reporting the distribution of the shortfall — how much is lost when the rule misprescribes — would say whether 29 per cent is an indictment or a curiosity.
The reason to report the hit rate anyway is that it is the claim the rule makes. The rule is stated as a prescription, not as a bound: use a cylindrical projection for the tropics asserts that this is the thing to do, and the natural test of an assertion of that form is how often it is right. A defence of the rule on the grounds that being wrong is cheap is a different rule, and a better one, and nobody teaches it.
It is recorded here as a shortfall rather than computed, because computing it is a rerun of the scoring with a different reduction and it belongs beside the number it qualifies.
Which is the rung’s honest position. The rule fails the test of its own claim, badly, on real regions. Whether it should nevertheless be taught depends on a number this rung has and did not report, and saying so is better than letting 29 per cent stand as the whole finding.
Where the ladder goes next
Seven rungs of this ladder have taken a confident and widely repeated claim and measured it. The claims so far have all been about projections. The one that has not been touched is about the reader: that a map’s distortion is something a person can see — which is a claim about perception, and the only one on the list this collection has no instrument for. It is also the claim Mercator against Peters was fought over, by people who all agreed about the numbers.
Named alongside this one
Essays reaching for the same objects. Nobody chose these; they are what the concept index makes visible.
- The pooled score abandons a region kavrayskiy's criterion · optimisation · projection selection · purpose · region
- The shortest route between two coasts azimuthal · optimisation · purpose · region
- A family is not closed under averaging conic · optimisation · projection family
- The aspect has three numbers, not one kavrayskiy's criterion · optimisation · projection selection
- The best flat picture is not a map azimuthal · optimisation · purpose
- The maps with no family are simply better optimisation · projection family · purpose
What links here
Every essay whose body links to this one.
The objects this essay names
Each one links to every other essay that touches it.
AuditAzimuthalBounding boxConicGeneralisationKavrayskiy's criterionOptimisationProjection familyProjection selectionPurposeRegionRule of thumb