What is taught wrongly

The region was a rectangle in a projection

Every verdict this collection hands out is a verdict about a region, and almost every region it has been given is a latitude–longitude box. A box discards a territory's shape, which costs a few per cent, and it records the territory's width in degrees of longitude, which costs up to a factor of 1.72 and changes which family wins on four of forty stated territories. Convert that width to arc and the box gives the right answer every time — so the comparison had a projection in it before it began.

Assumes The area weighting was a readership all along.

The ranking is not an order closes on the one thing eleven measurements of a projection comparison had never examined: the region is given. Every score in this field weights a square kilometre by where it is, every verdict is a verdict for some stated place, and eight of the ten places used were boxes somebody drew around a country.

A box is not how any territory is shaped. It is how a territory is recorded — the four numbers a gazetteer stores, a web service returns and anybody asking which projection for Chile is likely to type. The question this asks is how much of the answer is in the box, and the answer has two parts of very different sizes.

Thirteen measurements now stand between the received advice on this ground and a defensible verdict, and every one of them has found the comparer’s own freedom somewhere: the tolerance that decides the verdict in the threshold, which projection a weighting can make best in the weights, the average was a choice of norm in the aggregation and the area weighting was a readership all along in who the map is for. The region has stood outside all of them as the one thing supplied from the world. It is not supplied from the world. It is four numbers in a coordinate system, and the coordinate system is a projection.

A territory, and two enclosures somebody might score instead. A stated territory at 45° north — an ellipse twenty-eight degrees long and eight wide, its long axis bearing 45° — with its latitude–longitude bounding box and the smallest cap around it, on a globe. The territory fills 40.6 per cent of the box and 28.6 per cent of the cap. Scoring the library over each names a winner: an azimuthal for the territory itself, an azimuthal for its bounding box, an azimuthal for the smallest cap around it. The cost of working from the box rather than the territory, measured as the box's winner scored on the territory against the territory's own winner, is a factor of 1.0584.
Fig. 1 A stated territory at 45° north — an ellipse twenty-eight degrees long and eight wide, its long axis bearing 45° — with its latitude–longitude bounding box and the smallest cap around it. The territory fills 40.6 per cent of the box. Scoring the library over each names a winner, and here all three agree; the cost of working from the box is a factor of 1.0584.

What a reading is, and what it costs

A territory here is a stated ellipse on the sphere with a centre, two semi-axes and a bearing for its long axis — the shape this ground has supported since the population sweeps began and which nothing has used to ask this question. The bearing is the part that matters, and it is the part a box cannot hold: a box’s axes are the graticule’s, so a territory running north-east has a box with no north-east in it. Fitting the aspect to the region prices what a territory’s orientation is worth when it is known — a factor of 183 for a long thin country — and a box is the record that throws it away. Around it sit the readings somebody might work from instead: the bounding box, the box with a tenth added for margin, the smallest cap, and the box read back as an ellipse.

Each reading names a winner, by scoring the library over that reading and taking the best family. The number that matters is not that winner’s score on the reading — which is a score for a region nobody is mapping — but its score on the territory, against what the territory’s own winner reaches. That ratio is what working from the reading costs, in the criterion’s own units.

The rectangle costs nothing; reading longitude as distance costs everything. Five readings of one territory — an ellipse at 45° north lying along a parallel — each scored by taking its own winning projection and measuring that projection on the territory, against what the territory's own winner reaches. The bounding box costs a factor of 1.4277 and names a conic where the territory wants an azimuthal. The last row is the same box with its width converted from degrees of longitude to degrees of arc, and it costs 1.0000 and names an azimuthal — the right answer. So the error is not that a box is a rectangle. It is that a box is recorded in longitude, and a degree of longitude at 45° is 1.414 times shorter than a degree of arc.
Fig. 2 Five readings of one territory, an ellipse at 45° north lying along a parallel. The bounding box costs a factor of 1.4277 and names a conic where the territory wants an azimuthal. The last row is the same box with its width converted from degrees of longitude to degrees of arc: it costs 1.0000 and names an azimuthal — the right answer.

The last row is the finding and everything below is it in detail. A bounding box’s rectangularity costs nothing measurable. Its units cost a great deal.

A box is recorded as a range of longitude and a range of latitude, and those are not the same kind of number. A degree of latitude is a degree of arc everywhere. A degree of longitude is a degree of arc only at the equator, and at 45° north it is 1.414 times shorter. So a box that is forty degrees of longitude wide and eight of latitude tall is not a region five times wider than it is tall; it is a region 3.5 times wider than it is tall. Reading the first as the second exaggerates the east–west extent, and exaggerating east–west extent is exactly the thing that pushes a projection choice towards a conic.

There is a shorter way to say it. Plotting latitude against longitude on ordinary axes is a projection — the plate carrée, the projection nobody chooses is the essay about it, and its argument is that the map is what happens when nobody decides anything. A bounding box read as a rectangle on the ground is that projection applied to the region before any projection is chosen for it. So a comparison that takes the region as given has already made a projection choice, and made it silently, and made the one the third essay in this line describes as the worst available.

Correct that one factor and the box is right

Correct one factor and the box gives the right answer every time. The same territory — an ellipse lying along a parallel — moved from 5° to 70° of latitude. The upper curve is the cost of scoring its bounding box; the lower is the cost of scoring that same box with its width read as arc rather than as degrees of longitude. The uncorrected box names a conic at every latitude above 5° and costs up to a factor of 1.7174 at 15°. The corrected one names what the territory names at every latitude and costs 1.0041 at worst. Nothing about the rectangle changed; only the units of its width did.
Fig. 3 The same territory, lying along a parallel, moved from 5° to 70° of latitude. The upper curve is the cost of scoring its bounding box; the lower is the cost of scoring that same box with its width read as arc. The uncorrected box names a conic at every latitude above 5° and costs up to 1.7174. The corrected one names what the territory names at every latitude and costs 1.0041 at worst.

The two curves are the whole argument. Nothing about the rectangle changes between them — it has the same corners, the same area on the ground, the same fraction of territory inside it. Only the units of its width change, and with them the verdict at every latitude from 15° upward.

At 5° the two agree, because a degree of longitude at 5° is 0.4 per cent shorter than a degree of arc and the error is below the resolution of the comparison. Above that the uncorrected box names a conic while the territory and the corrected box name an azimuthal, and the cost peaks at 1.72 at 15° rather than at the pole — because the verdict changes where the two candidate families cross, and where that crossing sits depends on the region’s real proportions, not on how large the correction is.

That non-monotonicity is worth keeping. A correction whose size grows as 1/cos φ does not produce a cost that grows as 1/cos φ, because the cost is the gap between two discrete answers rather than a continuous quantity. Between 30° and 70° the correction more than doubles while the cost falls by a third.

It also explains a shape that looks wrong at 30°, where the uncorrected box costs only 1.0157 while naming the wrong family. Naming the wrong family and paying almost nothing for it is possible whenever the two families are nearly as good as each other on the territory — which is most of what the projections that are beaten on both counts is about from the other side. A disagreement is not automatically expensive, and a cost is not automatically a disagreement; this essay measures both and they are only loosely related.

The cost is not how empty the box is

Anybody asked in advance what makes a bounding box a bad stand-in for a territory would say: how much of the box is not territory. It is the obvious quantity, it is easy to compute, and it is not the one that decides.

The cost of a box is not how much of it is territory. One territory turned through ninety degrees, with the cost of scoring its bounding box instead of itself. The obvious expectation is that the cost tracks how empty the box is, and it does not: the box is fullest at bearings of 0° and 90° — 77 and 80 per cent territory — and the worst cost of the seven is at 90°, a factor of 1.4277, where the box is 80 per cent full. At 45°, where the box is emptiest at 41 per cent, the cost is only 1.0584. The box names a different family at 90°.
Fig. 4 One territory turned through ninety degrees, with the cost of scoring its bounding box instead of itself, and the share of the box that is territory printed beside each mark. The box is fullest at bearings of 0° and 90° — 77 and 80 per cent — and the worst cost of the seven is at 90°, where the box is 80 per cent full. At 45°, where the box is emptiest at 41 per cent, the cost is 1.0584.

The two quantities move opposite ways over most of the range. A territory turned diagonally has a box nearly two thirds of which is somewhere else, and reading it costs six per cent. A territory lying along a parallel fills four fifths of its own box and reading it costs forty-three per cent and the wrong family.

The reason is the same one as before. A diagonal territory’s box is nearly square, and a nearly square box has no strong east–west exaggeration to be misread; it is emptier, and emptiness only dilutes the score rather than pointing it somewhere else. A territory along a parallel has a wide shallow box, which is exactly where the longitude-versus-arc factor bites hardest, and the box’s own aspect ratio is what the family rule reads.

The rule of thumb, scored ran the standard advice — cylindrical near the equator, conic in the middle latitudes, azimuthal at the poles — over thirty regions and got it right nineteen times. Every one of those regions was a box. The rule reads a region’s latitude and its proportions, and this says that the proportions it was reading were wrong by 1/cos φ on every region away from the equator.

A longer country is worse served by its box, and only slowly. A territory at 45° north with its long axis at 45°, made steadily longer. The share of its box that is territory falls from 79 per cent at a circle to 24 at six to one, and the cost of working from the box rises from 1.0000 to 1.1867. It rises, so the emptiness of a box is not nothing. But it rises slowly: at three to one the box is 46 per cent territory and costs 1.0413, against the factor of 1.72 the recorded width costs at the worst latitude in the figure above. Shape is the small term here and units are the large one.
Fig. 5 A territory at 45° north with its long axis at 45°, made steadily longer. The share of its box that is territory falls from 79 per cent at a circle to 24 at six to one, and the cost rises from 1.0000 to 1.1867. Emptiness is not nothing — but at three to one the box is 46 per cent territory and costs 1.0413, against the factor of 1.72 the recorded width costs at the worst latitude.

Emptiness does cost something, and it is worth having the size of it. A territory six times longer than it is wide fills a quarter of its box and pays nineteen per cent for the substitution. That is a real number and it is an order of magnitude below the other one at its worst.

So there are two effects and they are separable. The shape a box loses is a few per cent, rising slowly with elongation and never changing which family wins in any case measured here. The units a box is recorded in are up to a factor of 1.72, and they change the family.

Over a population

Most boxes cost nothing and a few cost a great deal. Forty stated territories — four latitudes, five bearings and two elongations — each scored twice, once as itself and once as its bounding box. The median cost is 1.0013, so for half of them working from the box is free to four decimal places. 4 of the forty name a different family from the box than from the territory, and the worst costs a factor of 1.763. The mean share of a box that is territory is 68 per cent, which is the quantity anybody would reach for and which the figures above show is not the one that decides.
Fig. 6 Forty stated territories — four latitudes, five bearings and two elongations — each scored as itself and as its bounding box. The median cost is 1.0013, so for half of them working from the box is free to four decimal places. Four of the forty name a different family from the box than from the territory, and the worst costs a factor of 1.763.

The distribution is the practically useful part, and it is the shape these measurements have found before: mostly nothing, and then something large. Half the territories are served by their boxes to a tenth of a per cent. Ten per cent of them get a different family.

That is the same shape the rule scored out of sample found for the rule of thumb itself: right most of the time, wrong in a way that is not random, and quoted without a statement of which. It makes the box a defensible default and an indefensible convention. A cartographer who takes a box, scores it, and takes the answer will usually be right; a comparison that reports a verdict without saying which reading of the region it used has left a one-in-ten chance of naming the wrong family unstated, and has no way to tell a reader which case they are in.

And there is a cheap repair, which is the other reason this is worth having. Multiplying the box’s width by the cosine of its middle latitude before scoring it costs one line and removes the whole of the large effect. Nothing here recommends abandoning boxes; it recommends reading them as regions on a sphere rather than as rectangles.

What it does to the verdicts already on this ground

Two things, and they pull in opposite directions.

The regions the scoring on this ground has used are boxes, and the four it names as countries — Europe, Chile, the conterminous United States, and the band of the tropics — are all recorded the same way. Three of the four are wide and shallow away from the equator, which is precisely the configuration where the correction bites: Europe’s box spans fifty degrees of longitude and thirty-five of latitude, and at its middle latitude of 52.5° a degree of longitude is 1.64 times shorter than a degree of arc, so its true proportions are nearer square than the numbers say. Whether that changes the winner for Europe is a computation this essay does not run, because the box is the region those essays named and rescoring it here would be answering a different question from the one they asked.

The other direction is that this does not weaken any of them. A verdict reported for a stated region is correct for that region; what this adds is that the region was stated in a coordinate system, and that a reader taking the verdict as advice about a country is making a substitution nobody priced. Designing a grid for one region is the same distinction in the practice field — a national grid is designed for a declared box, and its optimality is optimality for that declaration.

What each number was compared against

A cap around a round territory is that territory, so the cap reading must name the same winner at a cost of exactly one. It does, to six decimal places — the control that says the arithmetic is not manufacturing a cost where there is none.

A box must contain its territory, checked at every boundary point at four bearings, and must be no smaller in solid angle. Both are enclosures and an enclosure that fails either is an arithmetic error rather than a reading.

A disc fills π/4 of its own box, which is 78.5 per cent, and the measured fill for a round territory is inside 60 to 90 per cent. That is the check on the area measure, and it caught the first version: summing the per-region sample weights compares a region with itself rather than with another, and every box came back at three per cent of its own territory.

Turning the territory must cost something, or the orientation a box discards was worth nothing and there is no essay here. The worst turn costs 1.43.

And the criterion is one criterion. Every cost above is measured on Kavrayskiy’s, which is the one the ranking on this ground uses throughout. The average was a choice of norm is the warning that another summary would reorder some of these, and it applies to the costs as well as to the verdicts.

What forty stated territories are not

They are ellipses. A real country is neither an ellipse nor a box, and the cost of the substitution measured here is the cost of one stated shape against another. What travels is the mechanism — a longitude range read as an arc range — which is a property of the recording rather than of the shape, and which no real coastline escapes.

The population is a grid rather than a sample of the world. Four latitudes, five bearings and two elongations are chosen to span the space rather than to represent the countries that exist, so the one-in-ten disagreement rate is a rate over that grid. A population weighted by real territories would give a different number and, given that most countries are not at high latitude, probably a smaller one.

Only three families compete. The verdict is which of cylindrical, conic and azimuthal wins, with the best member of each fitted, as the rule-of-thumb essays do. A comparison over named projections rather than families would change what counts as a disagreement — probably increasing it, since a family’s best member is refitted to whatever region it is handed and a named projection is not, so the family comparison is the one that absorbs a bad reading most gracefully. The rate of one in ten measured here should therefore be read as a floor on how often a bad reading of a region changes the answer, not as an estimate of it.

And the correction is not free of assumptions either. Multiplying by the cosine of the middle latitude is right for a box that is not very tall; for a box spanning fifty degrees of latitude there is no single factor, and the honest repair there is to score the territory. Chile’s box is thirty-nine degrees of latitude tall, which is exactly that case: its width in arc runs from 0.94 of its width in degrees at the northern edge to 0.56 at the southern, and no one cosine describes it.

Still open: whether the box was ever the region anybody meant

The correction above treats the box as a faithful record in bad units, and there is a reading on which it is not a record of the territory at all.

A bounding box is what a dataset covers, and a dataset covers what somebody decided to collect. A national mapping agency’s box is its jurisdiction, which includes territorial sea and excludes a neighbour’s enclave; a web service’s box is whatever tiles it has; an atlas sheet’s box is chosen so the sheet series tiles. None of those is the shape of the land, and each is a different object with its own defensible claim to be the thing a projection should be chosen for.

That turns the question from an error to be corrected into a choice to be named — which is the shape every other measurement on this ground has taken. The criterion was a choice, the weighting was a choice, the aggregation was a choice, and the readership was a choice; the region looked like the one thing given from outside, and it is a box drawn by somebody for a purpose.

What would settle it is not another sweep over stated shapes. It is a comparison of two boxes for one country that both exist — the jurisdiction’s and the land’s — and a measurement of whether the projection each names differs by more than the reading error priced here. Whether any published atlas states which of them its choice was made for, and what it would cost to have chosen for the other, are questions a synthetic population cannot ask.

Named alongside this one

Essays reaching for the same objects. Nobody chose these; they are what the concept index makes visible.

The objects this essay names

Each one links to every other essay that touches it.

AspectBounding boxConventionMeasurementPlate carréeProjection selectionPurposeRankingRegionRegional distortionRule of thumbVerification