What each projection optimises

The pooled score abandons a region

Fifteen rungs optimise for one region. An atlas is several, and pooling their samples into one area-weighted score is what everybody does — which on Britain and New Zealand serves Britain 1.2 times worse than it could be served alone and New Zealand 125 times worse. The worst-case objective makes them equal at 33 and 59, and the cost of sharing rises with separation from 1.4 to 59.

Fifteen rungs of this ladder optimise a projection for a region. The region is a box, a cap or an ellipse; the projection has parameters; a criterion turns the pair into a number and something searches.

Almost nothing anybody actually maps is one region. A national mapping agency covers a mainland and its overseas territories. An atlas plate carries two continents. A comparison of two countries has to draw them on one sheet or the comparison is not a comparison. In every one of those cases the same machinery is pointed at several regions, and the moment it is, a decision appears that nobody makes explicitly.

Two regions, three answers. Britain and New Zealand, 166° apart, with the pole of the best oblique conic under each of three objectives. Pooling the samples and taking an area-weighted score puts the pole in one place; refusing to let either region be worse than the other puts it somewhere else. The regions are drawn on Mollweide so that equal ground areas are equal page areas.
Fig. 1 Britain and New Zealand, 166° apart, with the pole of the best oblique conic under each of three objectives. Pooling the two regions’ samples and taking an area-weighted score puts the pole in one place; refusing to let either be worse than the other puts it a long way away. Both are correct answers to questions nobody has distinguished.

Three objectives, all defensible

Handed a set of regions and a criterion, there are three obvious things to minimise.

The pooled score treats the union as one region: put all the samples together and take the area-weighted mean, which is exactly what a regional criterion does if the samples arrive together. It is what happens by default, because it requires no decision at all.

The equal-weight score takes the mean of the regions’ own scores, so a small territory counts as much as a large one.

The worst case takes the largest of the regions’ own scores, so no region may be sacrificed for another. It is Chebyshev’s principle applied to a set of regions rather than to the points of one, and it is the objective a mapping agency with a statutory duty to every territory actually has.

The first is the default and the third is usually the intention. They are different problems.

It is worth being clear that none of the three is a mistake. Each is a complete and coherent answer to a question somebody might be asking, and a published optimisation is entitled to use any of them. What is not defensible is using one and reporting the answer as the best projection for the set, because the three disagree by factors that are larger than anything else in this ladder — larger than the choice of criterion, larger than the choice of projection family, and larger than the difference between a good aspect and a bad one for a single region.

What the default does

What each objective gives each region. Britain and New Zealand under an oblique conic, with each region's Kavrayskiy score on a logarithmic scale. The top row is what each could have alone. The pooled objective serves one of them almost as well as if the other did not exist and abandons the other outright; the worst-case objective makes them equal and makes both worse than either would be under the pooled answer's favourite.
Fig. 2 Britain and New Zealand under an oblique conic, with each region’s Kavrayskiy score on a logarithmic scale. The top row is what each could have alone. The pooled objective serves Britain almost as well as if New Zealand did not exist and abandons New Zealand outright; the worst-case objective makes them equal and makes both far worse than the pooled answer’s favourite.

Britain alone can reach 0.00054 and New Zealand alone 0.00030. The pooled optimum gives Britain 0.00066 — a cost of 1.2 times — and New Zealand 0.03793, which is 125 times worse than it could have been. The pooled objective has not compromised between them. It has decided, and the decision is not visible in the number it reports.

The mechanism is that the pooled objective is an area-weighted mean, and Britain’s sampled area is 84 per cent of the union against New Zealand’s 16. Improving Britain by a little buys more than improving New Zealand by a lot, so the search walks into Britain’s basin and stays there.

The comparison that makes the size of the abandonment concrete is with the worst-case answer, which costs Britain 33 times and New Zealand 59. Read one way that is much worse: both regions are far worse served than under the pooled optimum’s favourite. Read the other way it is the only answer in which the two regions are being treated as one problem at all — the pooled answer is not a compromise between them, it is New Zealand’s optimum abandoned in favour of a map that is essentially Britain’s own.

Where the pooled objective puts the damage. For each pair, how much worse each region is under the pooled optimum than it would be alone, on a logarithmic scale. The two bars in a row are almost never the same length: the pooled objective serves one region nearly as well as if it were alone and loads the whole cost onto the other. On Britain and New Zealand the ratio between the two is a hundred to one.
Fig. 3 For four pairs of regions, how much worse each region is under the pooled optimum than it would be alone, on a logarithmic scale. The two bars in a row are almost never comparable: the pooled objective serves one region nearly as well as if it were alone and loads the whole cost onto the other. The lopsidedness grows with separation and reaches a hundred to one.

It is not always the smaller region that loses, and that is worth being careful about. Japan and New Zealand have equal sampled weight — 50 per cent each — and the pooled optimum still costs Japan 5.1 times and New Zealand 1.0. The objective sacrifices whichever region’s score responds least steeply to being moved, which correlates with size and is not the same thing. A region with a sharp optimum defends itself; a region with a flat one is given away.

The one case where it is free

The Europe and eastern United States pair is worth a paragraph on its own, because it is the case a reader is most likely to have in front of them and it is the reassuring one.

At 72° of separation, sharing an oblique conic costs Europe 1.1 times its own optimum and the United States 1.4 times, and the three objectives put the pole within a few degrees of one another. That is the two-plate atlas spread, and the measurement says the convention of drawing both on one projection is not merely tolerable but nearly free — the shared optimum is a good map for both, not a compromise either would complain about.

What makes it free is that both regions are mid-latitude and on the same side of the world, so a single conic’s band of low distortion can cover both. The band is the thing being shared, and its width is what decides whether two regions fit inside it: what a standard parallel buys is that band’s shape, and the separation at which sharing stops working is roughly the separation at which the second region falls outside it.

That is also why the cost curve is so steep rather than gradual. It is not that the map slowly deteriorates as the regions move apart; it is that at some separation one of them leaves the band, and after that the map is doing nothing for it at all.

The trade, drawn

The reason the three objectives disagree is that the two regions are genuinely in competition, and the shape of that competition is the useful object.

The two regions trade against each other. Every aspect in the search, plotted by what it does to one region against what it does to the other. The lower-left frontier is the set of aspects that cannot be improved for one without damaging the other, and every objective is a way of choosing a point on it. The three marked points are the three objectives' answers; the corner of the plot, where both regions get their own optimum, is empty and always will be.
Fig. 4 Every aspect in the search, plotted by what it does to Britain against what it does to New Zealand. The lower-left frontier is the set of aspects that cannot be improved for one without damaging the other, and every objective is a way of choosing a point on it. The marked corner, where both regions get their own optimum, is empty and always will be.

The picture is a Pareto front, and it says three things at once. The corner where both regions reach their own optimum is unoccupied, so sharing is not free. The front is long and curved, so the choice of point on it matters a great deal. And each objective is a rule for picking a point: the pooled score picks where a weighted sum is smallest, the worst case picks where the two coordinates are equal, and neither is more correct than the other.

The ranking is not an order makes the neighbouring argument about criteria; the average was a choice of norm makes it about aggregation over space. This is the same freedom a third time, over regions, and it is the one with a name attached to each of its terms — because the regions are places with people in them.

The cost of sharing, against distance

The one thing that is not a matter of choice is that sharing costs something, and how much depends on one measurable quantity.

What sharing a projection costs. For four pairs of regions, how much worse the worst-served region is under a shared projection than it would be under its own. Under the worst-case objective it runs from 1.4× at 72° of separation to 58.8× at 166°. Under the pooled objective it is worse still, because that objective is free to abandon one of them.
Fig. 5 For four pairs of regions, how much worse the worst-served region is under a shared projection than it would be under its own, against the angular separation of the pair. Under the worst-case objective it runs from 1.4 times at 72° to 58.8 times at 166°. Under the pooled objective it is worse still, because that objective is free to abandon one of them entirely.
pair apart worst case pooled
Europe and the eastern United States 72° 1.4× 1.3×
Japan and New Zealand 84° 6.6× 5.1×
Chile and Europe 116° 22.6× 25.6×
Britain and New Zealand 166° 58.8× 124.6×

The rise is steep and it is monotone in the separation across all four pairs. At 72° apart — Europe and the eastern United States, which is the classic two-plate atlas spread — sharing an oblique conic costs the worse-served region 40 per cent, which is nothing. At 166° it costs a factor of 59, which is the difference between a national grid and a world map.

That gives a rule with a number in it, which is what how many sheets an atlas needs supplies for coverage and nothing has supplied for this: below about 80° of separation, share the projection; past about 120°, do not, and the sheet count that essay computes is the price of not sharing.

Why the answer jumps

The search behaves in a way that is worth seeing, because it explains why this is a decision rather than a dial.

The landscape a shared projection is searched on. The pooled score of Chile and Europe at every position of the oblique conic's pole, shaded on Mollweide. The surface has two basins rather than one — one favouring each region — and the optimum sits in whichever is deeper. That is why the answer moves discontinuously when the weighting changes: it does not slide from one region to the other, it jumps.
Fig. 6 The pooled score of Chile and Europe at every position of the oblique conic’s pole. The surface has two basins rather than one — one favouring each region — and the optimum sits in whichever is deeper. Changing the weighting does not slide the answer from one region to the other; past a threshold it jumps from one basin to the other.

Two separated regions produce a score surface with two basins, one around each region’s own optimum. A weighting decides which basin is deeper, and the optimum is in that basin — so as the weight is moved the answer sits still, sits still, and then jumps.

That is the same structure where the valley breaks in two found inside a single region when the region grows past a threshold, and the shape of the valley measures for one basin. Here the two basins are not an artefact of a difficult region; they are the two regions, and the threshold is the weighting rather than the size.

The practical consequence is that a sensitivity analysis on the weights is useless in the ordinary way. Perturbing the weight by ten per cent will almost always change nothing at all, and occasionally change everything.

What the equal-weight objective is for

The third objective has been in every figure and has not been argued for, and it deserves one paragraph because it is the one most likely to be what somebody actually wants.

Equal weighting says each region counts once, whatever its size. On Britain and New Zealand it gives 0.01315 and 0.02159 — worse than the pooled optimum for Britain, better for New Zealand, and not equal. It is not the worst case and it is not the pooled score; it sits between them, and it is the objective a federation of equal members would write down.

Its defect is that it is not scale-free in the way the other two are. Splitting one region into two halves doubles its weight under equal weighting and changes nothing under the pooled objective or the worst case. So an equal-weight optimisation can be gamed by how the territories are enumerated, which is a property that should disqualify it from anything statutory and does not disqualify it from being the honest description of many real decisions.

Three objectives, three defensible readings, three different maps. The number to take away is not which is right but that the difference between them is a factor of a hundred on the pair measured here, and that a published optimisation which does not say which of the three it used has not said what it computed.

What was computed, and how

The projection is a Lambert conformal conic in a free oblique aspect, searched over the position of its pole on a 24 × 13 grid followed by a compass walk from each objective’s own best grid point. The criterion is Kavrayskiy’s, normalised as every comparison in this file normalises — each candidate scaled so the geometric mean of its areal factor over the region is one.

Each region’s alone score is taken from the same grid, with the shared optima added as further starting points for its walk. That is not fastidiousness: running a separate search for each region returned a worse score for the eastern United States than the shared optimum achieved, which reads as sharing being free and is an artefact of two walks settling in different basins of a multimodal surface. A region cannot do better shared than alone, and a measurement saying it did is measuring two searches rather than two regions.

The separation is the angular distance between the regions’ centres, which is the crudest possible summary of how far apart they are and is deliberately so — a measure involving their extents would be fitted to this result rather than tested against it.

The assertions require four things separately: that no region score better shared than alone, anywhere; that the cost of sharing rise with separation across the four pairs; that the pooled optimum and the worst-case optimum be different points; and that the worst-case optimum be no worse on the worst region than the pooled one is, which is what minimax means and is the check that both searches are searching the same space.

Where the model stops

Two regions, one family. Everything here is a pair, and a national series covering five territories is a harder problem whose Pareto front is a surface rather than a curve. Nothing in the argument depends on there being two, and every number does.

One projection family. The oblique conic has two free parameters here and a third — the cone constant — held at whatever the family’s own fit chooses. Allowing more freedom would lower every score and would not obviously change the ratios, since both the alone and the shared optima would move together, but that is a claim this rung has not tested.

And the regions are boxes and ellipses. Chile is a box 10° by 39°, which is a fair caricature of a long thin country, and Britain is a box. A real territory’s boundary would change the sampled weights, which is the quantity the pooled objective runs on, so the 84-to-16 split is a fact about the caricature.

The generalisation

The rule is that pooling several things into one score is a weighting decision, and the weights are whatever the pooling happened to produce.

Nobody chooses to weight Britain at 84 per cent and New Zealand at 16. They put the samples in one array, and the array has more of one than the other, because one region is bigger and the sampler is uniform. The weighting is an artefact of the data structure — which is the same shape as the sample drawn on the page, where the domain of an average is inherited from the loop it was computed in.

The difference is that here the artefact has a constituency. A weighting that abandons a territory is a decision about that territory, and it is being made by the size of an array rather than by anybody. A pooled objective over regions is a political instrument that looks like an average, and the cheapest possible repair is to print the regions’ individual scores beside the total, which costs one line and makes the sacrifice visible.

The habit generalises past maps. Whenever an objective sums over cases of unequal size — a model fitted to datasets of different lengths, a benchmark averaged over problems of different difficulty, a design optimised for a mixture of users — the same question applies: what are the weights, did anybody choose them, and which case is being given away?

Who found it, and when

The mathematics is standard multi-objective optimisation and none of it is new. Pareto’s ideas date from 1896, minimax from earlier still, and the fact that a weighted sum reaches only the convex hull of a front is exactly what which projection a weighting can make best measures for criteria rather than for regions.

The cartographic literature has the neighbouring problems and not this one. Choosing a projection for a region is treated thoroughly; choosing a sheet layout for a territory is treated as a covering problem; and the case of one projection serving several separated regions is handled in practice by simply not doing it — national grids give overseas territories their own zones, and atlases give continents their own plates. That is the right answer and it is arrived at by convention rather than by measurement, which means nobody can say what the convention is worth or where its boundary lies.

The measured boundary is around 80 to 120 degrees of separation. Below it, sharing costs a few tens of per cent and the convention is over-cautious; above it, sharing costs a factor of twenty or more and the convention is right. Between them the answer depends on the objective, and that is the range in which somebody has to say which of the three they meant.

Where the ladder goes next

Sixteen rungs have chosen a projection for a region, for a line, for a sheet and now for a set. Every one of them takes the region as the thing to be served, and a region here is a box with an area. The thing anybody is actually mapping is a set of places with people in them, whose distribution is nothing like uniform over any box — and an objective that weights a square kilometre of tundra as heavily as a square kilometre of city is making the same unexamined weighting decision this rung has just measured, one level down.

Named alongside this one

Essays reaching for the same objects. Nobody chose these; they are what the concept index makes visible.

What links here

Every essay whose body links to this one.

The objects this essay names

Each one links to every other essay that touches it.

AggregationAspectKavrayskiy's criterionObjective functionOptimisationPareto frontProjection selectionPurposeRegionSheet layoutTrade-offVerificationWeighting