The pooled score abandons a region
Fifteen rungs of this ladder optimise a projection for a region. The region is a box, a cap or an ellipse; the projection has parameters; a criterion turns the pair into a number and something searches.
Almost nothing anybody actually maps is one region. A national mapping agency covers a mainland and its overseas territories. An atlas plate carries two continents. A comparison of two countries has to draw them on one sheet or the comparison is not a comparison. In every one of those cases the same machinery is pointed at several regions, and the moment it is, a decision appears that nobody makes explicitly.
Three objectives, all defensible
Handed a set of regions and a criterion, there are three obvious things to minimise.
The pooled score treats the union as one region: put all the samples together and take the area-weighted mean, which is exactly what a regional criterion does if the samples arrive together. It is what happens by default, because it requires no decision at all.
The equal-weight score takes the mean of the regions’ own scores, so a small territory counts as much as a large one.
The worst case takes the largest of the regions’ own scores, so no region may be sacrificed for another. It is Chebyshev’s principle applied to a set of regions rather than to the points of one, and it is the objective a mapping agency with a statutory duty to every territory actually has.
The first is the default and the third is usually the intention. They are different problems.
It is worth being clear that none of the three is a mistake. Each is a complete and coherent answer to a question somebody might be asking, and a published optimisation is entitled to use any of them. What is not defensible is using one and reporting the answer as the best projection for the set, because the three disagree by factors that are larger than anything else in this ladder — larger than the choice of criterion, larger than the choice of projection family, and larger than the difference between a good aspect and a bad one for a single region.
What the default does
Britain alone can reach 0.00054 and New Zealand alone 0.00030. The pooled optimum gives Britain 0.00066 — a cost of 1.2 times — and New Zealand 0.03793, which is 125 times worse than it could have been. The pooled objective has not compromised between them. It has decided, and the decision is not visible in the number it reports.
The mechanism is that the pooled objective is an area-weighted mean, and Britain’s sampled area is 84 per cent of the union against New Zealand’s 16. Improving Britain by a little buys more than improving New Zealand by a lot, so the search walks into Britain’s basin and stays there.
The comparison that makes the size of the abandonment concrete is with the worst-case answer, which costs Britain 33 times and New Zealand 59. Read one way that is much worse: both regions are far worse served than under the pooled optimum’s favourite. Read the other way it is the only answer in which the two regions are being treated as one problem at all — the pooled answer is not a compromise between them, it is New Zealand’s optimum abandoned in favour of a map that is essentially Britain’s own.
It is not always the smaller region that loses, and that is worth being careful about. Japan and New Zealand have equal sampled weight — 50 per cent each — and the pooled optimum still costs Japan 5.1 times and New Zealand 1.0. The objective sacrifices whichever region’s score responds least steeply to being moved, which correlates with size and is not the same thing. A region with a sharp optimum defends itself; a region with a flat one is given away.
The one case where it is free
The Europe and eastern United States pair is worth a paragraph on its own, because it is the case a reader is most likely to have in front of them and it is the reassuring one.
At 72° of separation, sharing an oblique conic costs Europe 1.1 times its own optimum and the United States 1.4 times, and the three objectives put the pole within a few degrees of one another. That is the two-plate atlas spread, and the measurement says the convention of drawing both on one projection is not merely tolerable but nearly free — the shared optimum is a good map for both, not a compromise either would complain about.
What makes it free is that both regions are mid-latitude and on the same side of the world, so a single conic’s band of low distortion can cover both. The band is the thing being shared, and its width is what decides whether two regions fit inside it: what a standard parallel buys is that band’s shape, and the separation at which sharing stops working is roughly the separation at which the second region falls outside it.
That is also why the cost curve is so steep rather than gradual. It is not that the map slowly deteriorates as the regions move apart; it is that at some separation one of them leaves the band, and after that the map is doing nothing for it at all.
The trade, drawn
The reason the three objectives disagree is that the two regions are genuinely in competition, and the shape of that competition is the useful object.
The picture is a Pareto front, and it says three things at once. The corner where both regions reach their own optimum is unoccupied, so sharing is not free. The front is long and curved, so the choice of point on it matters a great deal. And each objective is a rule for picking a point: the pooled score picks where a weighted sum is smallest, the worst case picks where the two coordinates are equal, and neither is more correct than the other.
The ranking is not an order makes the neighbouring argument about criteria; the average was a choice of norm makes it about aggregation over space. This is the same freedom a third time, over regions, and it is the one with a name attached to each of its terms — because the regions are places with people in them.
The cost of sharing, against distance
The one thing that is not a matter of choice is that sharing costs something, and how much depends on one measurable quantity.
| pair | apart | worst case | pooled |
|---|---|---|---|
| Europe and the eastern United States | 72° | 1.4× | 1.3× |
| Japan and New Zealand | 84° | 6.6× | 5.1× |
| Chile and Europe | 116° | 22.6× | 25.6× |
| Britain and New Zealand | 166° | 58.8× | 124.6× |
The rise is steep and it is monotone in the separation across all four pairs. At 72° apart — Europe and the eastern United States, which is the classic two-plate atlas spread — sharing an oblique conic costs the worse-served region 40 per cent, which is nothing. At 166° it costs a factor of 59, which is the difference between a national grid and a world map.
That gives a rule with a number in it, which is what how many sheets an atlas needs supplies for coverage and nothing has supplied for this: below about 80° of separation, share the projection; past about 120°, do not, and the sheet count that essay computes is the price of not sharing.
Why the answer jumps
The search behaves in a way that is worth seeing, because it explains why this is a decision rather than a dial.
Two separated regions produce a score surface with two basins, one around each region’s own optimum. A weighting decides which basin is deeper, and the optimum is in that basin — so as the weight is moved the answer sits still, sits still, and then jumps.
That is the same structure where the valley breaks in two found inside a single region when the region grows past a threshold, and the shape of the valley measures for one basin. Here the two basins are not an artefact of a difficult region; they are the two regions, and the threshold is the weighting rather than the size.
The practical consequence is that a sensitivity analysis on the weights is useless in the ordinary way. Perturbing the weight by ten per cent will almost always change nothing at all, and occasionally change everything.
What the equal-weight objective is for
The third objective has been in every figure and has not been argued for, and it deserves one paragraph because it is the one most likely to be what somebody actually wants.
Equal weighting says each region counts once, whatever its size. On Britain and New Zealand it gives 0.01315 and 0.02159 — worse than the pooled optimum for Britain, better for New Zealand, and not equal. It is not the worst case and it is not the pooled score; it sits between them, and it is the objective a federation of equal members would write down.
Its defect is that it is not scale-free in the way the other two are. Splitting one region into two halves doubles its weight under equal weighting and changes nothing under the pooled objective or the worst case. So an equal-weight optimisation can be gamed by how the territories are enumerated, which is a property that should disqualify it from anything statutory and does not disqualify it from being the honest description of many real decisions.
Three objectives, three defensible readings, three different maps. The number to take away is not which is right but that the difference between them is a factor of a hundred on the pair measured here, and that a published optimisation which does not say which of the three it used has not said what it computed.
What was computed, and how
The projection is a Lambert conformal conic in a free oblique aspect, searched over the position of its pole on a 24 × 13 grid followed by a compass walk from each objective’s own best grid point. The criterion is Kavrayskiy’s, normalised as every comparison in this file normalises — each candidate scaled so the geometric mean of its areal factor over the region is one.
Each region’s alone score is taken from the same grid, with the shared optima added as further starting points for its walk. That is not fastidiousness: running a separate search for each region returned a worse score for the eastern United States than the shared optimum achieved, which reads as sharing being free and is an artefact of two walks settling in different basins of a multimodal surface. A region cannot do better shared than alone, and a measurement saying it did is measuring two searches rather than two regions.
The separation is the angular distance between the regions’ centres, which is the crudest possible summary of how far apart they are and is deliberately so — a measure involving their extents would be fitted to this result rather than tested against it.
The assertions require four things separately: that no region score better shared than alone, anywhere; that the cost of sharing rise with separation across the four pairs; that the pooled optimum and the worst-case optimum be different points; and that the worst-case optimum be no worse on the worst region than the pooled one is, which is what minimax means and is the check that both searches are searching the same space.
Where the model stops
Two regions, one family. Everything here is a pair, and a national series covering five territories is a harder problem whose Pareto front is a surface rather than a curve. Nothing in the argument depends on there being two, and every number does.
One projection family. The oblique conic has two free parameters here and a third — the cone constant — held at whatever the family’s own fit chooses. Allowing more freedom would lower every score and would not obviously change the ratios, since both the alone and the shared optima would move together, but that is a claim this rung has not tested.
And the regions are boxes and ellipses. Chile is a box 10° by 39°, which is a fair caricature of a long thin country, and Britain is a box. A real territory’s boundary would change the sampled weights, which is the quantity the pooled objective runs on, so the 84-to-16 split is a fact about the caricature.
The generalisation
The rule is that pooling several things into one score is a weighting decision, and the weights are whatever the pooling happened to produce.
Nobody chooses to weight Britain at 84 per cent and New Zealand at 16. They put the samples in one array, and the array has more of one than the other, because one region is bigger and the sampler is uniform. The weighting is an artefact of the data structure — which is the same shape as the sample drawn on the page, where the domain of an average is inherited from the loop it was computed in.
The difference is that here the artefact has a constituency. A weighting that abandons a territory is a decision about that territory, and it is being made by the size of an array rather than by anybody. A pooled objective over regions is a political instrument that looks like an average, and the cheapest possible repair is to print the regions’ individual scores beside the total, which costs one line and makes the sacrifice visible.
The habit generalises past maps. Whenever an objective sums over cases of unequal size — a model fitted to datasets of different lengths, a benchmark averaged over problems of different difficulty, a design optimised for a mixture of users — the same question applies: what are the weights, did anybody choose them, and which case is being given away?
Who found it, and when
The mathematics is standard multi-objective optimisation and none of it is new. Pareto’s ideas date from 1896, minimax from earlier still, and the fact that a weighted sum reaches only the convex hull of a front is exactly what which projection a weighting can make best measures for criteria rather than for regions.
The cartographic literature has the neighbouring problems and not this one. Choosing a projection for a region is treated thoroughly; choosing a sheet layout for a territory is treated as a covering problem; and the case of one projection serving several separated regions is handled in practice by simply not doing it — national grids give overseas territories their own zones, and atlases give continents their own plates. That is the right answer and it is arrived at by convention rather than by measurement, which means nobody can say what the convention is worth or where its boundary lies.
The measured boundary is around 80 to 120 degrees of separation. Below it, sharing costs a few tens of per cent and the convention is over-cautious; above it, sharing costs a factor of twenty or more and the convention is right. Between them the answer depends on the objective, and that is the range in which somebody has to say which of the three they meant.
Where the ladder goes next
Sixteen rungs have chosen a projection for a region, for a line, for a sheet and now for a set. Every one of them takes the region as the thing to be served, and a region here is a box with an area. The thing anybody is actually mapping is a set of places with people in them, whose distribution is nothing like uniform over any box — and an objective that weights a square kilometre of tundra as heavily as a square kilometre of city is making the same unexamined weighting decision this rung has just measured, one level down.
Named alongside this one
Essays reaching for the same objects. Nobody chose these; they are what the concept index makes visible.
- The first break is mostly its denominator aspect · objective function · optimisation · purpose · trade-off · verification
- The maps with no family are simply better aspect · optimisation · purpose · trade-off · verification · weighting
- The aspect has three numbers, not one aspect · kavrayskiy's criterion · objective function · optimisation · projection selection
- The rule of thumb, scored aspect · kavrayskiy's criterion · optimisation · purpose · region
- The rule scored out of sample kavrayskiy's criterion · optimisation · projection selection · purpose · region
- Choosing for a line, not a region aspect · objective function · projection selection · purpose
What links here
Every essay whose body links to this one.
The objects this essay names
Each one links to every other essay that touches it.
AggregationAspectKavrayskiy's criterionObjective functionOptimisationPareto frontProjection selectionPurposeRegionSheet layoutTrade-offVerificationWeighting