What each projection optimises

The map that keeps the most ground inside a tolerance

A surveyor does not want the smallest worst case or the smallest mean square. They want as much of the territory as possible within a stated tolerance of true scale and are indifferent to how far outside the rest goes. That is a third criterion, it is answered by a third map, and on a 25° square at half a per cent it keeps 60.1 per cent of the area where the mean-square map keeps 45.4 and Chebyshev's 36.5 — bought by sending a tail out to 5.7 per cent where Chebyshev's worst is 3.0.

Assumes Chebyshev's map is the best at its worst, and not on average.

Chebyshev’s map is the best at its worst, and not on average put two criteria against each other on the same region. Chebyshev’s condition makes the scale constant on the boundary and minimises the worst departure anywhere; the mean-square condition minimises the area-weighted second moment of that departure. On a 25° spherical square the first is 8.3 per cent worse in root-mean-square scale error and the second 19.8 per cent worse in range, and neither is a rounding of the other.

It ended on the observation that a surveyor asked to choose between them would choose neither. What a national grid is specified by is a tolerance: the scale factor shall lie within one part in so many over the mapped area. Inside that band nothing is better or worse, and outside it the map is out of specification whether by a little or by a lot. Neither of the two criteria counts that, and the objective that does is a third one.

Three criteria, three maps, and the tolerance band each one encloses. A spherical square of 25° circumradius under three conformal maps, with the two contours at ±0.5% of true scale drawn heavy and the region between them the area each map keeps inside that tolerance. Chebyshev's map holds 36.5 per cent of the area there, the mean-square map 45.4, and the map fitted to the tolerance itself 60.1. The faint lines are further multiples of the tolerance, and they say where each map spends what it is not counting: the third map's outermost contours are much further out than the other two's.
Fig. 1 A spherical square of 25° circumradius under three conformal maps, with the two contours at ±0.5% of true scale drawn heavy and the region between them the area each map keeps inside that tolerance. Chebyshev’s map holds 36.5 per cent of the area there, the mean-square map 45.4, and the map fitted to the tolerance itself 60.1. The faint lines are further multiples of the tolerance, and they say where each map spends what it is not counting.

What a specification actually says

The criterion is not invented for the sake of a third curve. A national grid is defined by a scale factor and a statement about it, and the statement is almost always a band: the transverse Mercator zones of the Universal Transverse Mercator system carry a central scale factor of 0.9996 and a width chosen so that the scale stays within about one part in 2,500 of true across the zone, and a great many national grids are specified the same way with a different number in place of the 2,500.

Read as an objective, that says three things at once. Inside the band, nothing is preferred — a scale factor of 1.00002 is not better than 1.00035 for any purpose the specification has. Outside it, nothing is preferred either — 1.0008 and 1.0030 are both out. And what is wanted is as much of the mapped area inside as can be had.

That is a step function of the departure, integrated over area, and neither of the two criteria already fitted to a region’s condition is one. The worst case is a supremum and sees only the single worst point; the mean square is a smooth penalty and cares how far outside a point is, which the specification explicitly does not. Fitting the step is fitting what was asked for.

Why the objective is cheap, which is not obvious

A conformal map fitted this way is built from a series for the logarithm of its scale: the derivative is f=exphf' = \exp h, so f=exp(Reh)\lvert f'\rvert = \exp(\operatorname{Re} h) and

lnk(z)=Reh(z)+(the chart’s own metric term)\ln k(z) = \operatorname{Re}\,h(z) + \text{(the chart's own metric term)}

with hh a polynomial whose coefficients are exactly what the fit solves for. The map itself is a wildly nonlinear function of those coefficients — it is an exponential of a series, integrated — but its log-scale is affine in them. Perturbing one coefficient and taking second differences gives 10⁻¹⁷ against first differences of 10⁻⁵, which is rounding.

Two things follow, and the second is the one that makes the measurement possible at all.

The objective over a fixed sample of the region becomes one matrix–vector product, so evaluating a candidate map costs nothing.

And the line search becomes exact. Moving the coefficients along any direction changes each sample’s lnk\ln k affinely, so a sample is inside the tolerance on an interval of the step and outside it elsewhere. Collecting every interval’s two endpoints and sweeping them in order gives the exact best step along that line — not the best of a sampled few. That matters more than speed: it means a fit that has stopped moving has reached a point no single-coordinate move can improve, which is a property of the objective rather than of a step size, and it is what lets the next measurement say that several answers are genuinely several.

Every projection minimises something records the objectives cartographers have written down. This one is easy to write down and was, as far as this account can find, not fitted — and the reason is probably that it looks expensive. It is not.

It is worth being clear about what kind of object is being fitted, because the word condition has been used here for something narrower. A projection written as a condition means an exact requirement — the distance from these two places shall be right — which either has a solution or does not, and three conditions are one too many prices what happens when more are demanded than the construction has freedom for. A tolerance is not a condition in that sense. It is a score, and the whole content of this measurement is that scoring by area-inside-a-band gives a different answer from scoring by worst case or by mean square.

What the third criterion buys

The gain is largest where the tolerance bites hardest, and nothing at either end. The share of a 25° spherical square's area kept within a stated tolerance, by three conformal maps. At a tolerance of half a per cent the map fitted to the criterion keeps 60.1 per cent against the mean-square map's 45.4 and Chebyshev's 36.5 — 14.8 points of area bought by counting what is actually wanted. At a tolerance of 0.1 per cent the mean-square map is already nearly optimal and the gain is 0.11 points; at two per cent everything is inside and all three keep the region whole.
Fig. 2 The share of a 25° spherical square’s area kept within a stated tolerance, by three conformal maps. At a tolerance of half a per cent the map fitted to the criterion keeps 60.1 per cent against the mean-square map’s 45.4 and Chebyshev’s 36.5 — 14.8 points of area bought by counting what is actually wanted. At a tolerance of 0.1 per cent the gain is 0.11 points; at two per cent everything is inside and all three keep the region whole.

The gain is not uniform and its shape is the useful part. At a very tight tolerance almost nothing is inside for any map, and the best thing to do is what the mean-square map already does: flatten the scale field near the centre as hard as possible. The tolerance criterion agrees and buys a tenth of a point. At a very loose tolerance everything is inside for every map and there is nothing to buy.

In between — at half a per cent on this region — the criterion buys 14.8 points of area, which is a third more territory inside specification from the same twelve-term fit. That is the range a real specification sits in, because a specification loose enough to be met everywhere is not doing any work and one tight enough to be met almost nowhere is not either.

The criterion buys most where the region has corners to spend on. The share of each region's area kept within ±0.5% of true scale by the three criteria. a cap: 24.3, 24.3, 29.2; a spherical square: 36.5, 45.4, 60.1; an ellipse, axes 1 : 0.45: 65.0, 78.6, 84.6; a sliver, axes 1 : 0.2: 100.0, 100.0, 100.0. The gain runs from 0.0 points on the sliver, where every map already keeps the whole of it, to 14.8 on the square. A cap gains 4.9: its scale field is radial and the other two criteria are optimised by the same map there, but the tolerance criterion is free to reshape the radial profile and neither of them is.
Fig. 3 The share of each region’s area kept within ±0.5% of true scale by the three criteria, on four shapes. The gain runs from nothing on the sliver, where every map already keeps the whole of it, to 14.8 points on the square. A cap gains 4.9: its scale field is radial and the other two criteria are optimised by the same map there, but the tolerance criterion is free to reshape the radial profile and neither of them is.

Across shapes the gain tracks how much structure the region gives the fit to work with. A sliver is narrow enough that a twelve-term map holds all of it inside half a per cent whatever the criterion, so the criterion is idle. A square has corners, and corners are where the three maps disagree most.

Why a corner is where the three maps part

A cap gives the fit nothing to argue about, a square gives it a great deal, and the difference is worth saying in words because it is the whole of the shape dependence above.

On a cap the region is round and the scale field of every conformal map fitted to it is round too. All three criteria are then optimising one function of one variable — the radial profile — and the first two are optimised by the same profile.

A square’s boundary is at 25° from the centre at a corner and at about 19° along the middle of a side, so a map that is constant on the boundary must vary along any circle inside it. That variation is freedom, and freedom is what the three criteria spend differently: the worst case is set at the corners, the mean square is dominated by the sides because there is more area near them, and the tolerance band can be pushed out along the sides while the corners are abandoned. The 14.8 points the criterion buys on a square are bought almost entirely in the four wedges between the band’s edge and the corners.

What it pays

Each map is best on its own criterion and pays on the other two. The three maps of a 25° spherical square, each scored on all three criteria. The tolerance map keeps 60.1 per cent of the area within ±0.5% against 45.4 and 36.5, and it pays for that twice: its worst departure is 5.70 per cent against Chebyshev's 2.97, and its root-mean-square departure 1.110 against the mean-square map's 0.804. Each map wins the column it was fitted for and loses the other two, which is what makes these three criteria rather than three approximations to one.
Fig. 4 The three maps of a 25° spherical square, each scored on all three criteria. The tolerance map keeps 60.1 per cent of the area within ±0.5% against 45.4 and 36.5, and it pays for that twice: its worst departure is 5.70 per cent against Chebyshev’s 2.97, and its root-mean-square departure 1.110 against the mean-square map’s 0.804. Each map wins the column it was fitted for and loses the other two.

The table is the whole argument and it has no surprises in it, which is the point. Each map is best on the criterion it was fitted for and worse on both others — not slightly worse, in the case of the worst departure, but nearly twice Chebyshev’s. A surveyor who adopts this map has decided that a region a long way outside specification is no worse than a region just outside it, and the map takes them at their word: it doubles the worst case to buy the area.

Whether that is the right decision is not a question a measurement answers. What the measurement supplies is the exchange rate — 14.8 points of area for 2.7 percentage points of worst case on this region at this tolerance — and the average was a choice of norm is the standing statement that such a choice is the comparer’s and cannot be read off the geometry.

There is one case where the decision is not a matter of taste, and it is the case specifications are written for. If the tolerance is a legal limit — the grid shall be within one part in 2,500 — then ground outside it cannot be used at that grid whatever its scale factor, and the worst case is genuinely irrelevant: a parcel 0.08 per cent out and a parcel 0.30 per cent out are both surveyed on a different sheet. Where that is the situation, the third map is not a compromise between the other two but the only one of the three answering the question. Where the tolerance is a target rather than a limit, it is the worst of the three.

The pooled score abandons a region makes the neighbouring point about combining regions rather than combining criteria: an objective that adds two territories’ errors together will sacrifice the smaller one, and the sacrifice is invisible in the total. A tolerance criterion sacrifices the corners in exactly the same way and for the same reason — what is not counted is free to be spent.

The third map flattens the middle and lets the corner go. The departure from true scale along a ray from the centre of the square to a corner, for the three maps, with the tolerance band shaded between its two edges. Chebyshev's map leaves the ray at 0.56 per cent at the corner and the mean-square map at 2.00; the tolerance map reaches 3.29. What it buys for that is a longer flat stretch near the middle: it stays inside the band out to a larger share of the way to the corner than either rival, and everything past that point is not counted by the criterion at all.
Fig. 5 The departure from true scale along a ray from the centre of the square to a corner, for the three maps, with the tolerance band drawn. Chebyshev’s map leaves the ray at 0.56 per cent at the corner and the mean-square map at 2.00; the tolerance map reaches 3.29. What it buys for that is a longer flat stretch near the middle: it stays inside the band further out than either rival, and everything past that point is not counted by the criterion at all.

The profile shows the mechanism rather than the score. The tolerance map is flatter in the middle and steeper at the edge — it spends the series’ freedom on keeping a long stretch inside the band and then lets the corner go, because a corner outside the band by three per cent costs the criterion exactly what a corner outside by one per cent costs.

The three maps differ most exactly where the tolerance is drawn. For each of the three maps, the share of the square's area whose scale departs from true by less than the amount on the axis — so the height at 0.5%, marked, is the criterion itself. The tolerance map's curve is above the others at that mark by construction and crosses below them further out: it has bought area near the tolerance by sending a tail well past it. Chebyshev's curve reaches 1 first, because its worst case is the smallest; the tolerance map's reaches 1 last.
Fig. 6 For each of the three maps, the share of the square’s area whose scale departs from true by less than the amount on the axis — so the height at 0.5%, marked, is the criterion itself. The tolerance map’s curve is above the others at that mark by construction and crosses below them further out. Chebyshev’s curve reaches 1 first, because its worst case is the smallest; the tolerance map’s reaches 1 last.

The crossing is the clearest single picture of the trade. Read the three curves at the tolerance and the tolerance map wins; read them anywhere past about one per cent and it loses to both. A criterion that ignores a tail gets a tail.

Fitted on a sample, and scored off it

One objection has to be answered before any of these numbers can be quoted, and it has been met before. The objective is a count over a finite sample of the region, and a fit free to shape its scale field could in principle shape it to that sample’s own gaps — putting the band’s edge exactly between two rings of samples and collecting area that is not there. A condition imposed at points is not a condition is the standing instance of that failure here, and its own remedy is the one used here: measure on a set the fit did not use.

Scored on ground the fit never saw, the third map loses the least. Each map's share of the area within ±0.5% measured on the 6,400 samples the fit used, and again on 19,099 finer ones offset from them. The objective is a count over a finite sample, so a fit could in principle shape its scale field to that sample's own gaps. It has not: the tolerance map loses 0.63 points off the sample against Chebyshev's 1.68, which is the smallest of the three, and the ordering of the three is unchanged. Chebyshev's map loses most because its scale contour sits nearest the tolerance, so a small change in where the samples fall moves a large share across it.
Fig. 7 Each map’s share of the area within ±0.5% measured on the 6,400 samples the fit used, and again on 19,099 finer ones offset from them. The tolerance map loses 0.63 points off the sample against Chebyshev’s 1.68, which is the smallest of the three, and the ordering is unchanged. Chebyshev’s map loses most because its scale contour sits nearest the tolerance, so a small change in where the samples fall moves a large share across it.

The fitted map loses 0.63 points when scored on nineteen thousand samples it never saw — the smallest loss of the three, not the largest. Whatever it is doing with its extra freedom, it is not exploiting the sample; a map that were would show the reverse pattern, and the reverse pattern is what Chebyshev’s map shows for a quite different reason. Its scale contour lies near the tolerance over a long stretch, so a small change in where the samples fall moves a large share of area across the line.

The control, and it is a region where two of the three coincide

On a cap the first two criteria are one criterion, and the third is not. The same comparison on a spherical cap, where the scale field is radial and the worst-case and mean-square maps are the same map. Their two curves lie on one another to within 0.000 points at every tolerance, which is the control this whole comparison rests on: a construction that separated them here would be reporting itself. The tolerance map does separate, by up to 4.9 points — it is free to reshape the radial profile and neither of the other two is, so its gain on a cap is the criterion's own and not a difference of maps.
Fig. 8 The same comparison on a spherical cap, where the scale field is radial and the worst-case and mean-square maps are the same map. Their two curves lie on one another to within 0.000 points at every tolerance. The tolerance map does separate, by up to 4.9 points — it is free to reshape the radial profile and neither of the other two is.

A cap is the degenerate case and it is what makes the rest of this a measurement. Its scale field is radial, the worst-case and mean-square conditions are then optimised by the same map — a coincidence the two named criteria’s own comparison measured at 1.7 × 10⁻¹⁵ — and so their two shares must be equal at every tolerance. They are, to a thousandth of a point.

Had the construction separated them there, every number above would have been reporting the fitting arithmetic rather than the criteria. That the tolerance map does separate on a cap is not a failure of the control but the criterion’s own content: it reshapes the radial profile, piling area into the band and sending the rim out, which is a freedom the other two do not have because their objectives are minimised by the flattest radial profile there is.

What each number was checked against

The log-scale must be affine in the coefficients. Second differences of 10⁻¹⁷ against first differences of 10⁻⁵. Without this the exact line search is measuring a linearisation rather than the map, and every share above would carry an unstated error.

On a cap the first two criteria must keep exactly the same area, at 0.05 per cent and at half a per cent, and do so to within 0.005 of a point.

The fitted map must never keep less than either rival, at every tolerance tested. Both rivals are feasible starting points, so a search that came back below one of them would be reporting itself rather than the objective — which is exactly what an early version of this fit did, by sweeping the free level of the scale field last instead of first.

It must be beaten on at least one rival criterion, or it would not be a third criterion but a better solution to one of theirs. It is beaten on both.

The fit must not be exploiting its own sample. Every map is re-scored on 19,099 samples offset from the 6,400 it was fitted on. The tolerance map loses 0.63 points, the mean-square map 0.72 and Chebyshev’s 1.68, and the ordering of the three is unchanged.

And a tolerance of 50 per cent must be kept in full by all three, which is the other end of the same control: with no tail to ignore, a criterion that ignores tails has nothing to say.

What this fit does not settle

The objective is a step, so it has no gradient. Every fitting method used on a region’s condition elsewhere solves a linear system in one shot, because the criterion was a least square. This one cannot be: the derivative of the objective is zero almost everywhere and undefined on a set of measure zero, which is why the fit is a search rather than a solve. That is not a defect of the implementation; it is the criterion.

Twelve terms, one region size, one boundary. Every number is for a twelve-term series on a 25° region. A longer series would raise all three shares and it is not obvious it would raise them equally — the tolerance criterion has more to do with the extra freedom than the other two, so the gain quoted here may be a floor.

The tolerance is symmetric about true scale. A grid specification is often written as a one-sided or an asymmetric band, and the whole construction takes a two-sided one. Nothing in the method needs the symmetry; the figures do not test it.

Area is the weight. The criterion counts a square kilometre of empty ground as heavily as a square kilometre of city, which is the unexamined weighting every objective on this page shares.

And the share is the best this search found. The objective is not convex, so none of these figures is a proof that no map does better. That is not a caveat to be filed away: it is large enough, and structured enough, to be the whole of the next measurement.

Still open: how much of the answer is where the fit began

The fit above was run from two starts and the better was reported, which is what an honest search does and is also an admission. A convex objective would make the choice of start irrelevant and there would be nothing to admit.

This one is not convex — the band’s indicator is not a concave function of the coefficients, and nothing in the construction makes the composition with an affine map concave either. So the fit has local optima, and the number reported is the one some particular search reached from some particular beginning.

How many optima this region and this tolerance actually have, how far apart the answers they give are, whether the two starts anybody would naturally use — Chebyshev’s map and the mean-square map — reach the same place, and whether a straight line between two of the answers really runs downhill from both ends, are questions a search that reports its best cannot ask of itself.

The shape of that enquiry is already familiar from the other side of this field. The landscape the search walks on built the surface an aspect search moves over and found it neither a valley nor a basin, and where the valley breaks in two measured when it fractures. Those are searches over three numbers. This one is over twenty-five, and the objective is a step rather than a smooth score, so there is no reason to expect the answer to be gentler.

What this makes readable

Essays that name this one as a prerequisite.

Named alongside this one

Essays reaching for the same objects. Nobody chose these; they are what the concept index makes visible.

What links here

Every essay whose body links to this one.

The objects this essay names

Each one links to every other essay that touches it.

Chebyshev's criterionConformalityObjective functionOptimal conformalOptimisationPurposeRegionScale factorToleranceVerification