What each projection optimises

A criterion worth using is one whose answer is not unique

Forty-two starts on one region and one tolerance give forty-two different maps, keeping from 36.70 per cent of the area to 65.15. The two starts anybody would actually use — Chebyshev's map and the mean-square map — end 22.7 points apart, which is more than the criterion buys over either of them. And the spread is widest at exactly the tolerance where the criterion is worth using at all, because both facts are the same fact: an objective with a flat top and a sharp edge has somewhere to hide answers.

Assumes The map that keeps the most ground inside a tolerance.

The map that keeps the most ground inside a tolerance fitted a third criterion and reported what it bought: 14.8 points more of a spherical square inside half a per cent of true scale than the mean-square map manages, bought by doubling the worst departure. Every one of those numbers was qualified the same way — the best this search found — and the qualification was not modesty. The objective is a step function of the departure integrated over area, and a step is neither concave nor convex, so nothing in the construction promises one answer.

This measures how much that matters, and the answer is that it matters more than the criterion’s own gain.

Forty-two starts, forty-two answers. Every local optimum coordinate ascent reaches from 42 starts on the same region and the same criterion — a 25° spherical square at a tolerance of ±0.5% — plotted at the share of area it keeps. They run from 36.70 per cent to 65.15, a spread of 28.4 points, and no two of them are the same map: clustered on their scale fields rather than their coefficients, 42 of the 42 are distinct. The two marked starts are the ones anybody would actually use — Chebyshev's map reaches 37.43 per cent and the mean-square map 60.14, and neither is the best this search found.
Fig. 1 Every local optimum coordinate ascent reaches from 42 starts on the same region and the same criterion — a 25° spherical square at a tolerance of ±0.5% — plotted at the share of area it keeps. They run from 36.70 per cent to 65.15, a spread of 28.4 points, and no two are the same map: clustered on their scale fields rather than their coefficients, 42 of the 42 are distinct. Chebyshev’s map reaches 37.43 per cent and the mean-square map 60.14, and neither is the best this search found.

Forty-two starts, forty-two answers

The starts are not noise. Each one is a point drawn somewhere between Chebyshev’s solution and the mean-square solution and then perturbed by up to the distance between them — so every start is a conformal map of the same region fitted by the same series, and the search is being asked which of them it prefers rather than being handed nonsense.

It prefers all of them. Forty-two ascents give forty-two local optima, and the clustering that says so is done on the scale field rather than on the coefficients: two maps counted as different differ by more than 10⁻⁴ in lnk\ln k somewhere in the region, which is a difference a reader could find with a scale bar rather than one hidden in the fit’s parameterisation.

The distinction matters because a series fit has reparameterisations that change its coefficients without changing its map, and a count of distinct coefficient vectors would count those as separate answers. A hundredth of the tolerance being fitted is the separation used, which is generous: two maps this account calls different disagree about the scale somewhere by a hundredth of the band whose contents are the whole objective.

The range is the thing to carry. The worst optimum keeps 36.70 per cent of the area and the best 65.15 — a factor of 1.77, on one region, at one tolerance, from one objective.

For scale, the worst of the forty-two is barely better than Chebyshev’s map, which keeps 36.47 and was fitted to a quite different criterion. A search for the tolerance map that landed there would have done a great deal of work to arrive where the oldest of the three criteria already was.

Three of the answers, drawn — they are genuinely different maps. The tolerance contours of the best local optimum the search found, one from the middle of the range, and the worst, on the same square at the same tolerance. They keep 65.2, 49.5, 36.7 per cent of the area. The point of drawing them is that the spread is not a fit wobbling about one answer: the enclosed regions have different shapes, and the clustering that counted 42 distinct optima was done on the scale fields themselves rather than on the coefficients, so two maps counted as different are different everywhere a reader could look.
Fig. 2 The tolerance contours of the best local optimum the search found, one from the middle of the range, and the worst, on the same square at the same tolerance. They keep 65.2, 49.5 and 36.7 per cent of the area. The enclosed regions have different shapes, so the spread is not a fit wobbling about one answer.

Drawn, they are obviously three maps and not three readings of one. The best pushes its band out along the sides into four long lobes and abandons the corners entirely; the worst holds a rounder band and gives up the sides sooner.

It is not the search wobbling

Two explanations for a spread like that have to be ruled out before it can be attributed to the objective, and both are ruled out by construction rather than by argument.

Every ascent converges, and they converge to different places. The share kept after each sweep of coordinate ascent, for the two named starts and 9 of the seeded ones. Every trail is monotone and every one settles — the search is not failing to converge, which would be the other explanation for the spread. They settle at 37.3 to 60.1 per cent. Because each line search along a coordinate is exact rather than sampled, a trail that has stopped moving has reached a point no single-coordinate move can improve, which is what a local optimum of this objective is.
Fig. 3 The share kept after each sweep of coordinate ascent, for the two named starts and 9 of the seeded ones. Every trail is monotone and every one settles. They settle at 37.3 to 60.1 per cent. Because each line search along a coordinate is exact rather than sampled, a trail that has stopped moving has reached a point no single-coordinate move can improve.

The first explanation is that the ascents have not finished. They have: every trail is monotone and flat by its last sweep, and the search stops when a whole sweep moves nothing.

The second is that the line search is too coarse to find the improvement. It cannot be. Because lnk\ln k is exactly affine in the coefficients, a sample is inside the tolerance on an interval of the step along any direction, and the endpoints of all those intervals are the roots of linear equations. Sorting them and sweeping gives the exact best step. There is no step size in this fit, so there is no step size to blame.

What is left is that these are genuine local optima of the objective, and the third measurement settles that directly.

The line between two answers runs downhill from both ends

The line between two answers runs downhill from both ends. The criterion evaluated along the straight line in coefficient space between two of the local optima above. The two ends keep 37.26 and 36.71 per cent of the area; the worst point between them keeps 31.45, which is 5.26 points below the lower end. That settles it rather than suggesting it: a convex objective's superlevel sets are convex, so the segment joining two maps that each keep more than some share could never leave that set. This one does, so the criterion is not convex and the several answers above are a property of it rather than of the search.
Fig. 4 The criterion evaluated along the straight line in coefficient space between two of the local optima. The two ends keep 37.26 and 36.71 per cent of the area; the worst point between them keeps 31.45, which is 5.26 points below the lower end. A convex objective’s superlevel sets are convex, so the segment joining two maps that each keep more than some share could never leave that set. This one does.

This is the proof rather than the evidence. If the objective were concave — the property that would make a maximum unique and a hill-climb reliable — then the set of maps keeping at least some share would be convex, and the straight line between two members of it could not leave it. The line between these two leaves it by 5.26 points.

So there is nothing to fix. The several answers are a property of what was asked for, and any method whatever will have to choose among them: a different optimiser would find a different subset, not a single right one.

The reason is visible in the criterion’s own shape. An objective built from a step has a flat top — inside the band, moving a point further from true scale costs nothing — and a sharp edge, where moving it across the line costs its whole area at once. A flat top means a fit can drift a long way without the score changing, and a sharp edge means the score can be improved by a rearrangement that a smooth penalty would have refused. Between them they make a landscape of plateaux separated by cliffs, which is exactly a landscape with many tops.

It is worth contrasting that with the two criteria fitted before. Both are least squares: a sum of squared quantities, convex in the coefficients, with a unique minimum a single linear solve reaches. The nearest equal-area map to an impossible request is one such fit and reports a residual rather than a choice of answer. The step objective has no such guarantee and no such solve, and everything above follows from the difference.

A landscape of plateaux and cliffs, and what climbs it

The shape argued above is worth drawing in words because it explains both why the search works well and why its answer is not unique.

Inside the band, moving a sample further from true scale changes the score by nothing. Outside it, likewise. The score changes only when a sample crosses the band’s edge, and then it changes by that sample’s whole share of area at once. So the objective as a function of the coefficients is constant on open regions and jumps between them: a staircase in twenty-five dimensions, whose treads are the sets of coefficient vectors that put the same samples inside.

Two consequences follow. A gradient method is useless, because the gradient is zero wherever it is defined. And a method that works on the exact structure — jumping directly to the best tread along each direction, which is what the line search here does — is not merely a good heuristic but the natural algorithm: it never takes a step that does nothing, and it always lands where the score is highest along that line rather than somewhere on the way.

What neither of those gets round is that a staircase has many local tops. Climbing from any tread, one eventually reaches a tread with no higher neighbour along any coordinate, and which one that is depends on where the climb started.

The shape of the valley met the same question from the other side of this field, on the surface an aspect search walks over: it was called a valley and turned out, sampled densely, to be neither a valley nor a basin but a connected sheet that fractures. The threshold is not a percolation then removed one explanation of that fracture without supplying another. The landscape here is a different object — a step objective rather than a smooth score, twenty-five parameters rather than three — but the lesson those measurements arrived at holds: a search’s answer is a statement about the surface it walked on, and the surface has to be measured rather than assumed.

The spread is widest where the criterion is worth using

The answers spread widest where the criterion is worth using. The best and the worst local optimum the same search finds, against the tolerance. The band between them is how much of the answer is the starting point. It is 8.3 points at a tolerance of 0.05 per cent, widens to 27.8 at half a per cent, and closes to nothing at two per cent where every map keeps the whole region. The widest part of the band sits at exactly the tolerance where the criterion buys the most over its rivals, so the two headline numbers are the same measurement read twice: a criterion worth using is a criterion whose answer is not unique.
Fig. 5 The best and the worst local optimum the same search finds, against the tolerance. The band between them is how much of the answer is the starting point. It is 8.3 points at a tolerance of 0.05 per cent, widens to 27.8 at half a per cent, and closes to nothing at two per cent where every map keeps the whole region.

The band is narrow at a tight tolerance, widest in the middle and zero at a loose one — and that is the same curve, upside down, as the gain the criterion buys over its rivals. At a tolerance tight enough that almost nothing is inside, every map is doing the same thing and there is one answer. At a tolerance loose enough that everything is inside, every map keeps all of it and the answer is unique as a number however many maps attain it. In between, where the criterion has something to say, it says several things.

That is not a coincidence and it is the title of this measurement. The freedom that lets a criterion prefer a different map from its rivals is the same freedom that lets it prefer several maps to each other. A criterion whose answer is always unique is one that is not choosing.

Which map the fit is started from decides most of what it reports. The same coordinate ascent begun from Chebyshev's map and from the mean-square map, against the tolerance, with the best optimum any start reached drawn above them. At half a per cent the Chebyshev start ends at 37.4 per cent and the mean-square start at 60.1, a gap of 22.7 points — larger than the 14.8 the criterion buys over the mean-square map in the first place. A fit reported without saying where it began is a number with a free parameter in it, and this is how large that parameter is.
Fig. 6 The same coordinate ascent begun from Chebyshev’s map and from the mean-square map, against the tolerance, with the best optimum any start reached drawn above them. At half a per cent the Chebyshev start ends at 37.4 per cent and the mean-square start at 60.1, a gap of 22.7 points — larger than the 14.8 the criterion buys over the mean-square map in the first place.

The practical form of all this is the last figure, and it is the one a reader should take away. The two starts anybody would actually use are the two maps already fitted, and they end 22.7 points apart — more than the 14.8 the whole criterion is worth. A fitted tolerance map quoted without saying where the fit began is a number with a free parameter in it, and the free parameter is larger than the effect being reported.

There is a cheap partial remedy and it is what the fit that found the criterion’s gain did without comment: run from both named starts and keep the better. That recovers 60.1 of the 65.2 the wider search found, for twice the work, and it is reproducible — two people following the same recipe get the same map. What it does not do is find the best, and nothing here suggests 65.2 is the best either.

What this does to a published grid

A grid is not a private calculation. A published coordinate is a result makes the standing point that once a coordinate is issued it is a legal object rather than a best estimate, and a projection definition is the same kind of thing: it is published as parameters, and every coordinate computed from it afterwards depends on those parameters exactly.

So a criterion whose answer is not unique is a problem of a particular kind. It is not that the map is uncertain — each of the forty-two is a perfectly definite conformal map, computable to as many figures as anybody wants. It is that the specification does not pick one, so two competent people fitting the same stated criterion to the same stated region can publish different grids, and both are right.

That is worse than an ordinary numerical disagreement, because there is nothing to reconcile. Neither party has made an error. The difference is 22.7 points of area inside the tolerance, and the only thing that distinguishes the two answers is a choice — the starting map — that the specification never mentioned and that neither party is likely to have recorded.

Every projection minimises something is a catalogue of objectives, and the catalogue is written as though naming the objective named the map. For the two smooth criteria fitted here, it does: Chebyshev’s map is the best at its worst, and not on average compares two maps each determined by its condition. For this one, naming the objective is not enough, and what has to be named as well is the whole recipe.

What a practitioner should actually do

The measurements above are a complaint, and a complaint without a procedure is not much use. Three procedures are available and they are not equally good.

Run from the two named starts and keep the better. Reproducible, cheap, and it recovers 60.1 of the 65.2 the wider search found — 92 per cent of the available gain over the mean-square map. Anybody following the recipe gets the same map, which is the property a published grid most needs.

Run from many starts and keep the best. Better by about five points of area and not reproducible: the answer depends on the seed, the count and the perturbation scale, none of which is part of the criterion. A grid published this way would have to publish the seed to be reproducible at all.

Publish the criterion and the recipe together. Which is the honest version of the first, and is what the whole of this measurement argues for: the objective does not determine the map, so a definition that names only the objective is incomplete in a way that can be quantified at 22.7 points of area.

The third is not a compromise between the other two. It is the recognition that for a non-convex criterion the search is part of the specification, and leaving it unstated does not make it absent.

What each number was checked against

At a tolerance containing the whole region every start must reach the same value. With no tail to ignore there is nothing for a step objective to trade, and the twelve starts tested agree to better than 10⁻¹² of a point. A search reporting a spread there would be reporting its own failure to converge rather than a property of the objective, and every number in this essay would be unreadable.

The distinctness is measured on the maps, not on their parameters. Two optima count as one when their log-scale fields agree everywhere to 10⁻⁴. Clustering on coefficients instead would have counted reparameterisations of one map as several.

At a tolerance that bites, the achieved values must genuinely spread, and at half a per cent they spread by 27.8 points.

And the objective must be shown not convex rather than inferred to be. The straight line between two optima must dip below the lower of its two ends; the deepest such line found dips 5.26 points. Finding several optima is consistent with a convex objective and a broken search — the dip is not.

What forty-two starts do not establish

The best found is not the best. Nothing here is an upper bound on what a twelve-term conformal map of this square can keep inside half a per cent. The search is coordinate ascent, which cannot move diagonally, and a method that could would presumably find more.

The starts are drawn from one line. Every seeded start lies near the segment between the two named solutions, perturbed by the distance between them. That is deliberate — it keeps every start a plausible map rather than noise — and it means the forty-two optima are the ones reachable from that neighbourhood, not a sample of the whole space.

The perturbation scale is a choice. Each seeded start is displaced by up to the coefficient-wise distance between the two named solutions. A smaller perturbation would give starts closer together and, presumably, fewer distinct optima; a larger one would wander into maps nobody would fit. The scale was chosen to be the one distance this problem supplies of itself rather than tuned, and the count depends on it.

The clustering threshold is not the only choice in the count. Two optima also count as one when the ascent from one reaches the other, which never happens here because each ascent terminates. A method that restarted from a perturbed optimum would merge some of the forty-two, and the shape of the valley is the neighbouring case where exactly that kind of re-sampling changed a valley into a fractured sheet.

One region, one tolerance, twelve terms. The 27.8-point spread is for a 25° square at half a per cent. The shape of the curve against tolerance is measured; the shape against region and against series length is not.

The dip proves non-convexity and not multimodality in any particular direction. A dipping segment shows the superlevel set is not convex, which is exactly what rules out the guarantee a concave objective would give. It does not by itself say how many tops there are; the count comes from the starts, and the count is what it is.

And the count of distinct optima is a lower bound with a threshold in it. Forty-two starts gave forty-two distinct fields at a separation of 10⁻⁴ in lnk\ln k. A tighter threshold could only raise the count and a looser one lower it, and there is no natural threshold — 10⁻⁴ is a hundredth of the tolerance being fitted, which is the reason it was chosen and not a reason it is right.

Still open: whether a specification can be written that has one answer

The awkwardness here is not really about optimisation. It is that a specification written as a band asks for something that does not determine a map, and a surveyor handed two maps that both meet the specification over 60 per cent of the area has no way to prefer one from the specification alone.

It is the same complaint a condition imposed at points is not a condition made about collocation, one level up: a requirement that under-determines its object is not the requirement it appears to be, and the gap is filled silently by whatever the method happens to do.

That suggests a different question from which map maximises the area. A specification could be written to pick out a unique map — most simply by keeping the band and breaking ties with one of the two smooth criteria, so that among all maps keeping the most area the one with the smallest worst case is meant. Whether such a lexicographic rule has a unique answer, whether the tie-break is ever actually binding or whether the optima found here differ in their worst cases anyway, and how much area a rule that is reproducible gives up against a rule that is merely best, are questions a single objective cannot ask.

Named alongside this one

Essays reaching for the same objects. Nobody chose these; they are what the concept index makes visible.

The objects this essay names

Each one links to every other essay that touches it.

ConvergenceDegrees of freedomObjective functionOptimal conformalOptimisationRegionReproducibilityScale factorToleranceVerification