What each projection optimises

The basins have widths as well as depths

The previous rung measured the height of the pass and recorded a shortfall: shape means widths too. Measured, the basin has three of them — 32°, 16° and 7° at Japan — it gets wider rather than narrower as the region grows, and the exponent it predicts overshoots the measured one by half again.

Assumes Where the valley breaks in two.

The height of the pass between two basins measured the saddle directly, found it did not explain the exponent it was supposed to explain, and recorded a shortfall in one sentence: the shape of the aspect score surface has widths in it as well as heights, and the sweep that finds the pass was computing the widths on its way past and throwing them away.

This rung reads them out. Three things come back, and only the first was expected.

The basin has three widths, and they differ by a factor of 4.6. Two sections through the near-optimal basin of the Robinson aspect over Japan, drawn at one scale. Each ellipse is the set of aspects whose score is twice the optimum's, from the objective's own second derivative at the optimum: 32.1°, 16.0°, 6.9° along the three principal directions. A single number for "the width of the basin" is the cube root of their product, 15.3°, and it is not any of them.
Fig. 1 Two sections through the near-optimal basin of the Robinson aspect over Japan, drawn at one scale. Each ellipse is the set of aspects scoring twice the optimum. The basin has three principal widths — 32.1°, 16.0° and 6.9° — and no single number for “the width” is any of them.

What a width is here, and how it is got

The objective is a score over the three angles an aspect is: where the rotated pole goes, in longitude and latitude, and how far the page is then turned.

Near a minimum a smooth function is a quadratic form, so its second derivative at the optimum carries the whole local shape. The width along a principal direction is the displacement at which the score rises to twice the optimum’s — √(2 f₀/λ) for the eigenvalue λ in that direction — and it is in degrees of rotation, because that is what the parameters are.

Nineteen evaluations of the objective give the three-by-three second derivative, and its eigenvalues give three widths.

That is a great deal cheaper than the alternative, and the alternative was tried first. Reading the width off the sublevel set — sweeping the score in ascending order, unioning cells as they arrive and taking the cube root of the component’s volume — is the natural thing to do with machinery that already exists. On the 24 × 13 × 12 grid the whole ladder runs on, the near-optimal component is a single cell at every threshold up to half the fracture, so the cube root is a measurement of the grid spacing. Refining the grid enough to fix that is thirty seconds a region, six regions a sweep, and the answer would still be quantised.

The first finding: it is not round

The basin is never round, and how far from round is not steady. The ratio of the widest to the narrowest principal width, region by region. It runs from 4.04 to 11.28, so a search that steps the same distance in all three of its parameters is stepping across the basin in one direction and along it in another. The figure beside each bar is the single width the three collapse to.
Fig. 2 The ratio of the widest principal width to the narrowest, region by region. It runs from 4.04 to 11.28 across the sweep, so a basin is between four and eleven times longer in one direction than another, and the figure beside each bar is the single width the three collapse to.

A search that steps the same distance in all three of its parameters is therefore stepping across the basin in one direction and along it in another, by up to a factor of eleven. That is not a small remark about tuning. It is the reason the parameters a search reports are not reproducible while the map is: along the widest direction the objective barely changes, so where the walk stops is decided by its step schedule rather than by the surface, and that direction is eleven times longer than the direction the surface actually constrains.

The three widths are the quantitative version of what that rung established qualitatively, and they say which direction is the loose one.

The second finding, and it is the wrong way round

The plan for this rung assumed the basin narrows as the region grows. A larger region sees more of the projection’s own variation, so it should care more about where the aspect is put, so the near-optimal set should shrink.

It does the opposite.

The basin gets wider as the region grows, which is the wrong way round. The three principal widths of the near-optimal basin, over five regions of one shape at growing sizes. Every one of them is larger at the largest region than at the smallest — 21.9°, 11.5°, 5.2° at 6° against 103.4°, 48.7°, 25.6° at 40° — and only the narrowest rises monotonically; the widest peaks in the middle of the range. The plan for this rung assumed all three would fall, on the grounds that a larger region sees more of the projection's variation and therefore cares more about the aspect. It cares more in absolute terms and less in relative ones, and the width is measured where the score doubles.
Fig. 3 The three principal widths over six regions of one shape at growing sizes. Every one of them is larger at the largest region than at the smallest — 21.9°, 11.5° and 5.2° at a six-degree region against 103.4°, 48.7° and 25.6° at a forty-degree one. Only the narrowest rises monotonically; the widest peaks in the middle of the range.

The mechanism is in the definition and is not a defect of it. A width is measured where the score doubles, and what a larger region does is raise the whole surface: the best achievable distortion over a forty-degree region is far worse than over a six-degree one. The floor rises faster than the curvature does, so twice the floor is further away.

That is worth stating as a general caution rather than as a local fact. A relative threshold on an objective whose floor is moving measures the floor as much as the shape. Any near-optimal set defined as “within so many per cent of the best” inherits it, and this collection has now met the same trap in the tolerance that decides the verdict, where a fixed tolerance was found to be measuring the arithmetic’s noise floor rather than the map.

The width follows a power of the region's sharpness, at 1.09. The cube root of the three widths against the fraction of the projection's variation each region sees, on logarithmic axes, with the fitted line. The exponent is 1.086 at a coefficient of determination of 0.9441 — the basin widens, and it widens close to linearly. The fracture threshold over the same regions falls at -1.479, so the two move in opposite directions, which is what the quadratic model requires and is the only part of it that survives.
Fig. 4 The cube root of the three widths against the fraction of the projection’s own variation each region sees, on logarithmic axes, with the fitted line. The exponent is +1.086 at a coefficient of determination of 0.944 — the basin widens, and it widens close to linearly.

The threshold the widths are being compared against

The quantity the widths have to explain is the one where the valley breaks in two measured, and it is worth having it on the page rather than only in a fitted exponent.

Where Robinson's valley breaks, against how large the region is. The threshold at which the set of near-optimal aspects stops being one connected piece, for square regions of growing size at 38° north. It falls from 2.13 at 6° to 0.27 at 30°, a factor of 8.0. The previous rung measured this at one region and quoted "about twice the optimum"; that value belongs to a small region, and the prediction that it should fall as the region grows is what this tests.
Fig. 5 The fracture threshold against the fraction of the projection’s variation each region sees, over the same regions. It falls from 2.13 at a six-degree region to 0.27 at a thirty-degree one: a large region’s near-optimal set comes apart at a much tighter tolerance than a small one’s, because a large region can tell its aspects apart.

Two numbers now exist for the same six regions, measured by routes that share nothing: a threshold from a union-find sweep over a grid of scores, and a width from nineteen evaluations of the objective at a single point. Neither knows anything about the other, which is what makes comparing them a test rather than a restatement.

The third finding: the prediction overshoots

The reason the shortfall was recorded at all is that the width was the suspect for a missing exponent. The argument is two lines.

The set below a relative threshold t reaches a distance w₀√t from the optimum. Two near-optimal pieces become separate answers when their sets touch, so if the distance between them were fixed, w₀√t* would be a constant and

tw02.t^* \propto w_0^{-2}.

The width exponent is +1.086, so the fracture exponent should be −2.172. Over the same regions it measures −1.479.

The width does not close the gap — it opens one on the other side. Three exponents for the same quantity. The argument that produced the fracture threshold predicts one. The measurement over four regions gives 1.48. The basin width, which the previous rung recorded as the suspect, predicts 2.17 — past the measurement rather than up to it, by 0.69. The shortfall is paid and the answer is that the width over-explains.
Fig. 6 Three exponents for one quantity. The argument that produced the fracture threshold predicts one. The measurement gives 1.48. The basin width — recorded as the suspect — predicts 2.17, which is past the measurement rather than up to it.

The width does not close the gap. It opens one on the other side, and by more than the original: the argument was short by 0.48 and the width overshoots by 0.69.

That is a real result and it is worth being plain about why it is not a disappointment. The previous rung’s suspect has been measured, and measuring it has ruled it out as the explanation while establishing that it is a large part of the mechanism — the prediction moved from 1 to 2.17 and the truth is between them. What is left over is now a smaller and much better-specified question than “why is the exponent not one”.

Why it overshoots, which the same sweep answers

The quadratic argument has two quantities in it and only one of them has been measured. The other is the distance between the two pieces, which the argument assumes is fixed by the parameterisation.

There are not two pieces.

It does not break in two — it crumbles. The number of separate pieces the near-optimal set is in, just below its own fracture threshold, region by region. The two-basin picture that the quadratic prediction rests on needs this number to be two. It is between 4 and 14. Well above the fracture — three times the threshold — the same measurement returns one piece on every region, so the count is about the objective rather than about the grid. What the threshold marks is a percolation rather than a merge, and a percolation has its own exponent.
Fig. 7 The number of separate pieces the near-optimal set is in, just below its own fracture threshold. A two-basin model needs two. It is between four and fourteen. Well above the fracture — at three times the threshold — the same measurement returns one piece on every region, so the count is about the objective rather than about the grid.

At the threshold the near-optimal set of the Robinson aspect over Japan is in twelve pieces. Where the valley breaks in two already knew this — it reports one connected sheet at a loose threshold and fourteen basins at a tight one — and the two-basin language it used, which this rung inherited, is a description of the first thing that happens rather than of the thing that is being measured.

So the fracture threshold is not the merging of two basins at a fixed separation. It is a percolation: the threshold at which a set of many small components first connects into one. Percolation thresholds have their own scaling, set by the dimension and by the correlation structure of the field, and there is no reason for it to be the two-body exponent −2.

That reframing costs nothing already established. Everything the ladder has measured — the threshold, its region dependence, the pass height, the widths — is unchanged. What changes is which model those numbers should be fed to.

What the separation actually does

It is worth showing why the two-piece model could not have been rescued by measuring its separation instead of assuming it.

Taken just under the fracture, the distance between the two lowest pieces comes back at 151°, 151°, 122°, 66.5° and 216° across the sweep, with no order in it. Fed into the same quadratic model those give predicted thresholds of 47.5, 15.7, 0.68 and 5.9 against measured values of 2.13, 1.33, 0.69 and 0.27 — agreement at one region out of four and disagreement by a factor of twenty at another.

The numbers are not noisy because the measurement is poor. They are meaningless because “the two lowest pieces” is not a well-defined pair when there are twelve, and which two the sweep happens to find depends on the grid.

What is left of the two-line argument

It is worth separating what the crumbling destroys from what it leaves standing, because it is less than it looks.

The relation between the width and the threshold survives: the two move in opposite directions, at fits of R² 0.944 and 0.948, and they must, because a wider basin reaches its neighbours at a lower threshold whatever the neighbours are. That is asserted here and it is the part of the quadratic model that does not depend on how many pieces there are.

What does not survive is the factor of two. It comes from the square in w₀√t, which is a statement about a single pair of bodies approaching each other, and a percolation is not that. Twelve components arriving at a common level is a connectivity problem on a graph, and its exponent depends on how the field’s values are correlated between neighbouring cells rather than on the local curvature at one point.

So the honest reading of this rung is: the width is measured, it is a large part of the mechanism, and the exponent it predicts is wrong for a reason that the same sweep can see and that the ladder’s own previous rung had already reported without noticing what it was reporting.

What a search should do with all this

Three consequences, and the first two are immediately usable.

Step anisotropically. The three widths are known, cheaply, from nineteen evaluations at the current best point. A compass walk that scales its step to each width converges in the same number of steps in every direction instead of grinding along the loose one — and the loose one is up to eleven times longer.

Do not report the parameters. With a width of 32° along the loosest direction, two searches that agree about the map to a part in a thousand can report pole positions half a world apart, which is exactly what the reproducibility rung found and could not previously explain in one number.

And treat the threshold as a percolation. A search that wants to enumerate the genuinely distinct near-optimal aspects should not look for the height at which two basins meet; it should look for the height at which the largest component stops growing faster than the number of components falls, which is a different reading of the same sweep.

Three numbers rather than one

What this rung leaves the ladder with is a short list, and it is worth setting out because the next rung has to start from it.

A depth — how good the best aspect is, which is what every search reports.

Three widths — how well determined it is, direction by direction, from nineteen evaluations of the objective at the point the search stopped.

And a piece count — how many genuinely different answers there are at a stated tolerance, which is what the sweep produces and what nobody reads out of it.

The three are independent, they are all cheap, and only the first is ever printed. A search that reported all three would say what it found, how precisely, and whether it was the only one — which is the whole of what a reader of an optimisation result needs and is more than any published aspect search supplies.

Where the model stops

The Hessian is taken at 2.5° in each of three angles, which is coarse against a narrowest width of 5.2° at the smallest region — so the second difference there is averaging over a third of the basin. Halving the step moves the widths by under two per cent on every region tested, which is the reason the coarse step is kept, but the smallest region is the one to distrust.

One of the six regions produces no Hessian at all. At a twenty-degree span the search’s stopping point has a negative curvature in one direction, so it is on a ridge rather than in a bowl, and a local polish before the differencing does not rescue it. The row is dropped rather than repaired, and dropping it is why the fits below use four points rather than six.

The parameterisation is degenerate at a pole latitude of ±90°, where a change in pole longitude is a change in the third angle. None of the optima found here is within twenty-five degrees of that, and a region whose optimum was would need the widths read in a different chart.

The generalisation

Strip out the cartography and what is left is a statement about reading an optimisation landscape.

A basin has a depth, three widths and a shape, and every summary of it as one number throws away the part that decides how a search behaves. The depth says how good the answer is. The widths say how well determined it is, direction by direction. The anisotropy says whether stepping uniformly is sensible. And the number of components at a threshold says whether “the basin” is a thing at all.

The site’s own instrument is doing the same work here that it does on projections: the second derivative of an objective is the same kind of object as the second derivative of a map, and the flexion ladder is the essay about what a first-order description leaves out. An optimisation landscape described by its minimum alone is Tissot’s ellipse with the ellipse left off.

Who found it, and when

Reading a basin’s shape off the Hessian is the oldest idea in numerical optimisation and is the whole content of Newton’s method: a step scaled by the inverse second derivative is a step that treats every direction alike. That the eigenvalues of the Hessian are the reciprocal squares of the widths is a restatement, and quoting them as a condition number is standard.

Percolation as the thing that happens when a level set of a random-looking field first connects is younger — the theory dates from the late 1950s — and its arrival here is a consequence of the measurement rather than a hypothesis brought to it. Nobody chose to look for a percolation threshold; the piece count refused the alternative.

The width is the precision the answer should be printed to

There is a small use for these numbers that requires no further work and that the field gets wrong routinely.

A basin’s width is the distance over which the objective does not meaningfully change. So it is exactly the resolution at which the optimum is determined: two parameter values inside one width are two names for one answer, and reporting the difference between them is reporting noise.

That gives a rule for how many digits an optimal aspect should be published with. If the basin is three degrees across in one direction, the optimum in that direction is known to about a degree, and a paper reporting it to four decimal places is asserting a precision the landscape does not contain. The extra digits are a property of where the optimiser happened to stop.

And the widths are not equal, which is the first finding above, so the precision is not the same in every direction. An answer quoted to one degree in the narrow direction and one degree in the wide one is over-precise in one and under-precise in the other, and the honest report gives a different number of digits to each — or, better, gives the widths beside the optimum and lets the reader see the shape of what was found.

This is the cheap half of reporting the map rather than the parameters, and it is available to anybody whose optimiser can be run three more times. The expensive half is establishing that the parameters are identifiable at all; the cheap half is refusing to print digits the objective cannot support.

Where the ladder goes next

The exponent is now bracketed by two arguments rather than approached by one, and neither is right. The next thing to measure is the piece-count curve itself — how the number of components rises and falls through the threshold — because that curve has an exponent of its own and it is the one the percolation model actually predicts.

What this makes readable

Essays that name this one as a prerequisite.

Named alongside this one

Essays reaching for the same objects. Nobody chose these; they are what the concept index makes visible.

What links here

Every essay whose body links to this one.

The objects this essay names

Each one links to every other essay that touches it.

AnisotropyAspectAspect searchBasinExponentFracture thresholdHessianLevel setOptimisationOptimisation landscapePercolationShortfall