Weighted by what it disagrees with, the fix trusts the sight it leans on
Assumes A low sight is worth keeping only if it is weighted.
A low sight is worth keeping only if it is weighted gave each sight the error its altitude earns. A body five degrees above the horizon is seen through a long, layered column of air, its refraction correction is ten times a high body’s, and a tenth of that correction being uncertain makes the sight more than twice as noisy. Weighted by its variance, such a sight always helps the fix. Unweighted, below seven degrees it makes the fix worse than throwing it away.
The weights in that account came from outside the sights: from an altitude and a stated share of a correction. That share is the weak point. It depends on the temperature gradient over the water, which nobody on deck knows, and the earlier account could only sweep it from a fiftieth to a half and say which way a wrong guess fails.
A navigator with more sights than unknowns has another source of weights, and it looks like it should be better, because it comes from the sights themselves. With six lines of position and two unknowns there are four degrees of freedom left over, and the residuals say which lines disagree with the rest. So: solve, weight each line by the inverse square of its own residual, solve again, repeat until the weights settle. It is the iteratively reweighted least squares that robust estimation uses to push down an outlier, and it asks nothing about the weather.
It fails, and the way it fails is the reason it is worth measuring. It does not converge on the altitude weights, or anywhere near them. Where the geometry gives one sight a special position, it converges on the opposite of them.
The arrangement, and what a residual is
Five bodies stand between 44° and 61° up, spread between east-north-east and east-south-east, at bearings of 60° to 120°. The sixth is 5° up and due south. With the same stated errors as before, half a minute of arc for the sextant and a tenth of the refraction correction uncertain, the five high sights have standard deviations of 0.93 to 0.95 kilometres and the low one 2.05.
What makes the arrangement matter is not that the low body is behind the others. A line of position is a line, and a body due north gives the same line as one due south. What matters is the line’s direction. The five high bodies give five lines of position running between north-north-west and north-north-east, all nearly parallel, so between them they pin the ship’s east–west position well and its north–south position badly. The low body’s line runs east–west. It is the only sight that says much about how far north the ship is, and the fix depends on it.
The residual reports the error the fix was immune to set out what a residual actually is, and the whole of this account rests on it. A residual is not a sight’s error. It is the part of the error left once the fix has moved to accommodate all the sights. The share of a line’s own error that the fix absorbs is its leverage, the diagonal of the hat matrix that maps observations to fitted values. A line with high leverage has a small residual for a reason that has nothing to do with its quality: the fix has already moved to meet it.
The numbers make the problem concrete. Fitted with equal weights, the five high sights have leverages between 0.23 and 0.32, and their residuals are 86 to 105 per cent the size of their errors. Some are more than a hundred, because part of the low sight’s error has been pushed into them. The low sight’s leverage is 0.61, and its residual is 46 per cent the size of its error.
So the sight with the largest error has the smallest residual relative to that error. A rule that reads residuals as evidence of quality reads it as the best sight in the set.
One step lands in the right place
What happens next is not what that argument predicts, at least not at first.
From equal weights, the fix scatters 1.422 kilometres root-mean-square from the ship. With the altitude’s weights it scatters 1.110, which is the best any linear fix can do, and the 22 per cent between those two numbers is what the weights are worth.
One step of reweighting gives the low sight a median of 4.5 per cent of the weight. Its altitude earns it 4.0. That is almost exactly right. Its residual under equal weights is small relative to its error, but its error is so large that the residual still comes out about as large as the high sights’ residuals, and one inverse-square step lands its share, in the median, close to what its altitude would have given it.
The fix, though, barely improves: 1.398 kilometres, 2 per cent of the 22 on offer. The low sight’s share is right; the others’ are not. With four degrees of freedom shared among six residuals, each line’s residual is a single noisy draw, and the inverse square of one noisy draw is a very noisy weight. The five high sights deserve equal weights and do not get them. The step fixed the one line it was aimed at and scrambled the other five.
The leverage-corrected version does a little better on the first step, 1.336 kilometres. It divides each squared residual by one minus the leverage, the textbook repair, so that a high-leverage line is not credited with a smallness it did not earn. That buys a further 4 per cent. It is still only a little over a quarter of what the altitude weights buy.
Every further step walks it away
Then the iteration continues, and the low sight’s share climbs: 3.7 per cent at three steps, 11 at five, 25 at ten and 25 at thirty. The fix’s scatter climbs with it, past the equal-weight figure by the second step, to 1.517 kilometres at convergence. The leverage-corrected version climbs further, to 1.579.
The mechanism is a feedback loop and it has a direction. Each step gives more weight to the lines with small residuals. Giving a line more weight pulls the fix towards it, which makes its residual smaller, which gives it more weight next time. That is the point of the method: a line that agrees with the consensus is pulled into it and a line that disagrees is pushed out. For an outlier among equals it works.
Here the consensus is not made of equals. Among five nearly parallel lines and one crossing them, the crossing line decides where the fix sits north–south, and the loop cannot tell a line that agrees with the consensus from a line the consensus has been moved to agree with. Once the low sight’s weight rises a little, the fix moves to meet it along the one axis it controls, its residual shrinks, and the loop takes over. The iteration does not find the sight’s quality. It finds its leverage.
In every arrangement, worse than nothing
One arrangement could be special, so the same experiment is run on four: the low body alone in its direction; the low body behind the five, due west; six bodies round the compass, one of them low; and the low body among the five, on their side.
The verdict is the same in all four. Altitude weights are best, as the Gauss–Markov theorem says they must be. Residual weights iterated to convergence are worse than no weights at all: 1.52 against 1.42 kilometres with the low body alone, 1.08 against 0.94 behind, 1.17 against 0.99 round the compass, 1.08 against 0.95 among the five. The leverage correction does not rescue any of them.
What differs between the arrangements is where the low sight ends up. Alone in its direction, it ends with a quarter of the weight. In the other three it ends with far less than it earns, well under one per cent in the median. There its line is not special, its residual shows most of its error, and the loop pushes it out along with a good deal of the useful information it did carry.
So residual weights are wrong in both directions, depending on the geometry: over-trusting a sight that carries the fix, under-trusting one that does not. Neither error tracks the thing the weights were supposed to measure, which was the air.
The share follows the geometry, not the sight
The cleanest way to separate leverage from quality is to hold the geometry fixed and change the quality. The lone southern body can be raised from 3° to 45°, so that its sight improves from nearly three times as noisy as the others to exactly as good.
Its earned share rises steadily with its altitude, from 2.2 per cent at 3° to 16.5 per cent at 45°, where it is an ordinary sight and deserves an ordinary sixth.
The share residual weights give it does not follow. Above eight degrees it sits near a third at every altitude, twice an equal share, whether the sight is noisy or excellent. The residual weights are not measuring the sight. They are measuring the direction it points in. Only at 3°, where the error is so large that even 46 per cent of it outshouts everything else, does the share fall, to 1.6 per cent. That is still wrong, now in the other direction, below the 2.2 the sight earns.
That is the result stated plainly. Weights taken from residuals estimate a line’s leverage at least as much as its variance, and where the two disagree the leverage wins. A navigator who reweights by residuals is reweighting by the geometry of the sky, not the state of the air.
The weights split rather than scatter
A median can hide a great deal, and here it hides the most important part.
Across 4,000 independent sets of sights, the weight given to the low sight does not scatter about some wrong value. It splits into two groups. In 39 per cent of trials the iteration all but discards it, with under one per cent of the weight. In 55 per cent it gives it more than an equal sixth, and in 24 per cent more than forty per cent. Almost no trial gives it anything near the 4 per cent it earns.
That is what a feedback loop does, and the shape of the distribution says so directly. Early in the iteration, one draw of the errors tips the sight one way or the other. Once tipped, it is carried to one of two stable states: pushed out, or placed at the centre of the fix. The navigator does not see a distribution. They see one trial, and a method that throws out the most informative sight four times in ten and hands it nearly half the fix another quarter of the time is not a method that can be corrected afterwards.
What this says about weights a navigator can trust
It sharpens the conclusion of the earlier account. The weights are a guess the solve believes makes the general point that a weighted fit takes its weights at their word. The altitude weights were a guess about the air. They were a guess with a known shape, though: wrong in a direction a navigator could reason about, and safe to err on the side of distrusting the low sight.
Residual weights have neither property. They are not a better-informed guess about the air, because the residuals cannot see what the fix has already absorbed, and they are not wrong in a direction that can be reasoned about, because the direction depends on which way the first draw tipped. Iterated, they are worse than doing nothing.
Two things survive, and both are useful.
One step, leverage-corrected, is a modest improvement. It buys a little over a quarter of what the altitude weights buy, without knowing the weather, and it does not have time to run away. A navigator who wants a check on their refraction guess could compute it and compare. A single step is not the method robust estimation prescribes, however, and its value here comes from stopping before the loop starts.
The residuals remain good at the job they were designed for, which is finding an outlier among equals. A single gross blunder, such as a misread arc, a wrong body or a clerical slip in the reduction, is the kind of residual with more than one explanation only rarely; it stands out whatever its leverage, because its error is so large that even a fraction of it is visible. Two different questions have been run together here. Robust reweighting answers the question of which sight is simply wrong. It cannot answer which sight is somewhat noisier than the others, and that is the question altitude weights answer.
A cocked hat holds the ship one time in four and the ellipse that followed it both took the sights’ errors as known. This is the measurement of what happens when they are not, and it comes out on the side of knowing them: of stating an error for each sight from its altitude, and not asking the sights to report their own.
The controls on these numbers
Altitude weights must be best in every arrangement. They are the inverse variances; by the Gauss–Markov theorem no linear unbiased fix can beat them, and with normally distributed errors no unbiased fix of any kind can. A scheme that beat them by more than the trials’ own scatter would mean the trials or the fit were wrong. None does.
The claim that iteration makes things worse is required, not observed. In all four arrangements the converged residual weights must scatter further than equal weights, and they do.
The direction of the error must follow the geometry. The low sight must be over-trusted by more than a factor of three where it stands alone in its direction, and under-trusted where the bodies surround the ship. Both hold: a quarter against 4 per cent alone, and a small fraction of a per cent against 4 when surrounded.
The split must be a split. Fewer than one trial in five may land within a factor of two of the share the sight earns, or the claim that the weights do not scatter about any value would be an over-reading of a median. Fewer than that do.
And the linear model must be the fit. Every trial here uses the lines of position as straight lines fixed by their bearings. The earlier account showed the full sight reduction reproduces that linear algebra to a few per cent. The question is about weights, and the chart plays no part.
Where the result is narrower than it sounds
Six sights, two unknowns. With four degrees of freedom left over, each residual is a poor estimate of its own line’s variance. With twenty sights the per-line estimates would be better and a single step would buy more. The loop would still find leverage, because leverage does not shrink with more sights unless they fill in the directions the lone body covers.
One robust rule. The weights here are the inverse square of the residual with a floor. Other rules, such as Huber’s, Tukey’s biweight or a trimmed fit, downweight large residuals more gently and are built for outliers. None of them is told a line’s leverage unless the leverage correction is added, and the correction made things worse here. Whether any robust rule converges on altitude weights in this geometry is not measured, but the mechanism above gives no reason to expect it.
Stated errors. The sextant’s half-minute and the tenth of refraction are the same stated values the earlier account used. Different values move the numbers and cannot move the mechanism, which is about geometry.
The mean square, not the worst case. Every comparison is by root-mean-square distance from the ship. A navigator who cares about the worst of a hundred fixes cares about the split more than the mean, and the split is the worst feature of the method.
Still open: whether a sight’s scatter can be learned across fixes
Every fix above uses one set of six sights and forgets it. A navigator takes sights at every dawn and dusk twilight, for weeks, and the low bodies are low again and again. Across fixes, a line’s residual is no longer one noisy draw; it is a series. The leverage changes from fix to fix as the stars move, while the refraction’s share might stay much the same for days in a settled air mass.
That separation is what a single fix cannot make, and it might make the question answerable, in the way a common error solved for as an extra unknown becomes estimable only where the geometry lets it. A sight’s variance could be estimated from its residuals pooled over many fixes, each corrected for that fix’s own leverage, and grouped by altitude, and the result would be an estimate of the refraction share itself. Whether a week of twilights carries enough information to learn the tenth this account had to state, and whether a navigator who did so would beat one who simply guessed it, is a question one set of sights cannot ask.
Named alongside this one
Essays reaching for the same objects. Nobody chose these; they are what the concept index makes visible.
- A coordinate is the output of a solve covariance · least-squares · residual
- A residual cannot count the pieces of a map estimator · least-squares · residual
- A second pin is a measurement of the places covariance · least-squares · weighting
- The answer is a set estimator · least-squares · residual
- The blunder the network cannot see least-squares · residual · standard error
- The seven parameters have their own uncertainty covariance · least-squares · standard error
The objects this essay names
Each one links to every other essay that touches it.
CovarianceEstimatorLeast-squaresNavigationNoiseResidualStandard errorWeighting