Paths and directions

What the extra unknown costs where nothing can see it

Four lines of position can estimate a common error in every sight as well as the position. The estimate is honest — three kilometres put in comes back as 2.8 — and where the residual is blind, its scatter is plus or minus thirty-one, and the fix that carries it is scattered fourteen kilometres from the ship against the biased fix's two and a half. Letting the residual choose between the two is worse than either, because a test with no power is a one-in-twenty lottery on the worse answer.

Assumes The residual reports the error the fix was immune to.

The residual reports the error the fix was immune to left a price on the table and did not ask whether to pay it. A common error in every altitude — an index error, an unallowed dip — can be made a third unknown and solved for alongside the position, and four lines of position have room for that: three unknowns and one degree of freedom left over. The estimator exists, it is unbiased, and it removes the whole of a fault the two-unknown fit absorbs into the answer.

It also costs. How much is fixed by the same number that decided whether the residual could see the fault in the first place, and the two point opposite ways: the sky in which the error is invisible is the sky in which removing it is ruinous.

The estimator has been here before, which is worth noticing before anything is measured. A cocked hat holds the ship one time in four used the point equidistant from three lines of position and described it as the estimate that treats a common error as a third unknown. That is exactly this estimator at three sights, where three unknowns and three equations make it a construction rather than a fit. The fourth sight turns a construction into a choice.

Two defensible answers from the same four sights, and they are kilometres apart. Four bodies within sixty degrees of one another at azimuths of 60°, 80°, 100°, 120°, each sight carrying an independent error of about 2 km and the same common error of 3 km. The four lines are nearly parallel, so a step towards every body at once is almost exactly a step to the south. The two-unknown least-squares point takes that step and lands 3.66 km from the ship; the three-unknown fit, which solves for the common error as well, refuses it and lands 2.22 km away — its estimate of the common error being 1.87 km. Neither is wrong. They are the two ends of one exchange.
Fig. 1 Four bodies within sixty degrees of one another, each sight carrying an independent error of about 2 km and the same common error of 3 km. The four lines are nearly parallel, so a step towards every body at once is almost exactly a step to the south. The two-unknown least-squares point takes that step and lands 3.66 km from the ship; the three-unknown fit refuses it and lands 2.22 km away, estimating the common error at 1.87 km. Neither is wrong. They are the two ends of one exchange.

The exchange, and that it has no free parameter

Write κ\kappa for the share of a common error that survives into the residual of the two-unknown fit — the redundancy of the all-ones direction, computed from the four azimuths and nothing else. Solving for the common error instead inflates the position’s variance by a factor. The two are reciprocal, exactly.

The two costs are one quantity, and the line has no fitted parameter in it. Two hundred and sixty arbitrary layouts of four to seven bodies at random azimuths, and the five stated ones, with the share of a common error that reaches the residual against the factor by which solving for that error instead inflates the position's variance. The line is not a fit: it is the identity inflation × share = 1, drawn. The largest departure from it over all 264 layouts is 7.0e-14, which is arithmetic rounding. A navigator who can see a common error can remove it for nothing, and one who cannot see it cannot afford to remove it, because those are the same sentence.
Fig. 2 Two hundred and sixty arbitrary layouts of four to seven bodies at random azimuths, and the five stated ones, with the share of a common error reaching the residual against the factor by which solving for it inflates the position’s variance. The line is not a fit: it is the identity share × inflation = 1, drawn. The largest departure over all two hundred and sixty-four layouts is 7 × 10⁻¹⁴, which is arithmetic rounding.

Two hundred and sixty layouts of four to seven bodies at random azimuths, together with the five stated ones, sit on that line to within seven parts in a hundred million million. It is not a tendency and it is not an approximation for well-conditioned cases; it is one quantity written twice. A common error either lies inside the space the position can move through, in which case the fit absorbs it silently and asking for it back means asking the fit to distinguish two things it cannot distinguish, or it lies outside, in which case the residual reports it and taking it out costs nothing because the position never had it.

That is the whole logic, and it has a consequence worth stating before any measurement: there is no geometry in which a navigator both cannot see a common error and can cheaply remove it. The comfortable middle case does not exist. This is the same structure the blunder the network cannot see finds one level down, where the observation a network checks least is the observation whose error moves the coordinates most, and for the same reason: a redundancy number is simultaneously what a residual can see and what an adjustment has not already used.

What an unknown can imitate

The reason is best said in terms of what a column of the design can do. Each unknown contributes one column, and the fit can only remove from the observations whatever those columns can build. The two position unknowns build every pattern of the form step north by this much and east by that much, read along the four azimuths. The third unknown builds exactly one pattern: the same amount in every sight.

Whether that third pattern is anything new depends on whether it already lay inside the first two, and four nearly parallel lines make it nearly so. At azimuths of 60°, 80°, 100° and 120° the four unit vectors sum to a vector of length 3.70 pointing very nearly north, which means moving the position 1.08 kilometres south reproduces a one-kilometre common error in all four sights to within four parts in a thousand. The third column is a near-copy of one the fit already had, and asking a solve to separate two near-copies is asking it for a difference of large and nearly equal quantities. That is where the factor of two hundred and forty comes from, and it is the same collinearity the network’s answer is decided before it is measured finds in a survey design: what the observations can separate is settled by their arrangement, and no amount of care in taking them changes it.

Three conditions are one too many meets the reverse case in a quite different construction — two free numbers asked to satisfy three exact conditions — and the shared moral is that a count of equations against unknowns says whether a system is solvable and says nothing about whether the solution means anything.

What the estimate is actually worth

An unbiased estimator with no stated spread beside it is the thing the answer is a set refuses to call a measurement, so the first question about the third unknown is not whether it is right on average but how widely it scatters.

The estimate of a common error is honest, and at a bunch it is useless. A common error of exactly 3 km put into every sight, and what the three-unknown fit recovers, over 5,000 trials with independent errors of 2 km. The estimate is right at every layout — 2.97, 2.97, 2.93, 2.83, 2.97 km against the 3 put in. Its scatter is not: the whiskers are two standard deviations, and they run from ±2.0 km at the four quarters to ±31 bunched within sixty degrees. That scatter is σ/√(m·κ) exactly, so at a bunch the fit reports a 3 km sextant error as 2.8 plus or minus 31, which is a number with no information in it.
Fig. 3 A common error of exactly 3 km put into every sight, and what the three-unknown fit recovers, over five thousand trials with independent errors of 2 km. The estimate is right at every layout — 2.97, 2.97, 2.93, 2.83 and 2.97 km against the 3 put in. Its scatter is not: the whiskers are two standard deviations, running from ±2.0 km at the four quarters to ±31 bunched within sixty degrees.

The estimate is honest everywhere. Three kilometres put in comes back as 2.97 at the quarters, 2.93 on one side of the sky over a hundred and twenty degrees and 2.83 bunched within sixty — the small shortfall at the bunch being the ordinary one of a mean over a heavy-tailed sample rather than a bias.

The scatter is another matter, and it has a closed form. The standard deviation of the recovered common error is σ/mκ\sigma/\sqrt{m\kappa}, and over all five layouts the measured scatter agrees with that to within a tenth: 1.01 against 1.00 at the quarters, 3.49 against 3.42 on one side, 15.84 against 15.61 at the bunch. So the sextant error a navigator most wants to know about — three kilometres, a little over a minute and a half of arc — is reported at a narrow bunch as three plus or minus thirty-one. That is not a poor estimate. It is a number with nothing in it, and the arithmetic that produced it announced as much before the sights were taken.

Where the crossing is, and it is further out than instinct puts it

Neither fit dominates. The two-unknown fix carries a bias of (1κ)(1-\kappa) of whatever common error is present; the three-unknown fix carries none and 1/κ1/\kappa times the variance. The crossing is where the bias equals the extra scatter.

Solving for a common error is free in one sky and ruinous in the other. The mean distance of the fix from the ship, ignoring a common error and solving for it, against the size of that error. At the four quarters the two curves lie on one another at 1.78 km, flat across the whole range: the error moves neither fit. Bunched within sixty degrees, ignoring it starts at 2.42 km and climbs to 8.94, while solving for it is flat at 14.1 km — unbiased, and worse than ignoring it across the whole range drawn. At this geometry no sextant error a navigator would credit is large enough to pay for the variance; at the quarters there is nothing to pay for. The exchange between those two is the decision.
Fig. 4 The mean distance of the fix from the ship, ignoring a common error and solving for it, against the size of that error. At the four quarters the two curves lie on one another at 1.78 km and neither moves. Bunched within sixty degrees, ignoring it starts at 2.42 km and climbs to 8.94 at 8 km of common error, while solving for it is flat at 14.1 — unbiased, and worse than ignoring it across the whole range drawn.

At a bunch of sixty degrees the unbiased fit is flat at 14.1 kilometres and the biased one climbs from 2.42 to 8.94 over a common error running to eight kilometres, which is four times the per-sight scatter and more than four minutes of arc on a sextant. The two do not meet inside any range a navigator would take seriously. At the quarters they lie on one another and the question does not arise.

How large a sextant error has to be before it is worth solving for. The common error at which the three-unknown fit overtakes the two-unknown one, against the arc the four bodies span, with independent errors of 2 km. At 90° it is 5.53 km — nearly three times the per-sight scatter, and larger than most real index errors. At a half-turn it is 1.44, and at 240° 1.25. Below 90° the curve leaves the range: at sixty degrees no common error inside twelve kilometres makes solving worthwhile. At 270° the two fits are the same point, so there is nothing to cross and the question dissolves rather than being answered.
Fig. 5 The common error at which the three-unknown fit overtakes the two-unknown one, against the arc the four bodies span. At 90° it is 5.53 km — nearly three times the per-sight scatter. At a half-turn it is 1.44 and at 240° it is 1.25. Below 90° the curve leaves the range: at sixty degrees no common error inside twelve kilometres makes solving worthwhile. At 270° the two fits are the same point, so there is nothing to cross.

Swept over the arc the bodies span, the crossing falls from beyond twelve kilometres at sixty degrees to 5.53 at ninety, 1.44 at a half-turn and 1.25 at two hundred and forty. The shape is 1/κ1/\kappa again, seen through a square root: the crossing is roughly where the bias (1κ)b(1-\kappa)b reaches the extra scatter σ1/κ1\sigma\sqrt{1/\kappa-1} in the position, so it grows without limit as κ\kappa falls.

Instinct puts the crossing low. A navigator who suspects the sextant is out by a mile and a half feels that a mile and a half is a lot, and at the quarters or a half-turn that instinct is right. In the sky where the suspicion is hardest to check, it is wrong by a factor of several, and the error is in the direction that hurts: solving in the belief that one is being careful is where the fourteen-kilometre fix comes from.

It is worth putting the two ends of that curve beside what a navigator actually does about index error, which is to measure it directly — on the horizon, or by bringing the two images of a star together — and apply it before any reduction. The measurements here say what that practice is worth and where. Where the bodies surround the ship, a residual index error costs nothing, so measuring it is a precaution against a fault that would not have mattered. Where they are bunched, it costs a great deal, the sights cannot report it, and solving for it is worse than leaving it: the only remedy that works is the one taken before the sight. The geometry that makes the direct measurement redundant is the geometry in which it is optional, and the geometry that makes it essential is the one in which nothing else will do.

The residual cannot be asked to decide

The obvious escape is to stop choosing in advance. Fit two unknowns, look at the residual, and solve for the common error only if the residual says there is one. That is a pre-test estimator, and it is what most people do without naming it.

Letting the residual choose is worse than either choice it is choosing between. Three rules at a layout of four bodies bunched within sixty degrees, where the residual has almost no power: ignore a common error, solve for it, or solve for it only when a test on the residual rejects at one in twenty. The pre-test rule costs 3.87 km with no common error present against the 2.42 of simply ignoring one — a 60% penalty bought by a test that fires 5.1% of the time, because each time it fires it substitutes a fit scattered 14 km. It never approaches the unbiased rule either. A pre-test built on a test with no power is not a compromise between two estimators; it is the worse one with a lottery attached.
Fig. 6 Three rules at a layout of four bodies bunched within sixty degrees: ignore a common error, solve for it, or solve for it only when a test on the residual rejects at one in twenty. The pre-test rule costs 3.87 km with no common error present against the 2.42 of simply ignoring one — a 60 per cent penalty bought by a test that fires 5.1 per cent of the time, because each time it fires it substitutes a fit scattered 14 km.

At the bunch where the decision is hardest, the pre-test rule is worse than ignoring the error, at every size of error measured, by between fifty and sixty per cent. With nothing wrong at all it costs 3.87 kilometres against 2.42.

The arithmetic of that is simple once seen. The test has no power here — it fires 5.1 per cent of the time with no error present and 5.2 per cent with a four-kilometre one, which is the false-alarm rate twice over. So the rule is: ignore the common error nineteen times in twenty, and one time in twenty substitute a fit scattered fourteen kilometres for no reason. A five per cent chance of a fourteen-kilometre answer adds about a kilometre and a half to the mean, and buys nothing, because the times it fires are uncorrelated with the times an error is present.

The general statement is that a pre-test estimator inherits the weakness of the test it is built on rather than escaping it. Where the test is powerful the pre-test is a good rule and where it is not the pre-test is the worse of the two rules with a lottery attached — and where the test is not powerful is exactly κ\kappa small, which is exactly where somebody would want a rule. The weights are a guess the solve believes makes the neighbouring point about an adjustment’s own inputs: a solve cannot audit an assumption it uses as evidence, and a residual tested against a variance the navigator supplied is a check that borrows its authority from the thing it is checking.

The decomposition of the pre-test’s mean makes the size of the penalty predictable rather than surprising. Its expected error is the ignoring rule’s, weighted by the nineteen times in twenty the test stays silent, plus the solving rule’s for the one time in twenty it fires — and since firing is essentially independent of whether an error is present, that is 0.95×2.42+0.05×14.0=3.000.95 \times 2.42 + 0.05 \times 14.0 = 3.00 kilometres, against a measured 3.87. The gap between those two is itself informative: the test fires on precisely the draws whose residuals are largest, and those are the draws on which the unbiased fit is worst, so the rule substitutes the expensive estimator disproportionately often on the occasions it is most expensive. A rule that picks the worse option at random one time in twenty pays a twentieth of the gap between the options, and the gap here is a factor of six.

What a closed figure cannot see is the surveyed version of the same trap: a misclosure that cannot detect a systematic error is not merely a weak check, it is a check whose passing carries no information, and a procedure that branches on it is branching on noise.

What does work, and is not measured here because it is a different calculation, is deciding before the sights: κ\kappa is known from the azimuths, so a navigator can see at the moment of choosing bodies whether a common error will be checkable, and choose bodies that make it so. The decision that cannot be made after the fact can be made before it.

What the third unknown cannot reach at all

One more thing the extra unknown is asked to do and cannot. Both of the essays behind this one end on a plotting sheet whose longitude is uncorrected for latitude, which displaces the whole plot. It would be pleasant if a third unknown absorbed some of that.

The third unknown absorbs none of a wrong sheet, and pays for it anyway. Four bodies on one side of the ship over 120°, plotted on Mercator and on a sheet with no correction for latitude, with and without the common error solved for. On the uncorrected sheet the two-unknown fit reaches 12.4 km from the ship at 40 km of assumed-position error and the three-unknown fit 12.6 — worse, not better. The estimated common error stays at -0.20 km while the plot is 12 km out: the sheet's error is a displacement of the whole plot rather than a step along the azimuths, and the third unknown can only move along the azimuths. It buys nothing here and still costs its inflation.
Fig. 7 Four bodies on one side of the ship over 120°, on Mercator and on a sheet with no correction for latitude, with and without the common error solved for. On the uncorrected sheet the two-unknown fit reaches 12.4 km from the ship at 40 km of assumed-position error and the three-unknown fit 12.6 — worse, not better. The estimated common error stays at −0.20 km while the plot is 12 km out.

It absorbs none of it. Over assumed-position errors from nothing to eighty kilometres the recovered common error goes from −0.07 to −0.64 kilometres while the fix moves from 1.5 to 24.8 kilometres off the ship, and the three-unknown fit is worse than the two-unknown one at every offset — by its own inflation, which it pays whether or not there is anything to buy with it.

The reason is geometric and final. A common error moves each line along its own azimuth. The sheet’s error moves the whole plot in one direction, lines and ship together, and the third unknown has exactly one shape available to it — the all-ones direction in the space of intercepts — which is not that shape. An unknown can only remove a fault it can imitate.

That is also why no choice of bodies rescues the sheet. The all-ones column is the only extra column on offer, and the sheet’s fault is a translation of the plane, which four azimuths cannot express as a common intercept unless the azimuths are degenerate. A line of position is Newton’s method, but only on a conformal chart located the fault where it belongs — in the construction, before any estimation — and nothing downstream of the construction can reach back to it.

So the account of what four sights buy is now complete and it is narrower than it looked. They buy a residual, which sees disagreement among sights and is blind to anything they share. They buy the option of removing a common error, at a price that is the reciprocal of the residual’s ability to see it. And they buy nothing at all against a wrong sheet, which remains what it was when the intercept method was first drawn here: a fault that no amount of observation detects, and that only a correct construction prevents.

What each number was checked against

At the quarters the two fits must be one fit. With a five-kilometre common error the three-unknown and two-unknown positions must agree within two per cent of each other’s scatter, and do — the control for everything that follows, because a crossing computed there would be an artefact of sampling noise rather than a measurement.

The reciprocal must hold off the stated layouts. Over two hundred and sixty random layouts of four to seven bodies, the product of the surviving share and the inflation departs from 1 by at most 7 × 10⁻¹⁴.

The recovered error’s scatter must be the closed form. For every layout, the measured standard deviation of the estimated common error must be within a tenth of σ/mκ\sigma/\sqrt{m\kappa}, and is: 1.01 against 1.00, 1.03 against 1.01, 3.49 against 3.42, 15.84 against 15.61, 1.94 against 1.89.

A narrow bunch must not cross inside the range. At sixty degrees of span, no common error up to twelve kilometres may make solving worthwhile, and none does.

The pre-test must be the size it claims and must stay between its own two rules. It fires 5.0 per cent of the time with nothing wrong, and at every common error measured its mean error lies between those of the two rules it chooses from — bounded below by the better and above by the worse, as any mixture must be.

And it must be markedly worse than ignoring where the test has no power, which is the point of the figure rather than an incidental: at the bunch it must exceed the ignoring rule by at least thirty per cent, and exceeds it by sixty.

The third unknown must absorb nothing of the sheet. At forty kilometres of assumed-position error on the uncorrected sheet, the three-unknown fit must be no closer to the ship than the two-unknown one and the recovered common error must stay under a twentieth of the sheet’s own displacement; it is 12.6 against 12.4 kilometres, with an estimate of −0.20.

What this priced and what it assumed

The common error is exactly common. Everything here rests on the fault having the all-ones shape across all four sights. A sextant’s index error does; a wrong height of eye does; a refraction error shared by two low bodies and not by two high ones does not, and would need its own column in the design and its own redundancy.

The variance is known. The crossings and the test both use the per-sight scatter as a given. A navigator who estimates it from the four sights themselves has one degree of freedom to do it with at a four-sight fix, which is a variance estimate with a hundred per cent relative error.

The comparison is by mean distance. A navigator who cares about the worst case rather than the average would place the crossings differently, because the unbiased fit’s heavy tail is what a worst-case criterion punishes most. The average was a choice of norm is the general statement of that, and it applies here unchanged: the crossing curve is a property of the mean, not of the estimators.

And the sky is a free choice, which at sea it is not. Choosing bodies to make κ\kappa large assumes there are bodies to choose. Under a broken overcast with one clear patch, the layout is whatever the patch allows, and the measurements here then describe a situation rather than a decision.

Still open: how much of this a shrinkage rule recovers

Two estimators sit at the ends of one line, and the pre-test rule — which jumps between the ends — is worse than either at the geometry where a rule is wanted. The obvious remaining move is not to jump but to slide: estimate the common error and subtract some fraction of it, with the fraction fixed by how well it is known rather than by a test.

That is a shrinkage estimator, and its natural form here writes itself, because the scatter of the recovered common error is known in closed form before any sight is taken. A fraction near one where κ\kappa is large and near nothing where it is small would reproduce both good cases and cannot be worse than the worse of them by construction. Whether such a rule beats simply ignoring the error at a narrow bunch by enough to matter, whether the optimal fraction depends on the size of the error — which is the thing nobody knows — and whether a rule that does depend on it can still be stated in advance from the azimuths alone, are questions two estimators and a coin cannot answer.

What this makes readable

Essays that name this one as a prerequisite.

Named alongside this one

Essays reaching for the same objects. Nobody chose these; they are what the concept index makes visible.

What links here

Every essay whose body links to this one.

The objects this essay names

Each one links to every other essay that touches it.

AzimuthCovarianceDegrees of freedomEstimatorLeast-squaresNavigationRedundancy numberResidualTrade-offVerification