Grids, and what a survey does

The blunder the network cannot see

A least-squares adjustment has no concept of a mistake. The smallest blunder its test will find in the least-checked leg of a braced quadrilateral is 87 millimetres, and by the time it fires a station has moved by nearly ten times the accuracy the same adjustment reports for it.

The weights are a guess the solve believes measured what an adjustment does with an observation that is imprecise. This rung asks what it does with one that is simply wrong.

A least-squares adjustment has no concept of a mistake, and the coordinate it produces is the output of a solve rather than a measurement. It distributes every inconsistency across the residuals in the proportions the redundancy numbers fix, and a gross error is therefore partly shown and partly absorbed into the coordinates. The previous rung measured that split. What it did not ask is the question a surveyor actually has: how big does a blunder have to be before anything notices, and how far have the coordinates moved by then?

How large a blunder has to be before the test notices. A blunder of increasing size put into the least-checked observation of a braced quadrilateral, with the standardised residual it produces. The horizontal line is the critical value the test uses, and the vertical one is the minimal detectable bias — δ₀σ/√r, which is 87 millimetres for this observation and is computed from the network's DESIGN, before any observation is made. Below it nothing is flagged; above it everything is. The observation's redundancy number is 0.144, so it is checked by a seventh of an observation and hides six-sevenths of whatever is wrong with it.
Fig. 1 A blunder of increasing size put into the least-checked observation of a braced quadrilateral, with the standardised residual it produces. The horizontal line is the critical value the test uses; the vertical one is the minimal detectable bias, which is 87 millimetres for this observation and is computed from the network’s design before any observation is made. Below it nothing is flagged; above it everything is.

Two numbers, both from the design alone

Baarda’s answer is a test and a threshold, and neither of them needs any data.

The test is the standardised residual — a residual whose weight was a guess, scaled to be comparable across observations w=v/(σr)w = v/(\sigma\sqrt{r}): an observation’s residual, divided by its own standard deviation and by the square root of its redundancy number. Under the null hypothesis it is a standard normal, so comparing it against a critical value gives a test with a stated false-alarm rate.

The minimal detectable bias is MDB=δ0σ/r\text{MDB} = \delta_0\,\sigma/\sqrt{r}, the smallest blunder the test will find with a stated power. Baarda’s δ0=4.13\delta_0 = 4.13 is the value for a false-alarm rate of one in a thousand and a power of eighty per cent, and it is a property of those two probabilities rather than of any survey.

Both are computable from the design matrix and the weights, which means before the field work. A network can be assessed for what it will be able to detect before anybody has stood behind an instrument.

What the least-checked leg can hide

The smallest blunder each observation cannot hide. The minimal detectable bias of every observation in the same network, at a false-alarm rate of one in a thousand and eighty per cent power. It is δ₀σ/√r and nothing else, so the ordering here is exactly the ordering by redundancy number — printed beside each bar — and both are properties of the geometry rather than of anything measured. The least-checked leg can hide 87 millimetres of error and the best-checked one 48, in a network whose observations are good to eight.
Fig. 2 The minimal detectable bias of every observation in the same network. It is δ₀σ/√r and nothing else, so the ordering is exactly the ordering by redundancy number. The least-checked leg can hide 87 millimetres of error and the best-checked one 48, in a network whose observations are good to eight.

The observations here are distances good to 8 millimetres. The minimal detectable biases run from 48 to 87 millimetres.

So an error of eight centimetres in the least-checked leg — a transcription slip, a wrongly recorded instrument height, a prism constant applied twice — passes every test the adjustment makes, four times in five. The observation is ten standard deviations wrong and the network reports nothing.

That is a much weaker guarantee than the phrase “checked by least squares” suggests, and it is a property of the geometry rather than of the instrument’s quality. Buying a better instrument reduces σ\sigma and reduces the MDB in proportion; it does not change which observations are well checked and which are not.

Where the coordinates have got to

What the coordinates have done by the time the test fires. For each observation, how far a station moves under a blunder exactly at that observation's minimal detectable size — the largest error the network can carry undetected — against the formal standard error the same adjustment reports. The report is the lower band. The worst case is AB, which can hide a blunder that moves B by 71 millimetres against a quoted 7.4 — a factor of 9.7. And it is not the least-checked observation, which is DA: where the coordinates go depends on the geometry as well as on how well checked the leg is.
Fig. 3 For each observation, how far a station moves under a blunder exactly at that observation’s minimal detectable size — the largest error the network can carry undetected — against the formal standard error the same adjustment reports. The report is the lower band. The worst case can hide a blunder that moves a station by nearly ten times the quoted accuracy.

This is the number that matters and the one nobody quotes.

A network can be given a formal standard error of seven millimetres and, at the same time, be unable to detect a blunder that would move that station by seventy-one. The ratio is 9.66 for the worst observation here and above 2 for every one of the ten.

The two quantities come out of the same matrices, in the same solve, at the same moment. One of them appears on every adjustment report and the other does not.

The two orderings are different, and that is the finding

The minimal detectable bias is ordered by redundancy, exactly, because it is δ0σ/r\delta_0\sigma/\sqrt{r} and nothing else. So the least-checked observation has the largest MDB, always, by construction.

The external reliability is not ordered by redundancy. The worst case here is AB, at 9.66 times the formal accuracy, and the least-checked observation is DA at 7.85.

The reason is that where the coordinates go depends on the network’s geometry as well as on how well checked the observation is. A blunder in one leg propagates into the free stations along the directions that leg constrains, and two legs with similar redundancy numbers can constrain very different combinations.

So a redundancy number is not the whole answer. It says how much of a blunder is hidden and it does not say how much damage the hidden part does, and the second is what a client is paying for.

The network under the largest error it cannot detect. A braced quadrilateral with a blunder of 84 millimetres in AB — exactly the minimal detectable size, so the test finds it only four times in five. The displacements are drawn 200 times over life size. The largest is 71 millimetres at B, in a network whose report quotes 7.4 millimetres for that station. Every observation still passes every test; the adjustment is consistent; the coordinates are wrong.
Fig. 4 The network under the largest error it cannot detect: a blunder of exactly the minimal detectable size in the worst leg, with the displacements drawn two hundred times over life size. Every observation still passes every test; the adjustment is consistent; the coordinates are wrong.

What the redundancy number actually is

The quantity everything here depends on deserves a sentence of its own, because it is the least intuitive object in an adjustment.

For each observation, rir_i is the fraction of a unit error in that observation that appears in its own residual. It runs from zero to one. An observation with r=1r = 1 is fully checked — anything wrong with it shows up entirely in its residual and does not touch the coordinates. An observation with r=0r = 0 is unchecked — it is needed to determine the unknowns, nothing constrains it, and every error in it goes straight into the answer.

The redundancy numbers sum to the degrees of freedom, exactly, which is the identity the previous rung’s gate checks. So a network with three degrees of freedom and ten observations has an average redundancy of 0.3, and the average is not the point: the smallest one is, because it is where the undetected error lives.

In this network the smallest is 0.144. That observation shows a seventh of whatever is wrong with it and hides six sevenths.

A gross error splits in the redundancy number's proportion. Each of the ten observations was given a 150 mm gross error in turn, and the change in its own residual measured. It is −r∇ every time, to 1.2 micrometres, which is the nonlinearity of the distance equation over that displacement and not a fitting error. The rest, (1 − r)∇, is absorbed by the coordinates and reported as nothing at all. The worst-checked observation here hides 128 mm of 150.
Fig. 5 The identity underneath all of it, from an earlier rung: each blunder shows exactly its observation’s redundancy number’s share of itself in the residuals, and hides the rest in the coordinates. This rung takes that split and asks where the threshold is — how large the hidden part has to be before the shown part trips a test.
How much of each observation the network can see. An observation's redundancy number is the share of its own error that shows up in its residual; the rest goes into the coordinates. They run from 0.144 to 0.467 here and sum to 3.000000, which is the network's three degrees of freedom — not approximately, identically. The diagonals are the best checked because they are the only observations with two independent routes; the four sides of the quadrilateral are the worst, and a gross error in one of them shows barely a seventh of itself.
Fig. 6 The redundancy numbers themselves, for the same network. They sum to the degrees of freedom exactly, which is the identity this ladder checks, and the smallest of them is where every number in this rung comes from: an observation checked by a seventh of an observation hides six sevenths of whatever is wrong with it.

Why an adjustment cannot do better

The limitation is not a defect of least squares, and it is worth being clear about that.

An adjustment has nn observations and uu unknowns, so it has nun - u degrees of freedom and the residuals live in a subspace of that dimension. A blunder is a vector in observation space, and only its component in the residual subspace is visible. The redundancy number is precisely the fraction of a unit error in that observation that lands in the visible part.

There is nothing to improve. Given the observations, the visible part is the visible part, and any estimator that claimed to see more would be inventing information.

What can be improved is the design. Adding an observation raises the redundancy numbers of the observations it checks — which is what another common point buys in a different setting; arranging the network so that no observation is nearly unchecked raises the worst one. That is what reliability analysis is for and it happens before the survey rather than after.

What a report should say

A conventional adjustment report gives the coordinates, their standard errors, an error ellipse per station, and a variance factor. Every one of those describes what happens if all the observations are merely noisy.

Two numbers describe what happens if one of them is wrong, and both are available in the same solve:

The largest minimal detectable bias in the network, which says what size of mistake could pass unnoticed.

The largest external reliability, which says how far a station could move while that happened.

Neither is expensive. Both come out of matrices the adjustment has already formed, and the second requires one extra solve per observation on a small network.

The eight-centimetre error nobody made up

It is worth grounding the 87 millimetres in the kind of mistake that actually produces it, because the number is otherwise abstract.

A prism constant applied twice, or with the wrong sign, is typically 30 to 60 millimetres — comparable to the corrections a tape measurement carries before it becomes a distance. An instrument or target height entered to the wrong decimal place is a decimetre. A distance transcribed with two digits transposed can be anything. A tape correction omitted on a short baseline is centimetres.

None of those is exotic and all of them are the ordinary failures of a field day. Against a minimal detectable bias of 87 millimetres, the prism constant survives comfortably, the height entry is borderline, and the transposition is caught only if the digits happened to be far apart.

So the threshold sits exactly in the range of the mistakes it is supposed to catch, which is the uncomfortable place for a threshold to be. Raising it — by tightening the false-alarm rate, which surveyors do to avoid rejecting good observations — pushes it further into that range.

The relation to a closed figure

What a closed figure cannot see makes the corresponding argument for the classical method: a traverse closure is a single scalar and it is blind to whole classes of error, including any error that happens to lie along the direction of closure.

A least-squares adjustment is much better than a closure — it has many residuals rather than one, and each of them is checked separately — and it is not qualitatively different. The blindness is smaller and it is still there, it is quantified rather than unknown, and the quantification is the whole advance.

That is the honest summary of what the method bought. Not the elimination of undetectable error, but a number for it, computable in advance, which a closure never had.

The test’s own asymmetry

There is a practical trap in data snooping that the measurement above shows and does not dwell on.

The test flags the observation with the largest standardised residual, and that is not always the observation with the blunder. In this network a blunder in the least-checked leg first flags a different leg, because the least-checked leg shows so little of its own error that a neighbour’s share is larger.

A surveyor who rejects the flagged observation therefore removes the wrong one, re-adjusts, and gets a network that passes — with the blunder still in it and now better hidden, because removing an observation lowers everybody’s redundancy.

The remedy is the standard one and it is procedural rather than statistical: flag, investigate the field record rather than the residual, and re-observe rather than reject.

The design question, put properly

Reliability analysis turns network design from an art into an arithmetic, and the arithmetic is worth stating because it is the whole practical payoff.

A design is a set of stations and a list of observations to make. From those alone — no field work, no instrument readings — the design matrix is known, the weights are known from the instrument’s specification, and every redundancy number, every minimal detectable bias and every external reliability figure follows.

So the question is this network good enough has two answers rather than one. The precision answer is the error ellipses, which is what everybody computes. The reliability answer is the largest undetectable displacement, which is what the client actually cares about and which is usually five to ten times larger.

Improving the first means better instruments or more repetitions. Improving the second means more geometry: an extra tie, a diagonal, a redundant baseline. Those are different remedies with different costs, and a design assessed only on precision will reach for the wrong one.

Where the model stops

Everything here is one blunder in one observation, which is Baarda’s own assumption and is the case the theory covers. Two simultaneous blunders can cancel in the residuals and neither be detected, and the sizes at which that happens are not given by any single-outlier theory.

The network here is also small — ten observations, four free stations — which makes the redundancy numbers low and the effects large. A dense modern network has redundancy numbers near one for most observations and the MDBs are correspondingly close to the noise, which is the reason reliability analysis is not a daily concern in a well-designed survey and is a serious one in a sparse or improvised one.

What “reliable” means when a report says it

The word appears on adjustment reports and in specifications, and it is worth pinning down because it has a technical meaning here that is narrower than the everyday one.

Internal reliability is the minimal detectable bias: the smallest error in each observation the test will find. It is a statement about observations.

External reliability is the effect on the coordinates of an error at exactly that size. It is a statement about the answer, and it is what a user of the coordinates needs.

A network can be internally reliable and externally unreliable, or the reverse. The first happens when every observation is well checked but the geometry amplifies whatever gets through; the second when a poorly checked observation happens to constrain nothing anybody cares about.

That distinction has the same shape as the one between an error ellipse and what a difference of two coordinates carries. So a specification that asks for “reliability” without saying which is asking for one of two different things, and a report that supplies neither — which is the usual case — has not answered either.

Who found it, and when

Willem Baarda’s A testing procedure for use in geodetic networks is from 1968 and it is one of the few pieces of statistical machinery in geodesy that was invented for the subject rather than borrowed. The internal and external reliability measures, the δ0\delta_0 table and the whole apparatus of data snooping are his.

What is striking about it in retrospect is how completely it separates precision from reliability. A network can be precise and unreliable — small ellipses, large undetectable errors — and the two are computed from the same matrices and are not related by any inequality. Half a century later, adjustment software reports the first universally and the second almost never.

The number a client should ask for

If a single figure had to be added to an adjustment report, it would not be either of the two above. It would be their combination.

How far could this coordinate be wrong, given that everything passed? That is a different question from the one a published coordinate’s own accuracy statement answers. That is the external reliability of the worst observation, evaluated at that station, and it is a distance in millimetres directly comparable to the standard error printed beside it.

On this network the two numbers for the worst station are 7.4 millimetres and 71.1 millimetres. A client reading the first has been told what happens if the observations are noisy; a client reading the second has been told what happens if one of them is wrong and nobody noticed, which is the failure mode that actually produces a rebuilt wall.

The second is nearly ten times the first, it comes out of the same solve, and it is not on the report.

Why it is not on the report

The number is computed from the design matrix, which every adjustment package already forms, so its absence is not a technical obstacle. Two other things explain it better.

It is a number about failure rather than about quality. A standard error describes how well the work was done; an external reliability figure describes how badly it could have gone wrong without anybody noticing. Reports are written to convey the first, and a figure ten times larger sitting beside it invites a question the report was not designed to answer.

And it belongs to the design rather than to the survey. External reliability is fixed once the observation plan is chosen — it can be computed before a single measurement is taken, and no amount of care in the field improves it. That makes it awkward to present as an achievement, and it is why the honest place for it is the specification a client agrees to rather than the report they receive at the end.

Which is the practical recommendation this rung leads to: ask for the number before the network is observed, not after. At that point it is a design parameter with a remedy — one more leg, one more tie, one more independent observation — and after the fieldwork it is only a disclosure.

Where the ladder goes next

Nine rungs of this anchor take an observation through a chain of corrections to a coordinate, and the last two ask what the solve believes and what it cannot see. What none of them has asked is what happens when the observations arrive over years rather than in one campaign, so that the network is being adjusted against a ground that has moved between them.

What this makes readable

Essays that name this one as a prerequisite.

Named alongside this one

Essays reaching for the same objects. Nobody chose these; they are what the concept index makes visible.

What links here

Every essay whose body links to this one.

The objects this essay names

Each one links to every other essay that touches it.

AdjustmentBaardaBlunderData snoopingDesignLeast-squaresMinimal detectable biasNetworkRedundancy numberReliabilityResidualStandard error