Grids, and what a survey does

Naming the control station that moved takes a fifth

A local survey tied to distant control can tell that a monument has moved, and with three control stations it cannot say which. Each station gives the local lines one number they know well — its distance — and a shift of the whole control uses up two of them, so three stations leave one and four leave two. Five evenly spaced name a station moved five centimetres 83 times in a hundred; three name it 20. Better lines cannot stand in for the missing stations, and every station added also exposes more of the control's own published error.

Assumes Three numbers are all a local survey can check of its control.

Three numbers are all a local survey can check of its control found that a local network tied to three distant control stations can check the shape of the control triangle and nothing else, and that almost all of what it can check in practice is the triangle’s size. A monument moved five centimetres along its tie changes that size, and a test read on the right degrees of freedom finds it four times in five.

Monuments do move. A mark is knocked by a vehicle or rebuilt on a new foundation, and the ground under all of them moves by tens of millimetres a year in places where the epoch is part of the coordinate, so a mark observed decades ago and never since can be wrong without anybody having touched it. Finding a moved mark is not the same as knowing which monument it was. A surveyor who learns that the control triangle is too large has learned that one of three marks is wrong, or two, or all three by a little, and the repair — dropping the bad mark and refitting — needs a name. Holding a wrong mark is not a small error either: a coordinate is the output of a solve showed that an over-constrained network pushes the difference into residuals that are then reported as though they were measurement error. The usual instruction is to tie to more control than the minimum so that there is a check. What a check can do depends on how many stations it is made of, and the answer is not the one a count of spare coordinates suggests.

The test for one station

The survey is the one the three-order propagation built: five local stations a kilometre apart, twenty-five lines each good to three millimetres plus two parts per million, and control stations ten kilometres out that carry the regional network’s published error. What changes is the number of control stations. They are spaced evenly round the local cluster, three of them first and then four, five, six and eight, and the local adjustment ties every local station to every one.

For each control station there is one question worth asking: has it moved along its tie? A move across the tie cannot be asked about, for the same reason what a closed figure cannot see found a traverse blind to some errors: nothing observed depends on it. The ties measure distance, a station ten kilometres out moved sideways changes its distance to the cluster by almost nothing, and the fitted survey’s misfit sees a sideways move of five centimetres with a non-centrality of 0.06, which is no signal at all. That is a fact about long ties and it does not improve with the number of stations.

Along the tie the question is Baarda’s outlier test moved from an observation to a control coordinate. Suppose station j has moved a distance ∇\nabla along its tie. That changes every observation by a known pattern, aja_j times ∇\nabla, and the adjusted residuals rr carry some of it. The test statistic is the residuals projected onto that pattern,

wj=ajTP rajTP(I−H) aj,w_j = \frac{a_j^{\mathsf T} P\, r}{\sqrt{a_j^{\mathsf T} P (I-H)\, a_j}},

which is a standard normal variable when nothing has moved and the control is exact. It is the same statistic the blunder the network cannot see used for each distance in a braced quadrilateral, and it is used here the way data snooping always uses it: compute it for every candidate, take the largest, and if it passes the critical value — 3.29 for a test of size one in a thousand — name that candidate.

Two candidates’ tests have a correlation, their normalised inner product in the same metric, and it decides everything that follows. A correlation of one means a move of the first station changes the residuals in exactly the pattern a move of the second would. When that is so, no statistic built from those residuals can say which of the two it was.

Three stations are one test

Moved five centimetres, one station of three looks like all three, and one of five looks like itself. The local cluster at the centre, tied to three control stations (left) and to five (right) ten kilometres away, drawn on their true bearings with the distance shortened. The northern station, filled, has been moved 50 mm along its tie and nothing else is wrong. The bar at each station is the expected value of its own test for a displacement along its tie, noise-free. With three stations the three tests read 3.25, 3.22, 3.22: the data say only that the triangle is too large, and any of its corners could have made it so. With five, the moved station reads 4.31 and no other more than 2.44.
Fig. 1 The local cluster tied to three control stations and to five, drawn on their true bearings with the ten kilometres shortened. The northern station has been moved 50 millimetres along its tie and nothing else is wrong. The bar at each station is the expected value of its own test, without noise. With three stations the tests read 3.25, 3.22 and 3.22. With five, the moved station reads 4.31, and no other reads more than 2.44.

Move the northern of three stations five centimetres outward along its tie and compute each station’s test with no noise at all. The moved station’s test reads 3.25. The other two read 3.22 each. The data do not say that the northern station moved. They say that the control triangle is five centimetres too large at one corner, which is almost exactly what they would say if either of the other corners had moved instead.

The reason is a count, and it is the same count that governed the fitted survey’s misfit. Each control station gives the local lines one number they know well: its distance from the cluster. They know nothing much about where it sits across its tie. Three stations therefore offer three well-known numbers. But a shift of the whole control — the local survey sliding bodily north, say — changes those three distances too, and the local lines cannot tell a shift of their own position from a shift of the control, so two of the three well-known numbers are used up by the two directions a shift can take. What is left is one number. For three stations evenly spaced it is the triangle’s size, and all three stations’ tests are tests of that one number.

Two stations' tests can be told apart only when a shift cannot turn one into the other. The largest correlation between one control station's along-tie test and any other station's, for control stations spaced evenly round the local cluster. A correlation near one means a displacement of the one station is almost exactly as consistent with the data as a displacement of the other. Each station contributes one number the local lines know well, its distance from the cluster, and a shift of the whole control uses up two of them. Three stations: 0.992. Four: 0.994, because opposite stations pair off — one moved outward is the other moved inward plus a shift along their line. Five: 0.567. Six: 0.498. Eight: 0.332.
Fig. 2 The largest correlation between one control station’s along-tie test and any other station’s, for stations spaced evenly round the cluster, beside the count of well-known numbers left once a shift is removed. Three stations: 0.992. Four: 0.994, because opposite stations pair off — one moved outward is the other moved inward plus a shift along their line. Five: 0.567. Six: 0.498. Eight: 0.332.

The correlations are the count made visible. Three stations’ tests correlate at 0.992: they are one test written three times. It is not quite one, because the lines know a little about where each station sits across its tie, and that little is what separates them at all.

Four stations are no better, and that is the result the count of coordinates does not predict. Four stations offer eight coordinates against the three a rigid fit uses up, five spare, and a count of spare coordinates says the check has grown by two thirds. The count that matters is of well-known numbers: four distances, less two for a shift, leaves two. Evenly spaced, the four stations are two opposite pairs, and within each pair the northern station moved outward is exactly the southern station moved inward with the whole control shifted north — the eastern and western stations cannot see a northward shift at all, because it runs across their ties. The two tests of a pair correlate at 0.994.

Five stations leave three numbers, and five tests in three dimensions can be spread out. The largest correlation falls to 0.567. At six it is 0.498 and at eight 0.332. The change between four and five is not a gradual improvement. It is the first count at which the tests stop being copies of one another.

A fifth station turns noticing into naming

A correlation says what the data could distinguish in principle. How often a real test names the right station depends on the noise, and on how large the move is against it.

A fifth control station is what turns noticing into naming. One control station moved 50 mm along its tie, the local lines noisy, the other control perfect; 4,000 seeded trials for each count of control stations. Each trial tests every station along its tie at a size of 0.001 and names the largest test if it passes. Three stations: the moved one named 20 per cent of the time, another 30 per cent, nothing 49 per cent. Four: 44 per cent, 30 per cent, 25 per cent. Five: 83 per cent, 3 per cent, 15 per cent. Eight: 93 per cent, 0 per cent, 6 per cent.
Fig. 3 One control station moved 50 millimetres along its tie, the local lines noisy, the other control perfect, in 4,000 seeded trials for each count of stations. Each trial tests every station along its tie and names the largest test if it passes. Three stations: the moved one named 20 per cent of the time, another 30 per cent, nothing 49 per cent. Four: 44, 30 and 25. Five: 83, 3 and 15. Six: 89, 1 and 10. Eight: 93, under half a per cent, and 6.

With the other control perfect and the local lines carrying their ordinary noise, a station moved five centimetres along its tie is flagged about half the time with three stations — its test’s expected value, 3.25, sits almost exactly on the critical value, 3.29 — and when something is flagged the choice between three nearly identical tests is made by the noise. The moved station is named 20 times in a hundred and one of the other two 30. A surveyor who drops the named station and refits has dropped a good mark and kept the bad one three times in five that the test says anything.

With four stations more moves are flagged, because four ties see a move better than three, and the naming is no better: 44 per cent right and 30 wrong. With five it is 83 right and 3 wrong. The bad station, once flagged, is almost always the one named. At eight, 93 and under half a per cent.

The moves not flagged at all are a separate matter: 15 per cent at five stations, 6 at eight. They are moves too small for the test at that noise, and more stations shrink that fraction only slowly. Naming and noticing are different quantities, and a fifth station changes the first far more than the second.

No arrangement of four does better

Evenly spaced stations are the worst case for pairing, since every station has an exact opposite, and the natural suspicion is that the four-station failure is a failure of symmetry that an irregular layout would avoid.

No arrangement of four stations separates them, because four numbers less two leaves two. How often a station moved 50 mm along its tie is named correctly, for four control stations on four different arrangements of bearings and for five evenly spaced, with the largest correlation between two stations' tests beside each. Evenly spaced four: named 44 per cent of the time, correlation 0.994. Pulled out of symmetry: 53 per cent (0.985), 47 per cent (0.984), 46 per cent (0.994). Five evenly spaced: 83 per cent (0.567). Four stations leave two well-known numbers once a shift is removed, and four tests in two dimensions cannot all point different ways.
Fig. 4 How often a station moved 50 millimetres along its tie is named correctly, for four control stations on four arrangements of bearings and for five evenly spaced, with the largest correlation between two stations’ tests. Evenly spaced four: 44 per cent, correlation 0.994. Pulled out of symmetry: 53 per cent at 0.985, 47 at 0.984, 46 at 0.994. Five evenly spaced: 83 per cent at 0.567.

It does not. Pulling the four stations off their right angles by ten, twenty and thirty degrees moves the naming rate between 46 and 53 per cent and leaves the largest correlation between 0.984 and 0.994. Four tests living in two dimensions cannot all point in different directions: however the four stations are arranged, some pair of them makes nearly the same pattern once the shift is taken out. Symmetry decides which pair it is. The count decides that there is one.

That makes the instruction to tie to one more mark than the minimum exactly half right. A fourth station does improve noticing: the smallest move found four times in five falls from 66 millimetres to 53. It does not buy a name. For a name the survey needs as many well-known numbers left over as there are candidates to separate, in enough dimensions to keep their tests apart, and that starts at five stations for a survey whose ties give it one good number each.

Better lines notice more and name little more

The other way to strengthen a check is to measure better. A surveyor tied to three marks could re-observe every line, or use a better instrument, and ask whether precision can stand in for the missing stations.

Better lines make three stations flag the move every time, and name it only slowly. A station moved 50 mm along its tie, named correctly, against the precision of every local line, from three millimetres plus two parts per million at the left to eight times better at the right. Solid: five control stations. Dashed: three. Dotted: three, how often anything was flagged at all. With three stations, lines twice as good flag the move 100 per cent of the time and name the right station 50 per cent; four times as good, 68 per cent; eight times, 91 per cent. With five, lines twice as good name it 100 per cent of the time.
Fig. 5 A station moved 50 millimetres along its tie, named correctly, against how many times better every local line is. Solid: five control stations. Dashed: three. Dotted: three, how often anything was flagged at all. With three stations, lines twice as good flag the move every time and name the right station 50 per cent of the time; four times as good, 68 per cent; eight times, 91. With five stations, lines twice as good name it every time.

It cannot, cheaply. With three stations, lines twice as good flag a five-centimetre move in every trial — the test has become sensitive — and name the right station exactly half the time. Four times better, 68 per cent. Eight times better, 91. The correlation of 0.992 is not one, so precision does separate the three tests eventually, but it has to overcome a difference between patterns that is under a tenth of their size, and the lines have to get eight times better to do what two more stations do outright. With five stations, lines only twice as good name the moved station in every trial.

Making every line better by the same factor is a common scaling of the weights, the one change the weights are a guess the solve believes found its standard check cannot see; here it sharpens every test at once and leaves the correlations between them exactly where they were, 0.992 at every precision drawn. This is the shape the network’s answer is decided before it is measured keeps finding. What an adjustment can distinguish is a property of its geometry; the precision of its observations decides how sharply it distinguishes the things the geometry already separates, and cannot separate what the geometry makes identical.

Every station added shows more of the control’s own error

Every trial above held the other control stations perfect. They are not. Each carries the regional network’s published error, and a test that treats them as exact will sometimes read that error as a moved monument.

Every station added lets the local lines see more of the control's own published error. The same tests, run with every control station carrying its published error as the three-order propagation gives it, and the tests still treating the control as exact. Dotted: how often a good station is flagged when nothing has been moved — 2 per cent with three stations, 16 per cent with five, 20 per cent with eight, against at most half a per cent with perfect control. Solid: a station moved 50 mm named correctly, 21 per cent, 63 per cent and 80 per cent. Dashed: the wrong station named, 30 per cent, 17 per cent and 9 per cent.
Fig. 6 The same tests with every control station carrying its published error and the tests still treating the control as exact. Dotted: how often a good station is flagged when nothing has moved — 2 per cent with three stations, 8 with four, 16 with five, 12 with six and 20 with eight, against at most half a per cent with perfect control. Solid: a station moved 50 millimetres named correctly — 21, 38, 63, 75 and 80 per cent. Dashed: the wrong station named — 30, 31, 17, 9 and 9 per cent.

With perfect control and nothing moved, the tests flag something at their stated rate: at most half a per cent with eight stations, which is eight tests of size one in a thousand. With the published error in place and nothing moved, three stations flag something 2 per cent of the time, five stations 16 per cent and eight stations 20 per cent. The good stations have not moved. The regional network’s relative error between them is showing.

That is the finding of three numbers are all a local survey can check extended in the direction it pointed. With three stations the local lines can check one number of the control and the published error barely moves it. Each station added gives them another number to check, the published error has a component in each, and the regional network’s twenty-six millimetres at each station — most of it shared, but not all — becomes visible a number at a time. The count that lets a test name a moved station is the same count that lets it see the control’s ordinary error, and a test that assumes the control exact will blame a good mark for it. A published coordinate is a result, with an error of its own, and the more of them a survey leans on the more of that error it has to explain.

The naming rates move accordingly. With five stations and the published error present, a five-centimetre move is named correctly 63 per cent of the time and the wrong station 17, against 83 and 3 with perfect control. At eight, 80 and 9. More control still names better, and the gap between it and perfect control is the regional network’s own imprecision appearing where the test was not built to expect it.

More stations shrink the smallest move that can be noticed, and slowly. The smallest displacement of a control station along its tie that its own test finds four times in five at a size of 0.001, for each count of evenly spaced control stations: 66.1 mm with 3, 52.7 mm with 4, 48.5 mm with 5, 46.2 mm with 6, 43.8 mm with 8. The local station's own reported error falls from 11.2 mm to 7.2 mm over the same range. A move across a tie is not in the table because no count of distant stations makes it visible.
Fig. 7 The smallest displacement of a control station along its tie that its own test finds four times in five at a size of one in a thousand, for evenly spaced stations: 66.1 millimetres with three, 52.7 with four, 48.5 with five, 46.2 with six and 43.8 with eight. The local station’s own reported error falls from 11.2 to 7.2 millimetres over the same range. A move across a tie is not in the table, because no count of distant stations makes it visible.

The smallest detectable move falls slowly — 66 millimetres at three stations, 49 at five, 44 at eight — because each station adds one more well-known distance and a displacement at one station is seen mostly through its own tie. The local survey’s own reported error falls from 11.2 to 7.2 millimetres, since more ties locate it better. Neither curve has a step in it. The step is in the naming, between four and five, and it is invisible in every quantity a survey report normally prints.

Who worked out the test

The statistic is Willem Baarda’s, published in 1968 by the Netherlands Geodetic Commission as a testing procedure for geodetic networks, together with the idea of internal reliability — the smallest error a test will find with a stated probability — that gives the last figure its numbers. His procedure tests one observation at a time, on the hypothesis that one is wrong, and names the one whose test is largest. The correlation between two observations’ tests, and the fact that a correlation near one makes them inseparable, is part of the same account; it was worked out for observations inside a network, where two nearly parallel lines give nearly identical tests, and it applies without change to a control coordinate treated as one more observation.

What the count adds is the reason a particular configuration of control produces inseparable tests. Long ties measure distance, so each station contributes one number, and a shift consumes two. That is a statement about control ten kilometres from a cluster a kilometre wide, and it would change if the ties were short enough to locate the control across them, as at the two-kilometre spacing where the control carries the error of every network above it found the ties fanning widely.

How the numbers were checked

The default survey must be the earlier one. With three control stations every number of the three-order propagation is reproduced — a local station’s full error of 59.381 millimetres to a thousandth — so the generalisation to more stations changed nothing it was not asked to change.

A test with nothing to find must fire at its stated rate. With perfect control and no move, the tests flagged something in 0.1, 0.3, 0.5, 0.4 and 0.5 per cent of 4,000 trials for three to eight stations, against at most eight in a thousand for eight tests of size one in a thousand.

Three stations must be one test and five must not. The largest correlation among three stations’ tests must exceed 0.98 and among five must stay under 0.7; they are 0.992 and 0.567.

The expected test values must follow from the correlations. The noise-free test at each station equals the moved station’s own expected value times the correlation between the two, which is how the first figure’s bars are computed and how they agree with the trials: the moved station’s expected 3.25 against a critical 3.29 is why three stations flag a five-centimetre move about half the time, and do — 51 per cent.

Where the stated survey stops

One moved station at a time. Data snooping assumes one error, and two moved stations can mimic one moved station or none. With five stations two simultaneous moves leave one well-known number to spare, and whether that is enough to name both is not measured here.

The test treats the control as exact. That is how a local adjustment is usually run, and it is the reason for the false alarms in the sixth figure. A test that carried the control’s published covariance would not blame good marks for the regional network’s error, and would find moved marks less readily for the same reason. The trade is the question below.

Stations evenly spaced at one distance. Real control is where it is. Two control stations on nearly the same bearing give nearly the same tie direction, and their tests will correlate whatever the count, so five stations in two tight clusters behave much like two.

Distances only. A local survey with directions or satellite baselines of its own would know each control station across its tie as well, and each station would then contribute two good numbers rather than one. The count would change from k − 2 to 2k − 3, and three stations would leave three.

Still open: a test that knows the control is not exact

The false alarms come from asking a question the test was not built for. It asks whether any control station has moved, and treats every station as perfect unless it has; the control’s ordinary published error, which grows more visible with every station added, then looks like movement. The remedy is standard in form: carry the control’s published covariance into the test, so that the regional network’s twenty-six millimetres are expected rather than flagged.

What that costs is the thing the last figure measured. A test that expects the control to disagree by its published error must see a move clear that disagreement before it fires, so its smallest detectable move grows, and it grows most where the published error is largest relative to the local lines — which is exactly where more stations had made the control’s error visible. Whether a test that knows the control’s covariance still names a five-centimetre move at five stations, how many stations it takes to recover the naming rate of perfect control, and whether a survey is better served by a test that is honest about its control or one that is sensitive to it, are questions a test that assumes the control exact cannot ask.

Named alongside this one

Essays reaching for the same objects. Nobody chose these; they are what the concept index makes visible.

The objects this essay names

Each one links to every other essay that touches it.

AdjustmentControl pointsCovarianceDegrees of freedomHierarchyLeast-squaresMinimal detectable biasNetworkRedundancyReliabilityVerification