Grids, and what a survey does

One survey cannot tell an optimistic report from a moved mark

A local survey that carries its control's published covariance can estimate, from its own residuals, how far that covariance understates the truth. With five control stations the estimate has a standard deviation of 3.8 times the published scale, and a single monument moved five centimetres adds about as much again — four times what a report twice too optimistic would add. Run at the scale it estimated for itself, the honest test stops flooding good surveys with alarms, and names a moved mark seven times in a hundred instead of twenty-five: the move inflates the estimate and the estimate absorbs the move. The report can be checked only across surveys, and telling one twice too optimistic from an honest one takes about two hundred of them with five stations each, or fifty with twelve.

Assumes Honest about its control, the test needs eight centimetres where it needed five.

Honest about its control, the test needs eight centimetres where it needed five repaired the station-naming test by telling it the truth about its control. A local survey that holds its regional control exact reads the regional network’s ordinary published error as movement and flags good surveys sixteen times in a hundred; carrying that published covariance into the adjustment as the control’s weight brings the false alarms back to their stated rate, at the price of thirty millimetres of reach.

Every number there rested on one assumption, stated at the end: that the regional network’s published covariance is right. A regional adjustment reports the covariance of its own solution under its own assumptions, and the weights are a guess the solve believes is the standing warning about taking such a number at its word. A local survey tied to many control stations has, in principle, a way to check it: its residuals at the control are a sample of the published error’s differential part, and their size is evidence about whether the report was honest.

The difficulty is the one the honest test was built around, turned round. A residual larger than expected could be a moved monument or an understated covariance. The question is whether one survey can tell them apart, and if not, how many it takes.

From one survey an optimistic report and a moved mark look alike. The scale on the regional network's published covariance, estimated from the residuals of one survey of five control stations, over 3,000 seeded surveys each: honest report, nothing moved, mean 0.92 and median 0.53; report twice too optimistic, mean 1.91 and median 1.43; honest report, one mark moved 5 cm, mean 4.84 and median 4.48. One survey's estimate has a standard deviation of 3.80, and a five-centimetre move along a tie adds between 3.7 and 4.2 to it: about one standard deviation, and four times what a report twice too optimistic adds.
Fig. 1 The scale on the regional network’s published covariance, estimated from the residuals of one survey of five control stations, over 3,000 seeded surveys of each kind. An honest report with nothing moved gives a mean of 0.92 and a median of 0.53; a report twice too optimistic, 1.91 and 1.43; an honest report with one mark moved five centimetres, 4.84 and 4.48. One survey’s estimate has a standard deviation of 3.80.

One number for the control’s honesty

The model is the smallest one that can ask the question. The regional network’s published covariance is the one the control carries the error of every network above it propagated down three orders, and its true error is that covariance multiplied by an unknown scale ss. A scale of one is an honest report; two is a report whose real error variances are twice what it published, which is to say its standard deviations are 41 per cent larger; four is a report optimistic by a factor of two in standard deviation.

The local survey is the one the earlier essays built: a five-station local order tied by distances to regional control stations spaced round it, adjusted with the control carried as observations weighted by the published covariance. Its weighted sum of squared residuals, q=vTPvq = \mathbf v^{\mathsf T} P \mathbf v, has an expectation that is linear in the scale:

E q  =  a+b s,\mathrm E\,q \;=\; a + b\,s,

where aa is the part of the survey’s redundancy the lines supply and bb the part the control supplies — both traces computable from the design before a single observation is made. So s^=(q−a)/b\hat s = (q - a)/b is an unbiased estimate of the scale, and its variance, for Gaussian errors, is twice the trace of the square of the residuals’ weighted covariance, divided by b2b^2. This is the estimator Helmert set out for groups of observations of unknown relative precision, with one group — the control — whose precision is in question.

The first figure shows what it gives. With five control stations and an honest report, one survey’s estimate of the scale has a mean of 0.92 and a median of 0.53 — its distribution is lopsided, a sum of squares with few degrees of freedom — and a standard deviation of 3.80. An estimate of one, with an uncertainty of nearly four, says almost nothing about the report.

The control supplies only two degrees of freedom

More stations sharpen the estimate slowly, and one moved mark stays about as large as its spread. Solid: the standard deviation of the scale estimated from one survey, against the number of control stations — 10.13, 5.19, 3.80, 3.72, 2.62, 2.21, 1.92, 1.76 at 3, 4, 5, 6, 8, 12, 16, 24. Dashed: what one five-centimetre move along a tie adds to the estimate, 9.49, 5.84, 4.18, 4.83, 3.39, 2.90, 2.46, 2.14. The control carries only 0.5, 1.2, 1.9, 2.1, 3.4, 4.9, 6.6, 8.8 of each survey's 15, 20, 25, 30, 40, 60, 80, 120 degrees of freedom, and the estimate can only use those.
Fig. 2 The standard deviation of the scale estimated from one survey, against the number of control stations: 10.13 at three, 3.80 at five, 2.62 at eight, 1.76 at twenty-four. Dashed: what one five-centimetre move along a tie adds to the estimate, from 9.49 at three stations to 2.14 at twenty-four. Of each survey’s 15 to 120 degrees of freedom, the control carries 0.5 to 8.8.

The estimate is poor because the survey barely sees its control. Of the 25 degrees of freedom in a survey tied to five stations, the lines between the local stations account for 23.1 and the control for 1.9. The distances within the local order are observed to three millimetres and the regional control is uncertain to several centimetres, so the adjustment believes the lines and lets the control take up whatever misfit is left — and the misfit it can be asked to take up is small. Three numbers are all a local survey can check of its control found the same poverty from the other direction: a survey tied to three stations can check three numbers about them.

More stations help slowly. The control’s share of the redundancy grows with the count, from 0.5 at three stations to 8.8 at twenty-four, and the estimate’s standard deviation falls with it, from 10.1 to 1.8. At twenty-four stations — more control than any local survey is tied to — one survey still cannot tell an honest report from one whose variances are twice what it published, because the difference is little more than half a standard deviation.

A moved mark is worth an optimistic report and more

The dashed line in the same figure is the confound, drawn to the same scale. A monument moved along its tie adds to the expected sum of squares exactly as a larger control error would — the move is a residual the adjustment could not absorb — and in units of the scale, a five-centimetre move at five stations adds between 3.7 and 4.2. That is about one standard deviation of the estimate and four times what a report twice too optimistic adds.

The first figure puts the three cases side by side. A survey with an honest report and one moved mark gives an estimate centred at 4.84, well above the 1.91 of a doubled report with nothing moved. A reader who saw a large estimate from one survey would have no way to say which had happened, and the likelier reading, if monuments move at all, is the moved mark: a doubled report shifts the estimate by one, a single five-centimetre move by four.

That is not a failure of the estimator; it is correctly reporting what it sees. It measures how much larger the control’s misfit is than the published covariance predicts, and a moved monument makes the misfit larger. What the estimator cannot do is say why, because both causes enter the same residuals in the same way.

What an optimistic report does to the honest test

An optimistic report floods the honest test with alarms, and the survey's own estimate stems most of them. Good surveys of five control stations, nothing moved, flagged by the honest test, against how much larger the regional network's error really is than it published; 2,000 seeded surveys a point. Weights as published: 0.0, 0.5, 3.9, 18.7, 45.1 per cent; weights at the true scale: 0.4, 0.5, 0.4, 0.3, 0.3 per cent; weights at this survey's own estimate: 0.5, 1.8, 4.1, 5.5, 3.5 per cent at 0.5, 1, 2, 4, 8 times. At four times, the published weights flag 18.7 per cent of good surveys and the survey's own estimate 5.5 per cent.
Fig. 3 Good surveys of five control stations, nothing moved, flagged by the honest test, against how much larger the regional network’s error really is than it published. With the published weights: 0.0, 0.5, 3.9, 18.7 and 45.1 per cent at 0.5, 1, 2, 4 and 8 times. At the true scale: 0.3 to 0.5 per cent throughout. At the scale each survey estimated for itself: 0.5, 1.8, 4.1, 5.5 and 3.5 per cent.

The honest test’s whole virtue was that its false alarms fall to its stated rate, and that virtue depends on the report. Carrying the published covariance when the real error variances are twice as large, the test flags 3.9 per cent of good surveys; at four times, 18.7 per cent, which is back where holding the control exact had left it; at eight, 45 per cent. An optimistic report quietly returns the honest test to the failure it was designed to cure.

Run at the true scale, which no survey knows, the test holds its size everywhere: between 0.3 and 0.5 per cent. And run at the scale each survey estimated from its own residuals, it recovers most of the way: 4.1 per cent at twice, 5.5 at four times, 3.5 at eight. An inflated estimate widens every test, which is exactly what an optimistic report needs. That much of the variance component works as intended.

At an honest report it costs something: the survey’s own estimate is sometimes small, and a test run at a small scale is too sharp, so the false alarms rise from 0.5 to 1.8 per cent. A noisy estimate of a quantity that was already right makes the test worse.

Estimated from the survey it tests, the scale explains the move away

Estimated from the survey it is testing, the scale explains the moved mark away. Surveys of five control stations with one mark moved five centimetres along its tie, the move named correctly by the honest test, against the regional network's true error as a multiple of what it published; 2,000 seeded surveys a point. Weights as published: 22.1, 24.9, 26.6, 26.9, 25.8 per cent; weights at the true scale: 44.5, 24.9, 11.8, 4.3, 1.6 per cent; weights at this survey's own estimate: 9.3, 7.0, 4.5, 2.3, 0.8 per cent at 0.5, 1, 2, 4, 8 times. With an honest report the published weights name the move 24.9 per cent of the time and the survey's own estimate 7.0 per cent: the move inflates the estimate, and the inflated estimate absorbs the move.
Fig. 4 Surveys of five control stations with one mark moved five centimetres along its tie, the move named correctly by the honest test. With the published weights: 22.1, 24.9, 26.6, 26.9 and 25.8 per cent at 0.5, 1, 2, 4 and 8 times. At the true scale: 44.5, 24.9, 11.8, 4.3 and 1.6. At the scale each survey estimated for itself: 9.3, 7.0, 4.5, 2.3 and 0.8.

The cost falls where the test was meant to earn its keep. With an honest report and one mark moved five centimetres, the test at the published weights names the moved mark 24.9 times in a hundred, which is what the earlier essay found. The same test at the scale the survey estimated from its own residuals names it 7.0 times in a hundred.

The mechanism is a loop with no way out. The move inflates the residuals; the inflated residuals inflate the scale; the inflated scale widens every station’s test by the same factor; and the widened test no longer finds the residual that started it. A survey that estimates its control’s honesty from the same residuals it uses to look for a moved mark has explained the moved mark as dishonesty in advance. The more the mark moved, the more completely it is explained.

The curve for the true scale shows the other half of the trade. When the report really is optimistic, the control’s genuine error is larger and a five-centimetre move is genuinely harder to see: at four times the published variances, even the test that knows the truth names it 4.3 times in a hundred. At the published weights the naming rate stays near 25 per cent however optimistic the report, but by four times the same test is also flagging a fifth of the surveys in which nothing moved, and a flag that fires that often is weak evidence of anything. The blunder the network cannot see is the general form: a test’s power is only as good as its knowledge of what else could have produced the same residuals.

Across surveys the report can be caught

What separates an optimistic report from a moved monument is not in any one survey’s residuals. It is in how they recur. The regional network’s error is a property of the regional network: every local survey tied to it draws its control error from the same true covariance, and if that covariance is twice the published one, every survey’s estimate is pulled upward. A moved monument is an event in one survey and none of the others.

So the scale can be estimated across surveys, in the way a week of twilights learns what a sight is worth estimated a sight’s variance across nights rather than within one. Pool the estimates from many independent local surveys tied to the same regional network and the report’s scale emerges from their average, while each survey’s own moved marks contribute to its own estimate only.

Checking a regional report takes two hundred local surveys of five stations, or fifty of twelve. How often a report whose true error is twice what it published is told from an honest one, at a false-alarm rate of 5 per cent, when the scale is estimated from many independent local surveys and pooled; one survey in ten has a mark moved five centimetres. Five stations, pooled by the mean: 22, 43, 66, 90, 99 per cent at 20, 50, 100, 200, 400 surveys; five stations, pooled by the median: 20, 27, 50, 71, 92 per cent at 20, 50, 100, 200, 400 surveys; twelve stations, by the mean: 59, 91, 100, 100, 100 per cent at 20, 50, 100, 200, 400 surveys. Eighty per cent is first reached at 200 surveys (five stations, pooled by the mean), 400 surveys (five stations, pooled by the median), 50 surveys (twelve stations, by the mean).
Fig. 5 How often a report whose true variances are twice what it published is told from an honest one, at a false-alarm rate of 5 per cent, when the scale is pooled across independent surveys, one in ten of them with a mark moved five centimetres. Five stations by the mean: 22, 43, 66, 90 and 99 per cent at 20, 50, 100, 200 and 400 surveys. Five stations by the median: 20, 27, 50, 71 and 92. Twelve stations by the mean: 59, 91 and then 100.

The numbers are large. With five control stations a survey, and one survey in ten containing a monument moved five centimetres, a report whose variances are twice what it published is caught four times in five only once about two hundred surveys are pooled: 66 per cent at a hundred, 90 at two hundred. With twelve stations a survey, fifty surveys are enough.

The median, which a single moved monument cannot drag the way it drags a mean, does worse rather than better — 71 per cent at two hundred surveys — because each survey’s estimate is itself so lopsided that its median sits well below its mean even with nothing moved, and a pooled median is a less efficient summary of a skewed quantity than a pooled mean. The moved marks inflate the mean by a known amount, and a threshold drawn from honest surveys with the same rate of moves already allows for it.

Two hundred surveys tied to one regional network is not unusual over a decade in a busy area. It is an order of magnitude more than any single project has, and — as the network’s answer is decided before it is measured would predict, since the two hundred follows from the design’s traces alone — the conclusion follows directly: a regional report’s honesty is a question for whoever holds the local surveys in bulk, not for any one survey. A mapping agency receiving local surveys tied to its network could estimate its own network’s scale from them; a surveyor tied to that network cannot.

A scale borrowed from other surveys

The pooled estimate is also what rescues the single survey. The trouble with a scale estimated from the survey being tested was that the same residuals served twice; a scale estimated from other surveys of the same regional network serves once, and the survey’s own residuals are left to look for its own moved mark.

A scale borrowed from other surveys restores the test's size and costs it only what their moved marks cost. Surveys of five control stations, the honest test run four ways, against the regional network's true error as a multiple of what it published: one mark moved five centimetres, named correctly. Weights as published: 24.9, 26.6, 26.9, 25.8 per cent; weights at the true scale: 24.9, 11.8, 4.3, 1.6 per cent; scale pooled from 50 other surveys: 18.4, 8.5, 3.7, 1.5 per cent; from 200 other surveys: 16.9, 7.8, 3.1, 1.3 per cent at 1, 2, 4, 8 times. Good surveys flagged, the same four ways: 0.5, 3.9, 18.7, 45.1; 0.5, 0.4, 0.3, 0.3; 0.4, 0.2, 0.2, 0.3; 0.2, 0.4, 0.5, 0.7 per cent. One pooled survey in ten had a moved mark of its own, which lifts the pooled scale and costs the test between a quarter and a third of its naming even when the report is honest.
Fig. 6 Surveys of five control stations, one mark moved five centimetres, named correctly by the honest test run four ways, against the regional network’s true error as a multiple of what it published. At the published weights: 24.9, 26.6, 26.9 and 25.8 per cent at 1, 2, 4 and 8 times. At the true scale: 24.9, 11.8, 4.3 and 1.6. At a scale pooled from 50 other surveys: 18.4, 8.5, 3.7 and 1.5; from 200: 16.9, 7.8, 3.1 and 1.3. Every pooled test flags between 0.2 and 0.7 per cent of good surveys.

Run at a scale pooled from fifty other surveys, the honest test holds its size at every true scale measured: between 0.2 and 0.4 per cent of good surveys flagged, where the published weights let an optimistic report push the rate to 45 per cent. Its naming of a five-centimetre move follows the test that knows the true scale — 8.5 per cent against 11.8 at twice the published variances, 3.7 against 4.3 at four times — rather than collapsing, as the survey’s own estimate did, to a few per cent at an honest report.

It pays one price, and the price is the other surveys’ moved marks. One pooled survey in ten had a monument moved, each of those adds about four to its survey’s estimate, and so the pooled scale sits near 1.4 when the report is honest. The test is run a little too wide, and names the move 18.4 times in a hundred instead of 24.9. Pooling two hundred surveys instead of fifty does not help, because the bias is not noise: it is the rate at which monuments move, carried into everyone’s weights. Naming the control station that moved takes a fifth found that more stations make a moved mark easier to name; here more surveys make the regional report easier to judge and every survey slightly worse at naming its own mark, in proportion to how often marks move elsewhere.

That is a fair trade for any regional network whose report might be optimistic, and a poor one for a network whose report is already honest. Which of those a network is, the pooled estimate itself answers — after two hundred surveys.

What was checked

The estimator must be unbiased. Over 3,000 seeded surveys with an honest report and nothing moved, the mean estimate must be one to within three standard errors; with a report twice too optimistic, two. It is, both times.

Its spread must be the formula’s. The seeded estimates’ standard deviation must fall within ten per cent of the square root of twice the trace of the squared weighted residual covariance, over bb. At five stations the formula gives 3.80 and the trials agree.

At the true scale the test must hold its size. Whatever the report’s scale, a test that knows it must flag no more than about one good survey in a hundred. It flags between 0.3 and 0.5 per cent.

The finding must be there to fail. A five-centimetre move must add more to the estimate than a doubled report does, which adds one; it adds 3.7 to 4.2. And a test run at its own survey’s estimate must name the move less than half as often as the test at the published weights; it names it 7.0 times in a hundred against 24.9.

Where the model stops

One scale for the whole covariance. A regional report could be optimistic about its shared error and honest about its differential error, or the reverse. A local survey sees only the differential part, so a single scale here is a scale on what the survey can see, and a report wrong only in its shared part would never be caught by any number of local surveys.

Independent surveys. The pooling assumes each survey is tied to different control stations whose errors are independent draws from the regional covariance. Surveys tied to the same stations share one realisation of those stations’ errors, and pooling them adds nothing about the scale however many there are.

One move, along its tie, of a stated size. Moves across a tie are invisible to distances, as what a closed figure cannot see describes, and add nothing to the estimate; larger moves add in proportion to their square.

Gaussian errors. The estimator’s spread is computed for Gaussian errors, and the trials draw them. A regional error with heavier tails would spread the estimate further and push every count of surveys up.

Still open: whether the scale can be estimated by the network that published it

The pooled estimate needs a place where many local surveys meet, and the obvious one is the regional network’s own agency, which receives them. That agency could run the pooled estimate on every survey tied to its network and learn its own scale, region by region.

Whether the local surveys an agency actually receives are numerous enough and independent enough — tied to enough different stations — to reach the two hundred this model needs, whether the scale they would reveal varies across a region in a way that says where the regional adjustment’s assumptions failed, and whether a scale found that way should be fed back into the published covariance or kept as a separate caution, are questions one local survey cannot ask.

Named alongside this one

Essays reaching for the same objects. Nobody chose these; they are what the concept index makes visible.

The objects this essay names

Each one links to every other essay that touches it.

AdjustmentControl pointsCovarianceLeast-squaresMinimal detectable biasNetworkRedundancyReliabilityVerificationWeighting