One survey cannot tell an optimistic report from a moved mark
Assumes Honest about its control, the test needs eight centimetres where it needed five.
Honest about its control, the test needs eight centimetres where it needed five repaired the station-naming test by telling it the truth about its control. A local survey that holds its regional control exact reads the regional network’s ordinary published error as movement and flags good surveys sixteen times in a hundred; carrying that published covariance into the adjustment as the control’s weight brings the false alarms back to their stated rate, at the price of thirty millimetres of reach.
Every number there rested on one assumption, stated at the end: that the regional network’s published covariance is right. A regional adjustment reports the covariance of its own solution under its own assumptions, and the weights are a guess the solve believes is the standing warning about taking such a number at its word. A local survey tied to many control stations has, in principle, a way to check it: its residuals at the control are a sample of the published error’s differential part, and their size is evidence about whether the report was honest.
The difficulty is the one the honest test was built around, turned round. A residual larger than expected could be a moved monument or an understated covariance. The question is whether one survey can tell them apart, and if not, how many it takes.
One number for the control’s honesty
The model is the smallest one that can ask the question. The regional network’s published covariance is the one the control carries the error of every network above it propagated down three orders, and its true error is that covariance multiplied by an unknown scale . A scale of one is an honest report; two is a report whose real error variances are twice what it published, which is to say its standard deviations are 41 per cent larger; four is a report optimistic by a factor of two in standard deviation.
The local survey is the one the earlier essays built: a five-station local order tied by distances to regional control stations spaced round it, adjusted with the control carried as observations weighted by the published covariance. Its weighted sum of squared residuals, , has an expectation that is linear in the scale:
where is the part of the survey’s redundancy the lines supply and the part the control supplies — both traces computable from the design before a single observation is made. So is an unbiased estimate of the scale, and its variance, for Gaussian errors, is twice the trace of the square of the residuals’ weighted covariance, divided by . This is the estimator Helmert set out for groups of observations of unknown relative precision, with one group — the control — whose precision is in question.
The first figure shows what it gives. With five control stations and an honest report, one survey’s estimate of the scale has a mean of 0.92 and a median of 0.53 — its distribution is lopsided, a sum of squares with few degrees of freedom — and a standard deviation of 3.80. An estimate of one, with an uncertainty of nearly four, says almost nothing about the report.
The control supplies only two degrees of freedom
The estimate is poor because the survey barely sees its control. Of the 25 degrees of freedom in a survey tied to five stations, the lines between the local stations account for 23.1 and the control for 1.9. The distances within the local order are observed to three millimetres and the regional control is uncertain to several centimetres, so the adjustment believes the lines and lets the control take up whatever misfit is left — and the misfit it can be asked to take up is small. Three numbers are all a local survey can check of its control found the same poverty from the other direction: a survey tied to three stations can check three numbers about them.
More stations help slowly. The control’s share of the redundancy grows with the count, from 0.5 at three stations to 8.8 at twenty-four, and the estimate’s standard deviation falls with it, from 10.1 to 1.8. At twenty-four stations — more control than any local survey is tied to — one survey still cannot tell an honest report from one whose variances are twice what it published, because the difference is little more than half a standard deviation.
A moved mark is worth an optimistic report and more
The dashed line in the same figure is the confound, drawn to the same scale. A monument moved along its tie adds to the expected sum of squares exactly as a larger control error would — the move is a residual the adjustment could not absorb — and in units of the scale, a five-centimetre move at five stations adds between 3.7 and 4.2. That is about one standard deviation of the estimate and four times what a report twice too optimistic adds.
The first figure puts the three cases side by side. A survey with an honest report and one moved mark gives an estimate centred at 4.84, well above the 1.91 of a doubled report with nothing moved. A reader who saw a large estimate from one survey would have no way to say which had happened, and the likelier reading, if monuments move at all, is the moved mark: a doubled report shifts the estimate by one, a single five-centimetre move by four.
That is not a failure of the estimator; it is correctly reporting what it sees. It measures how much larger the control’s misfit is than the published covariance predicts, and a moved monument makes the misfit larger. What the estimator cannot do is say why, because both causes enter the same residuals in the same way.
What an optimistic report does to the honest test
The honest test’s whole virtue was that its false alarms fall to its stated rate, and that virtue depends on the report. Carrying the published covariance when the real error variances are twice as large, the test flags 3.9 per cent of good surveys; at four times, 18.7 per cent, which is back where holding the control exact had left it; at eight, 45 per cent. An optimistic report quietly returns the honest test to the failure it was designed to cure.
Run at the true scale, which no survey knows, the test holds its size everywhere: between 0.3 and 0.5 per cent. And run at the scale each survey estimated from its own residuals, it recovers most of the way: 4.1 per cent at twice, 5.5 at four times, 3.5 at eight. An inflated estimate widens every test, which is exactly what an optimistic report needs. That much of the variance component works as intended.
At an honest report it costs something: the survey’s own estimate is sometimes small, and a test run at a small scale is too sharp, so the false alarms rise from 0.5 to 1.8 per cent. A noisy estimate of a quantity that was already right makes the test worse.
Estimated from the survey it tests, the scale explains the move away
The cost falls where the test was meant to earn its keep. With an honest report and one mark moved five centimetres, the test at the published weights names the moved mark 24.9 times in a hundred, which is what the earlier essay found. The same test at the scale the survey estimated from its own residuals names it 7.0 times in a hundred.
The mechanism is a loop with no way out. The move inflates the residuals; the inflated residuals inflate the scale; the inflated scale widens every station’s test by the same factor; and the widened test no longer finds the residual that started it. A survey that estimates its control’s honesty from the same residuals it uses to look for a moved mark has explained the moved mark as dishonesty in advance. The more the mark moved, the more completely it is explained.
The curve for the true scale shows the other half of the trade. When the report really is optimistic, the control’s genuine error is larger and a five-centimetre move is genuinely harder to see: at four times the published variances, even the test that knows the truth names it 4.3 times in a hundred. At the published weights the naming rate stays near 25 per cent however optimistic the report, but by four times the same test is also flagging a fifth of the surveys in which nothing moved, and a flag that fires that often is weak evidence of anything. The blunder the network cannot see is the general form: a test’s power is only as good as its knowledge of what else could have produced the same residuals.
Across surveys the report can be caught
What separates an optimistic report from a moved monument is not in any one survey’s residuals. It is in how they recur. The regional network’s error is a property of the regional network: every local survey tied to it draws its control error from the same true covariance, and if that covariance is twice the published one, every survey’s estimate is pulled upward. A moved monument is an event in one survey and none of the others.
So the scale can be estimated across surveys, in the way a week of twilights learns what a sight is worth estimated a sight’s variance across nights rather than within one. Pool the estimates from many independent local surveys tied to the same regional network and the report’s scale emerges from their average, while each survey’s own moved marks contribute to its own estimate only.
The numbers are large. With five control stations a survey, and one survey in ten containing a monument moved five centimetres, a report whose variances are twice what it published is caught four times in five only once about two hundred surveys are pooled: 66 per cent at a hundred, 90 at two hundred. With twelve stations a survey, fifty surveys are enough.
The median, which a single moved monument cannot drag the way it drags a mean, does worse rather than better — 71 per cent at two hundred surveys — because each survey’s estimate is itself so lopsided that its median sits well below its mean even with nothing moved, and a pooled median is a less efficient summary of a skewed quantity than a pooled mean. The moved marks inflate the mean by a known amount, and a threshold drawn from honest surveys with the same rate of moves already allows for it.
Two hundred surveys tied to one regional network is not unusual over a decade in a busy area. It is an order of magnitude more than any single project has, and — as the network’s answer is decided before it is measured would predict, since the two hundred follows from the design’s traces alone — the conclusion follows directly: a regional report’s honesty is a question for whoever holds the local surveys in bulk, not for any one survey. A mapping agency receiving local surveys tied to its network could estimate its own network’s scale from them; a surveyor tied to that network cannot.
A scale borrowed from other surveys
The pooled estimate is also what rescues the single survey. The trouble with a scale estimated from the survey being tested was that the same residuals served twice; a scale estimated from other surveys of the same regional network serves once, and the survey’s own residuals are left to look for its own moved mark.
Run at a scale pooled from fifty other surveys, the honest test holds its size at every true scale measured: between 0.2 and 0.4 per cent of good surveys flagged, where the published weights let an optimistic report push the rate to 45 per cent. Its naming of a five-centimetre move follows the test that knows the true scale — 8.5 per cent against 11.8 at twice the published variances, 3.7 against 4.3 at four times — rather than collapsing, as the survey’s own estimate did, to a few per cent at an honest report.
It pays one price, and the price is the other surveys’ moved marks. One pooled survey in ten had a monument moved, each of those adds about four to its survey’s estimate, and so the pooled scale sits near 1.4 when the report is honest. The test is run a little too wide, and names the move 18.4 times in a hundred instead of 24.9. Pooling two hundred surveys instead of fifty does not help, because the bias is not noise: it is the rate at which monuments move, carried into everyone’s weights. Naming the control station that moved takes a fifth found that more stations make a moved mark easier to name; here more surveys make the regional report easier to judge and every survey slightly worse at naming its own mark, in proportion to how often marks move elsewhere.
That is a fair trade for any regional network whose report might be optimistic, and a poor one for a network whose report is already honest. Which of those a network is, the pooled estimate itself answers — after two hundred surveys.
What was checked
The estimator must be unbiased. Over 3,000 seeded surveys with an honest report and nothing moved, the mean estimate must be one to within three standard errors; with a report twice too optimistic, two. It is, both times.
Its spread must be the formula’s. The seeded estimates’ standard deviation must fall within ten per cent of the square root of twice the trace of the squared weighted residual covariance, over . At five stations the formula gives 3.80 and the trials agree.
At the true scale the test must hold its size. Whatever the report’s scale, a test that knows it must flag no more than about one good survey in a hundred. It flags between 0.3 and 0.5 per cent.
The finding must be there to fail. A five-centimetre move must add more to the estimate than a doubled report does, which adds one; it adds 3.7 to 4.2. And a test run at its own survey’s estimate must name the move less than half as often as the test at the published weights; it names it 7.0 times in a hundred against 24.9.
Where the model stops
One scale for the whole covariance. A regional report could be optimistic about its shared error and honest about its differential error, or the reverse. A local survey sees only the differential part, so a single scale here is a scale on what the survey can see, and a report wrong only in its shared part would never be caught by any number of local surveys.
Independent surveys. The pooling assumes each survey is tied to different control stations whose errors are independent draws from the regional covariance. Surveys tied to the same stations share one realisation of those stations’ errors, and pooling them adds nothing about the scale however many there are.
One move, along its tie, of a stated size. Moves across a tie are invisible to distances, as what a closed figure cannot see describes, and add nothing to the estimate; larger moves add in proportion to their square.
Gaussian errors. The estimator’s spread is computed for Gaussian errors, and the trials draw them. A regional error with heavier tails would spread the estimate further and push every count of surveys up.
Still open: whether the scale can be estimated by the network that published it
The pooled estimate needs a place where many local surveys meet, and the obvious one is the regional network’s own agency, which receives them. That agency could run the pooled estimate on every survey tied to its network and learn its own scale, region by region.
Whether the local surveys an agency actually receives are numerous enough and independent enough — tied to enough different stations — to reach the two hundred this model needs, whether the scale they would reveal varies across a region in a way that says where the regional adjustment’s assumptions failed, and whether a scale found that way should be fed back into the published covariance or kept as a separate caution, are questions one local survey cannot ask.
Named alongside this one
Essays reaching for the same objects. Nobody chose these; they are what the concept index makes visible.
- A coordinate is the output of a solve adjustment · covariance · least-squares · network · redundancy
- A second pin is a measurement of the places control points · covariance · least-squares · verification · weighting
- A low sight is worth keeping only if it is weighted covariance · least-squares · verification · weighting
- A confidence ellipse is honest only where the sheet keeps angles covariance · least-squares · verification
- An adjustment hides the sphere only in a strip of equal triangles adjustment · least-squares · verification
- The difference of two coordinates covariance · least-squares · redundancy
The objects this essay names
Each one links to every other essay that touches it.
AdjustmentControl pointsCovarianceLeast-squaresMinimal detectable biasNetworkRedundancyReliabilityVerificationWeighting