An agency can map where its published covariance is optimistic
Assumes One survey cannot tell an optimistic report from a moved mark.
One survey cannot tell an optimistic report from a moved mark gave a local survey a way to check the control it was tied to, beyond the three numbers three numbers are all a local survey can check of its control allowed it. A survey that carries its control stations’ published covariance into its adjustment can estimate, from its own residuals, a scale s: how many times larger the control’s true variances are than published, one for an honest report. With five control stations the estimate is poor. Its standard deviation is 3.8, and a single monument moved five centimetres along its tie adds about four to it — more than a report twice too optimistic would add. A survey cannot tell one cause from the other, and a test run at the scale its own survey estimated names a moved mark seven times in a hundred instead of twenty-five.
What separates the two causes is repetition. The control’s error is a property of the regional network — the control carries the error of every network above it — and recurs in every survey tied to it; a moved mark is an event in one. Pooled across two hundred surveys, that essay found, a report twice too optimistic is caught nine times in ten. And it named the obvious place for the pooling: the agency that published the regional network, which receives the local surveys tied to it and could learn its own scale from them.
It also named the complication. A regional adjustment’s assumptions do not fail evenly. A network is honest where its observations were good and its model fitted, and optimistic where a deformation went unmodelled, a weak link carried too much, or an old adjustment was stitched to a new one — the places where, as the weights are a guess the solve believes put it, the solve believed a guess. The scale is not one number for a region. It is a map, and the question is whether the surveys an agency receives can draw it.
A region with one optimistic patch
The model keeps the earlier essay’s survey exactly and puts many of them on a map. The region is a square. Its published covariance is honest everywhere except in a disc covering 15 per cent of the area, where the control stations’ true errors have four times the variance published for them. Local surveys fall at random over the square, each tied to five control stations, and one in ten of them has a monument moved five centimetres along its tie. Each survey reports the scale its own residuals estimate; every estimate is drawn from four thousand of the earlier essay’s simulated surveys at the true scale where the survey falls, so every number about a single survey here is that essay’s number.
The agency’s map is built from counts rather than averages. Near every point of a grid of 21 by 21, it counts the surveys within a fifth of the region’s width whose estimate exceeds 2, and compares the count with what an honest region would give: a binomial, since in an honest region four surveys in ten estimate the scale above 2. A cell is flagged when its count is so high that the largest such count anywhere on an honest region’s map would exceed it only one time in twenty. The threshold is the map’s, not the cell’s, so a flag anywhere means the same thing as a flag on a map with one cell.
The first figure is one agency’s map at 800 surveys, chosen as the middle of nine by how much of the patch it flags. It flags 18 cells, all inside the disc, in a block a little off the disc’s centre. The dots show why nothing simpler works: estimates above 2 are everywhere, 357 of the 800, and the patch is visible only as a slightly greater density of them.
One survey’s estimate is a coin weighted a little by the truth
The count is chosen because of what a single estimate looks like. The estimator is unbiased, but its spread is so large against its centre that individual estimates below zero are routine — the middle half of an honest region’s estimates runs from −1.5 to 3.8. Read as a yes-or-no answer to the question “is the scale above 2?”, an honest survey says yes 40 times in a hundred, a survey under a report twice too optimistic 50, four times 63 and eight times 80. A survey is a coin, and the truth weights it only a little.
A count of such answers near a point is therefore a count of weighted coin tosses, and its behaviour under an honest report is known exactly. That is what makes the threshold computable, and it is what keeps the moved monuments in their place: a moved mark adds to its survey’s estimate, but it can only turn one coin to heads, and one survey in ten does it everywhere, honest region or not. A median of a dozen nearby estimates would do the same job less well, because a dozen draws from a spread this wide move their median by more than a patch at four times moves it.
How many surveys find the patch
The answer depends on how optimistic the patch is, and the dependence is steep. A patch whose true variances are eight times those published is found by most agencies with two hundred surveys and by all of them with four hundred. A patch at four times is found by 30 per cent of agencies with two hundred surveys, 69 per cent with four hundred and 98 per cent with eight hundred. A patch at twice its published variances is found by a third of agencies with eight hundred surveys and two thirds with sixteen hundred.
The earlier essay’s two hundred surveys were enough to catch a report twice too optimistic everywhere. Confined to 15 per cent of the region, the same optimism needs more surveys than a busy area produces in a decade, because only 15 per cent of them fall where it is. The count that matters is the surveys inside the patch — a property of where the trouble is rather than of what was measured, as the network’s answer is decided before it is measured found for every number of this kind — and a patch with 800 surveys over the region has about 120 inside it — the same order as the two hundred the earlier essay needed for a whole network twice too optimistic.
Finding the patch is not mapping it
An agency that knows something is wrong somewhere has not yet learned where. The flags that find a patch at four times cover little of it at first: a median 7 per cent of its cells at four hundred surveys, 35 at eight hundred and 72 at sixteen hundred. They sit near the patch’s middle, where the count draws only on surveys inside it, and spread outward as the counts grow.
What the flags almost never do is lie outside the patch. Among agencies that flag anything, the median share of flagged cells inside the disc is between 83 and 100 per cent at every count for a patch at four times. For a patch at eight times the map eventually overshoots: at sixteen hundred surveys it covers the whole patch, and 31 per cent of its flags lie just outside the edge, where the count’s radius reaches into the disc. A flag says where the scale is wrong; the absence of a flag, until the counts are large, says almost nothing.
One number notices trouble only while the surveys are few
A map is more work than a number, and an agency might reasonably ask whether a number would at least tell it that something is wrong, so that the map need only be drawn when it is. The regional median can be tested the way the map is: against the regional medians honest regions’ agencies produce, with a threshold they exceed one time in twenty.
While surveys are few, the number does as well as the map or better. At two hundred surveys it notices a patch at four times 33 times in a hundred against the map’s 30, and a patch at twice 16 times against the map’s 8. The map spends its evidence on asking where, through a threshold that has to hold over 441 cells at once, and a single number pools every survey into one question. As surveys accumulate the order reverses. At eight hundred the map finds the four-times patch 98 times in a hundred and the number 64; at sixteen hundred the map finds even the twice patch two times in three, where the number finds it one time in three. The number’s question is diluted by the 85 per cent of the region that is honest, and more surveys of the honest region sharpen its estimate of the wrong thing.
The patch’s size moves everything. Over 5 per cent of the region instead of 15, a patch at four times is found by a third of agencies at eight hundred surveys and two thirds at sixteen hundred. Over 30 per cent it is found by nearly every agency at four hundred. What an agency can see is set by how many surveys fall inside the trouble, and a small patch of optimism in a large network is invisible to any pooling of a decade’s surveys.
One number for the region dilutes the patch
An agency could instead pool every survey into one regional scale, as the earlier essay pooled the surveys tied to one network. Here the regional number is the median of every survey’s estimate — the mean is pulled up by the moved marks, to about 1.4 even for an honest region. With the patch at four times, eight hundred surveys give a regional scale of 1.30. With the patch at eight times, 1.53. A network whose control is eight times more uncertain than published over a seventh of its area reads, pooled, as half again too optimistic everywhere.
The map’s own values are also short of the truth, for a reason that is the price of the count. The median of the estimates near each point inside the patch is 2.52 for a true 4 and 3.95 for a true 8, because the neighbourhood of a point inside the patch reaches outside it, and because the median of a skewed quantity sits below its mean even when nothing is moved — as the earlier essay found for the pooled median. Outside the patch the map reads 1.06 to 1.20 for a true 1. Every value on the map understates a large scale and overstates an honest one, by a known and modest amount.
Fed back as a map, the scale does its job
The point of learning the scale is to run each local survey’s test at the right one. Honest about its control, the test needs eight centimetres where it needed five showed what goes wrong otherwise: a test that holds the control at its published uncertainty reads the control’s real error as movement. Inside a patch at four times, a survey run at the published weights flags an undisturbed survey 19.9 times in a hundred.
Fed back as one regional number, 1.30, the scale barely helps there: false alarms fall from 19.9 to 13.7 per cent. And it costs where the published covariance was right. Outside the patch, where the report was honest, a test run at 1.30 is run a little too wide, and names a five-centimetre move 18 times in a hundred instead of 26 — nearly a third of its power gone to correct an error it did not have.
Fed back as a map, the scale does what it is for. Inside the patch the test runs at the map’s own value, about 2.5, and false alarms fall to 2.9 per cent. Outside, the map’s value is close to 1, and the test names the move 22 times in a hundred against 26 at the published weights. The decision the earlier essay left open — feed the scale back into the published covariance, or keep it as a separate caution — has a measured answer in this model: feed it back, and feed it back as a map, because a single number corrects the wrong place too little and the right place too much.
What the honest test buys inside the patch
The same figure says something less comfortable about the patch itself. At the true scale, inside it, the honest test names a five-centimetre move only five times in a hundred. That is not a defect of the map. It is naming the control station that moved takes a fifth and its successor read in the other direction: a control station whose position is uncertain by twice its published amount cannot have five centimetres of movement attributed to it with any confidence, and a test honest about that uncertainty mostly declines to.
So the map’s value, about 2.5 against a truth of 4, is a compromise the agency did not choose. Run at the map, the test names the move 14 times in a hundred with 2.9 per cent false alarms; run at the truth, 5 times with 0.6. Whether a local surveyor prefers to be told of more moves with more false alarms or fewer with fewer is a question about the survey’s purpose, and the map leaves the surveyor somewhere sensible between them. What the map cannot do is restore the patch to an honest network’s sensitivity. Only better control can.
What each number was held to
The threshold must hold its size. Over two hundred honest regions not used to set it, an agency of 200 surveys must flag something about one time in twenty. It flags 6 times in 200.
The surveys are the earlier essay’s surveys. Every estimate is drawn from seeded pools of that essay’s simulated surveys at the local true scale, with one in ten carrying a moved mark. The honest pool’s median must be within 0.15 of one; it is 1.01.
The finding must be there to fail. A patch at four times must be found by most agencies of 800 surveys and by few of 50. It is found by 98 per cent and 4 per cent.
A fed-back scale is the scale an agency produced. Each test run at a regional number or a map value draws that number from the three hundred agencies’ actual results, so the spread of the agency’s own estimate is in the false-alarm and naming rates, not only its centre.
Where the map stops
One patch, one shape. A disc of stated area at stated optimism in an otherwise honest region. A real network’s scale probably varies smoothly, with several patches of different sizes, because what it publishes is the output of one adjustment at one time, as a published coordinate is a result found, and a map built for one shape of patch — counts within a fixed radius — will blur a small patch and split a large irregular one.
Surveys at random. Real surveys cluster where building happens, and the patch an agency can see is the patch its surveys fall in. A deformation zone in empty country could be optimistic for decades without a survey to say so.
Five control stations a survey. The earlier essay found twelve stations a survey made fifty surveys do the work of two hundred. With more stations each survey’s coin is weighted more heavily, and every count in this essay would fall.
Independent surveys. Two surveys tied to the same control stations share those stations’ errors, and their estimates are not independent. An agency’s surveys in one district often share stations, so its effective count is smaller than its count of surveys.
Still open: whether the map can be drawn from the network’s own structure
The map here is drawn from where the surveys fell, with a radius that knows nothing about the regional network. But the scale’s variation is not arbitrary: it follows the network’s own weak places — a long chain, a junction of two adjustments, an area of unmodelled motion — and the network’s own adjustment knows where those are. A map whose cells are the regional network’s own regions of shared error, rather than discs on a square, would pool exactly the surveys that share an error.
Whether such a map finds a patch with fewer surveys than a geometric one, whether the regional adjustment’s own redundancy numbers predict where the scale will be wrong before any local survey is received, and whether a patch found that way can be traced to the observations that produced it, are questions a square with a disc on it cannot ask.
Named alongside this one
Essays reaching for the same objects. Nobody chose these; they are what the concept index makes visible.
- A low sight is worth keeping only if it is weighted covariance · estimator · least-squares · verification · weighting
- A second pin is a measurement of the places control points · covariance · least-squares · verification · weighting
- A confidence ellipse is honest only where the sheet keeps angles covariance · estimator · least-squares · verification
- A week of twilights learns what a sight is worth covariance · estimator · least-squares · weighting
- Weighted by what it disagrees with, the fix trusts the sight it leans on covariance · estimator · least-squares · weighting
- What the extra unknown costs where nothing can see it covariance · estimator · least-squares · verification
The objects this essay names
Each one links to every other essay that touches it.
AdjustmentControl pointsCovarianceError budgetEstimatorLeast-squaresReliabilityVerificationWeighting