Measuring distortion

The difference of two coordinates

Three essays give a single coordinate a width. Every practical use of one is a difference of two — a distance, a bearing, a movement, an area — and the width of a difference is not the two widths combined, because the errors are not independent. Far from its datum a one-leg baseline is six times more certain than the positions it joins.

A published coordinate comes with an accuracy statement. This ladder has taken that seriously: a written coordinate has a width set by its own digits, an instrument’s covariance becomes an ellipse on the page, and averaging a population of them moves the answer.

Every one of those is about a single point. Almost nothing anybody does with coordinates is.

A distance is a difference of two. So is a bearing, an area, a displacement between epochs, a check between a design and what was built. And the width of a difference is not the widths of its two ends combined, because the two errors are not independent: most of what is wrong with a published coordinate is wrong with its neighbour, in the same direction, by nearly the same amount. They came out of one adjustment, on one datum, tied to the same control.

A chain of stations, its positions' uncertainty and its baselines'. Nine stations held at the left-hand one, with every neighbour, second neighbour and third neighbour observed. The filled ellipses are each station's own uncertainty, which grows without limit as the chain runs away from the point it is held at — 7.5 mm at the near end and 574 mm at the far. The open ellipses above each leg are the uncertainty of the baseline, drawn at the same scale: they hardly grow at all, because almost everything that is wrong with one end is wrong with the other in the same direction.
Fig. 1 Nine stations in a chain, held at the left-hand one, with every neighbour, second neighbour and third neighbour observed. The filled ellipses are each station’s own uncertainty, which grows without limit as the chain runs away from the point it is held at. The open ellipses above each leg are the uncertainty of the baseline between two adjacent stations, drawn at the same scale, and they hardly grow at all.

What the cofactor matrix says

A coordinate is the output of a solve, and the solve produces more than coordinates: it produces the cofactor matrix Q, which carries the covariance of every unknown with every other. The absolute error ellipse of station i is read off the two-by-two block Q_ii, which is what every adjustment report prints.

The uncertainty of the vector from i to j is read off a different block:

Qrel=Qii+QjjQijQji.Q_{\text{rel}} = Q_{ii} + Q_{jj} - Q_{ij} - Q_{ji}.

The cross terms are the whole content of this essay. A calculation that treats the two positions as independent drops them, which amounts to assuming Q_ij = 0 — and in a network solved as one problem it is very far from zero for stations near each other.

Along the chain above, at the end furthest from the held station, the two positions have absolute ellipses of 574 and 574 millimetres and the baseline between them has a relative ellipse of 154. Combining the two absolutes as though they were independent gives 906 millimetres, which is off by a factor of 5.9.

Where the saving comes from, and where it goes

The uncertainty of a baseline, against the uncertainty its ends would imply. One end held at the far station, the other walked back along the chain. The upper curve is what a reader gets by combining the two published uncertainties as though they were independent; the lower is the truth, which uses the covariance between them. They differ by a factor of 5.9 for the shortest baseline and by 1.00 for the longest, because the correlation weakens with separation. The independent combination is not conservative and it is not wrong by a constant: it is wrong by an amount that depends on how far apart the two points are.
Fig. 2 One end held at the far station of the chain, the other walked back along it. The upper curve is what a reader gets by combining the two published uncertainties as though they were independent; the lower is the truth. They differ by a factor of 5.9 for a one-leg baseline and by 1.00 for one spanning the whole chain.

The two curves meeting is the mechanism made visible. Correlation between two stations’ errors comes from the observations that tie them together and from the datum they share, and both weaken with distance: separated far enough, two stations in the same network have errors that are as good as independent, and the naive combination becomes correct.

So the saving is a local property. It is largest exactly where it is most used — between neighbouring marks, over the distances a site survey works at — and it disappears over the distances at which nobody was going to difference two coordinates anyway.

The other half of the picture is in the first figure. Walking a one-leg baseline outward along the chain, the absolute ellipses grow from 7.5 mm to 574 mm — a factor of 76 — while the relative ellipse grows from 85 mm to 154, a factor of 1.8. The absolute uncertainty is dominated by the accumulated swing of the whole chain about its held station, and that swing is common to both ends of a short baseline, so it subtracts out.

Parts per million, and why specifications are written that way

Quoted as a fraction of the baseline’s own length, the relative uncertainty tells a different story again: 140 parts per million at the near end of the chain and 254 at the far end for a one-leg baseline, and falling steadily to 167 for a baseline spanning the whole chain.

That is why survey specifications are written in parts per million rather than in millimetres. A statement that a network is good to 10 mm is a statement about its datum as much as about its observations; a statement that it is good to 5 ppm is a statement about the observations alone, and it transfers between networks. The classification schemes national mapping agencies publish — order and class, with a constant term and a distance-dependent term — are exactly this pair of quantities kept apart.

Four datums, one set of residuals. The same 10 observations solved four times, holding a different station and bearing each time. Every coordinate in the network moves, by up to 11.0 mm. Not one residual moves: the four sets agree to 4.5e-10 mm, which is arithmetic noise. That is the arithmetical content of the datum is not a measurement — the observations decide the shape and the datum decides where the shape is put, and a reader can tell the two apart without seeing the constraint list, because an over-constrained network fails this test.
Fig. 3 The same observations held at each of four stations in turn. Every coordinate moves and every absolute ellipse changes shape, because an absolute ellipse is a statement about the datum as much as about the observations. The residuals do not move at all, which is the property the next section extends from residuals to baseline lengths.

What survives a change of datum, and what does not

What a change of datum moves, and what it cannot. The same chain solved twice, held at one end and then at the other. Each pair of bars is one baseline: the upper bar is the uncertainty ACROSS it, which is its bearing, and the lower is the uncertainty ALONG it, which is its length. The across bars change by up to 100 per cent between the two solutions; the along bars are identical to every digit, because a network of distances knows its own distances and has no idea which way round it is. That is why survey specifications are written in parts per million of a baseline.
Fig. 4 The same chain solved twice, held at one end and then at the other. Each pair of bars is one baseline: the upper is the uncertainty across it, which is its bearing, and the lower is the uncertainty along it, which is its length. The across bars change by up to 41 per cent between the two solutions; the along bars are identical to every digit.

This is where the essay stops agreeing with the tidy version of itself.

The tidy version says relative quantities are datum-free: hold the network anywhere, and the vector between two stations is the same vector with the same uncertainty. That is not quite true, and the way it fails is instructive.

Resolve the relative ellipse along and across the baseline. The along component — the uncertainty of the length — is identical under the two datums to every digit printed: 4.3198 millimetres held at one end, 4.3198 held at the other, over five baselines. The across component — the uncertainty of the bearing — is 50.65 mm one way and 71.61 the other.

The reason is that a network of distances knows its own distances and has no idea which way round it is. Rotating the whole figure changes no observation, so the orientation is supplied by the datum, and every quantity that depends on orientation inherits the datum’s arbitrariness. A length does not depend on orientation. A bearing is nothing but orientation.

So the correct statement is narrower and more useful than the tidy one: an estimable function of the coordinates has a datum-free uncertainty, and a function that is not estimable does not. In a distance network the lengths are estimable and the bearings are not. Add one observed direction to the network and the bearings become estimable too, and their uncertainties stop moving.

A worked case: setting out from two marks

The abstraction is worth grounding in the operation that uses it most. To set a new mark out, a surveyor occupies one control station, sights another, and measures an angle and a distance. Every quantity in that operation is a difference of the two control coordinates: the bearing between them orients the instrument, and the distance between them scales it.

Suppose the two marks are one leg apart at the far end of the chain above. Their published coordinates each carry an absolute ellipse of 574 millimetres. A user who combines those independently concludes that the baseline is uncertain to 906 millimetres — nearly a metre — and that setting out from it is hopeless at any tolerance worth having.

The truth is 154 millimetres, and resolved along the baseline it is 4.3. The new mark inherits that, plus its own observation errors, and lands where it should relative to everything else in the network. What it does not inherit is any improvement in its absolute position: the whole local figure, new mark included, is still 574 millimetres from where the datum says it is, and it moves as a rigid body if the datum is ever re-realised.

That is the working distinction between the two numbers, and it is why setting out runs the chain backwards is a different problem from positioning: setting out needs the relative number, and only the relative number.

What was computed, and how

The network is a chain of nine stations with every neighbour, second neighbour and third neighbour observed, and the observations are generated from stated coordinates with stated noise from a stated seed. Nothing here is a field measurement, and the caption of every figure says so: what is being measured is the adjustment, and for that the observations have to be ones whose errors are known, which no real observation ever is.

The third-neighbour observations are not decoration. With neighbours and second neighbours alone the chain has exactly as many observations as unknowns — zero degrees of freedom, a variance factor of nought over nought, and error ellipses that are a statement about the arithmetic rather than about the network. A network with no redundancy has no uncertainty to estimate, which is what a closed figure cannot see reaching the same conclusion from the other side.

The comparison against a naive combination uses the same solve and the same scale factor for both numbers, so the ratio is a pure statement about the cross-covariance. And the ladder is measured from the second station rather than the first: the first is the one the network is held at, so it has no uncertainty of its own by construction, and every baseline from it would report the far end’s absolute ellipse under another name — a ratio of exactly one, which is a fact about the datum and not about the correlation.

What the network knows about where each station is. The inverse of the normal matrix has a 2 × 2 block for every unknown station, and each block is an ellipse — in metres already, because the design matrix of a distance is a pair of direction cosines and carries no units. Drawn 2600×, the semi-major axes run from 7.36 mm to 11.12 mm against observations of 8 mm, with axis ratios up to 2.08. B is the exception and is not a measurement: its bearing from the held station is a datum constraint, so it has no freedom at all across that line and its ellipse is a segment. They are the same object as an error ellipse pushed through a projection, arriving from the other end — there a known covariance is mapped, here an unknown one is inferred from the geometry of what was observed.
Fig. 5 The absolute error ellipses of a braced quadrilateral, which is the shape most of this ladder’s neighbours are measured on. Every one of these is a statement about the datum as much as about the observations, and holding a different station moves all of them — which is the property the along-baseline quantity does not have.
One instrument, one covariance, drawn where Mercator puts it. The same measurement is made at every point: 5 m east by 5 m north, uncorrelated, which on the ground is a circle. Each ellipse is that covariance pushed through the projection's own Jacobian and drawn 26,000 times life size. On Mercator every one of them is still a circle — the axis ratio never exceeds 1.000000 — and only the size changes, by up to 2.0 times.
Fig. 6 The same covariance, drawn as it appears on a map: an instrument’s five-metre circular accuracy pushed through a projection, which turns it into an ellipse whose axis ratio is the projection’s own. Every ellipse in this essay is a ground quantity; this is what happens to one when it is drawn.

Why the report prints the number that is less useful

An adjustment report prints absolute ellipses, one per station, and almost never prints the cofactor matrix they came from. That is a choice about format with a consequence about arithmetic: from the printed numbers alone the relative uncertainty cannot be recovered, because the cross-covariance is exactly the part that was not printed.

The information loss is not marginal. A cofactor matrix for a fifty-station network is 100 by 100 numbers; the ellipses are 150. Everything about how the stations lean together — which is most of what the adjustment learnt — lives in the part that does not fit in a table with one row per mark.

Some of it can be reconstructed if the report also prints a relative ellipse per baseline, which good reports do for the baselines that matter, and which is why a specification asks for accuracy “relative to adjacent control” as a separate line from accuracy “relative to the datum”. Those are two different columns of the same solve, and a user given only the second cannot derive the first.

Where the model stops

Distances only. The network observes distances, so its datum defect is three — two translations and a rotation — and scale is fixed by the observations. A network of directions has a different defect and a different set of estimable functions, and its lengths are the quantities that move.

The correlation is the model’s, not an instrument’s. The cross-covariance here comes entirely from the geometry of the adjustment. Real observations have correlated errors of their own — a common atmospheric delay, a shared instrument constant, a temperature that drifts across an afternoon — and those add correlation the cofactor matrix cannot see unless it is told about them.

One scale of network. The chain is 4.8 kilometres long with 600-metre legs, and the transition from strong correlation to none happens across it. A continental network has the same structure over a thousand times the distance, and the same statement in parts per million; nothing here establishes the constant.

Nothing here is about accuracy. Every number is a precision — an internal consistency of one adjustment. A network can be exquisitely precise and sit two metres from where it says it is, which is what a published coordinate is and is a different quantity entirely.

The generalisation

The rule is one sentence with a condition attached: the uncertainty of a difference is smaller than the uncertainty of its parts by however much they are correlated, and they are correlated by however much they share.

Two coordinates from one adjustment share almost everything, so their difference is far better than either. Two coordinates from different adjustments — one from a national network, one from a satellite fix, one read off a map — share nothing, and the naive combination is then correct. Which means the dangerous case is not the one this essay is about; it is the mixed case, where a user differences a coordinate from one source with a coordinate from another and applies a rule of thumb learnt from the first.

The practical version of that is a warning about the wrong direction. A surveyor who knows relative accuracy is better than absolute will happily difference two published coordinates and quote a tight number, and that number is right only if the two came out of the same solve. A published coordinate list does not usually say which of its entries were adjusted together, and the covariance is almost never published at all — so the information needed to compute the number correctly is the part that gets discarded before publication.

Who found it, and when

The distinction between absolute and relative accuracy is as old as least squares in geodesy, and the algebra above is standard: relative error ellipses appear in Baarda’s work in the 1960s and in every adjustment textbook since.

The deeper statement — that only estimable functions have datum-free uncertainties, and that which functions are estimable depends on what was observed — is the theory of S-transformations, published by Baarda in 1973. It is the reason a modern adjustment report distinguishes a minimally constrained solution from an over-constrained one, and it is why holding a second station changes every coordinate and no residual.

What is not standard is the habit of publishing the covariance, and that is a matter of format rather than of theory. A coordinate list is a table of positions because a table of positions is what a printed sheet could hold, and the convention outlived the constraint by about forty years.

What a consumer can do with only the absolute figures

Most users will never be given the covariance, so it is worth saying what can honestly be done with the numbers a published list does carry.

Combining two absolute uncertainties in the usual way gives an upper bound. Adding the variances of two coordinates and taking the root is what everybody does, and it is exactly right when the two are uncorrelated. Neighbouring marks in one network are not: they share the datum, they share the observations that tie them to it, and their errors move together.

The overstatement is a factor of 1/1ρ1/\sqrt{1-\rho}, since the relative variance is 2σ2(1ρ)2\sigma^2(1-\rho) against the naive 2σ22\sigma^2. At a correlation of 0.9 that is a factor of about three; at 0.99, ten. Those are ordinary correlations for marks a few hundred metres apart in one adjustment.

So the bound is safe and is grossly pessimistic at short range, which is the range at which most work is done. A contractor told that two adjacent control marks are each good to twenty millimetres, and who therefore budgets thirty for the distance between them, may be working with a pair whose relative uncertainty is three. The specification is met with room to spare and the work is planned as though it were not.

And the direction of the error is the useful part. The naive combination never understates the relative uncertainty, so nothing built on it is unsafe. What it does is make short-range work look harder than it is, which is paid for in unnecessary observation rather than in rebuilt walls — a cheaper failure than the alternative, and one nobody notices they are paying.

There is one case where the pessimism does cause harm, and it is worth naming. A tolerance that cannot be met on paper gets a job redesigned — a denser control network, a tighter observation scheme, sometimes a different construction method — and the redesign is chosen against a figure that was never the real one. The cost there is not extra observation but a decision made on a number three times too large.

The number that would fix it is one correlation coefficient per pair, which the adjustment computed and the report discarded.

Where the ladder goes next

Four rungs give a coordinate a width, an ellipse, a bias and now a correlation. What is still missing is the width of a coordinate that has been transformed: pushed through a datum shift, a projection and a grid, each of which has parameters with uncertainties of their own.

Some of that is already here — the seven parameters have their own uncertainty and it propagates into a positional covariance that grows with distance from the centroid. What that essay does not do is difference two such coordinates, and the correlation between them is enormous, because the transformation parameters are shared exactly.

Named alongside this one

Essays reaching for the same objects. Nobody chose these; they are what the concept index makes visible.

What links here

The 8 essays that link to this one and share the most of its objects, of 10 that link here.

The objects this essay names

Each one links to every other essay that touches it.

CorrelationCovarianceDatumDegrees of freedomError ellipseInvariantLeast-squaresMeasurementNetwork strainPrecisionRedundancyReference frame