What the numbers refer to

What another common point buys

Rung three finds that a seven-parameter datum fit leaves a pattern rather than noise. Six per cent of the residual it reports is the transformation's own error and the other ninety-four is distortion no seven parameters can follow — so adding common points improves a term that was already small and cannot touch the one that is quoted.

Assumes Where a fit leaves residuals.

Where a fit leaves residuals establishes the fact this rung is about. Seven parameters carry a rigid motion and a size; a triangulation network is neither, so the best possible transformation between two datums leaves metres on the table, and it leaves them in a pattern rather than as noise.

A pattern has a consequence that essay did not draw, and it is the one every user of a published transformation is affected by. If the leftover were noise, adding common points would average it away and the transformation would get better as the reciprocal square root of the count. It is not noise. So what does another common point buy?

The answer is a number, and it is small.

What a seven-parameter fit's residual is made of. A published transformation accuracy is the root-mean-square residual at the common points, and here it is 1.51 metres. Almost all of it — 1.51 — is the network's own distortion, which a rigid motion and a scale cannot follow and which is present at every point of the country whether it was used in the fit or not. The transformation's own error, measured as the disagreement between the fitted parameters and the true ones over a clean grid, is 0.091 metres: 6 per cent of the quoted figure. The last bar is the control — the same fit with the distortion switched off, at 2.9e-4 metres.
Fig. 1 A published transformation accuracy is the root-mean-square residual at the common points: here, 1.51 metres. Almost all of it — 1.51 — is the network’s own distortion, which a rigid motion and a scale cannot follow and which is present everywhere in the country whether it was used in the fit or not. The transformation’s own error is 0.091 metres, six per cent of the quoted figure.

Two things are being added together

The residual at a common point is a sum of two quantities with completely different behaviour, and no published accuracy separates them.

The transformation’s own error. How far the fitted seven parameters put a point from where the true seven would. This is what a bigger fit reduces, it is measured here by applying both parameter sets to a clean grid with no distortion anywhere on it, and it comes to nine centimetres.

The distortion the seven cannot represent. The part of the network’s own strain that no rigid motion and scaling can follow. It is at every point of the network, used in the fit or not, and no number of common points touches it. It comes to a metre and a half.

The check that the decomposition is real rather than a story is the control at the bottom of the figure: switch the distortion off and the whole residual falls to 2.9 × 10⁻⁴ metres, which is the linearisation of the rotations in the transformation itself and is four orders of magnitude below the distorted case.

More points, and the same answer

Now the ladder, with the two error sources given the same size and different shapes.

More common points buy a better transformation, and only from noise. The transformation's own error against how many common points it was fitted to, for two populations whose displacements have the same root-mean-square size and different shapes. With independent noise it falls by a factor of 3.8 between 9 points and 196, which is what averaging does. With a smooth strain field of the same size it falls by 1.53, because the seven parameters are converging on the best possible approximation of something they cannot represent, and the approximation does not get better once it has been found.
Fig. 2 The transformation’s own error against the number of common points, for two populations whose displacements have the same root-mean-square magnitude. With independent noise at each marker it falls by a factor of 3.8 between 9 points and 196. With a smooth strain field of the same size it falls by 1.53, and most of that is not averaging.

A factor of 3.8 against a factor of 1.53 is suggestive and is not proof, because both curves fall. The normalisation settles it.

The normalisation that tells the two apart. The same two curves multiplied by the square root of the point count. If an error averages away as 1/√n the product is a constant, and the noise curve is: it runs from 1.36 to 1.68 with no trend. The strain curve climbs by a factor of 3.1 over the same range, so its error is not averaging away at all — the apparent improvement in the previous figure is the fit sampling the strain field more finely rather than the strain field going away.
Fig. 3 The same two curves multiplied by the square root of the point count. Anything that averages away as 1/√n has a constant product, and the noise curve does: 1.36 to 1.68 with no trend. The strain curve climbs by a factor of 3.1 over the same range, so its error is not averaging at all — the fall in the previous figure is the fit sampling the strain field more finely rather than the strain field going away.

That is the finding. A structured distortion and a random error of exactly the same size behave differently under the one operation everybody applies to reduce error, and nothing in a residual table distinguishes them. A fit given a hundred common points on a distorted network has not measured the transformation a hundred times; it has measured the same wrong thing more precisely.

The number that gets published

There is a second, smaller finding underneath, and it is about which residual is quoted.

Every marker, left out in turn. Each common point held back, the seven parameters fitted to the others, and the error at the held-back point measured — the standard estimate of out-of-sample error, and the one no published datum transformation carries. The root-mean-square is 1.763 metres against an in-sample residual of 1.636, a ratio of 1.078. The circles are drawn to the error at each marker: the corners are worse than the middle, because a corner has less of the network on the far side of it to hold the fit in place.
Fig. 4 Every common point held back in turn, the seven parameters fitted to the others, and the error at the held-back point measured. The root-mean-square is 1.763 metres against an in-sample residual of 1.636, a ratio of 1.078. The circles are drawn to the error at each marker, and the corners are worse than the middle.

Leave-one-out is the standard estimate of out-of-sample error in every field that has thought about the question, it takes as long as one fit per marker, and it appears in no published datum transformation. The gap here is eight per cent — small, because the fit has forty-nine points and seven parameters and is nowhere near overfitting.

Fitting on some of the points and scoring on the rest. The same population of common points split five ways, each time fitting on most of them and scoring on the remainder. The held-back error runs from 1.60 to 2.74 metres against in-sample residuals of 1.33 to 1.49. The held-back number is the larger in every split, and it is the one a user of the transformation is in — because a user is at a place that was not a common point.
Fig. 5 The same population split five ways, each time fitting on most of the markers and scoring on the rest. The held-back error is larger than the in-sample residual in every split. It is also the one a user of the transformation is in, because a user is at a place that was not a common point.

The eight per cent is not the interesting part. What is interesting is that the in-sample and out-of-sample numbers are close, which says the fit is well determined — and yet the quantity both of them measure is ninety-four per cent something the fit cannot reduce. A well-determined fit to the wrong model is still the wrong model, and the two diagnoses look identical in a residual table.

Why the seven cannot follow it, in one line

The reason is worth putting plainly because it is the whole mechanism and it is short.

A seven-parameter transformation is three translations, three rotations and a scale. Applied to a set of points, it can move them bodily, turn them, and make them bigger or smaller — and that is all. Written as a displacement field over the country, those seven degrees of freedom span exactly the constant and linear parts of the field: a constant displacement is a translation, a uniform expansion is the scale, and an antisymmetric linear part is a rotation.

So a distortion field that is constant is absorbed entirely. One that is linear is absorbed to the extent that it is a rotation plus a scale — the symmetric traceless part, which is a pure shear, is not, and there are two of those. And a field with any second-order structure at all is not touched.

That is why the strain field in these measurements is a quadratic: it is the lowest order at which the question has an answer, and putting a constant or a linear field in would produce a residual of zero and prove nothing. The choice is stated in the machinery and it is the same discipline that makes an assertion that has never rejected anything worthless.

What this says about published accuracies

An agency publishing a seven-parameter transformation typically states an accuracy of a metre or two over a country. That number is real, it is honest, and it is almost entirely a statement about the old network rather than about the transformation.

Two practical consequences follow and they point in opposite directions.

More common points are worth much less than they look. Doubling the count improves the part that is nine centimetres and leaves the part that is a metre and a half. Any effort spent gathering more common points for a seven-parameter fit is buying a fraction of a fraction, and the fraction is measurable in advance from the residual’s own structure.

And a grid shift is worth much more than it looks. When a formula is not enough sets out why national agencies distribute datum shifts as grids of numbers rather than as parameters, and prices the spacing such a table needs. This rung supplies the other half of the argument: the grid is not a refinement of the seven parameters, it is an attack on the ninety-four per cent, and it is the only thing that can be.

Between them the two essays say something a table of parameters does not. The seven parameters and the grid are not two accuracies of the same kind of object. The seven are a transformation and the grid is a map of a network’s distortion, and a user who has the first and not the second has a transformation whose error is nine centimetres and an answer whose error is a metre and a half.

The rule of thumb this replaces

There is a rule everybody working with datum transformations has heard: use as many common points as are available, spread as widely as possible. Half of it is right and the essay above says which half.

Spread widely: yes, and for a reason this ladder has already measured. Two parameter sets, one transformation shows that markers confined to a small region leave whole parameter combinations undetermined — a translation and a rotation trade against each other, and the network cannot see the difference. Spreading the markers is what conditions the fit, and it acts on the nine-centimetre term through the geometry rather than through the count.

As many as are available: no, or rather, past a point it stops mattering. The count enters only through the averaging, the averaging only works on the part that is random, and the part that is random is six per cent of what gets reported. Doubling from twenty-five markers to fifty moves the transformation’s error from about 0.10 metres to about 0.09.

Two parameter sets 100 metres apart, over the region they were fitted to. Each marker is drawn at a size proportional to how far the two transformations put it apart. The second set differs from the first by 100 metres of translation along the direction this network can least see, with the rotations and the scale re-fitted to absorb it — which is what a second agency's adjustment does when it chooses a different constraint. The worst disagreement anywhere in the region is 5.64 metres and the mean is 3.61. Applied at south-eastern Australia the same two sets differ by 193 metres, because the rotation that absorbed the translation here is a rotation of the whole Earth.
Fig. 6 Why spread matters and count does not: the combination of the seven parameters a country-sized network can least determine, and how much ground movement it produces. The weakest direction is a translation the markers cannot see, and adding markers inside the same region leaves it weak — the eigenvalue is a property of where the markers are.

So the rule survives with its reason changed. Gather markers to condition the fit and to check it; do not gather them expecting the published accuracy to improve, because the published accuracy is measuring something else.

Fitting on some of the points and scoring on the rest. The same population of common points split five ways, each time fitting on most of them and scoring on the remainder. The held-back error runs from 1.45 to 1.81 metres against in-sample residuals of 1.32 to 1.38. The held-back number is the larger in every split, and it is the one a user of the transformation is in — because a user is at a place that was not a common point.
Fig. 7 The same five splits on a larger population — 144 common points instead of 64. Every held-back number falls a little and every in-sample number falls a little, and the gap between them narrows, which is what a better-determined fit looks like. What does not move is that both of them are close to one and a half metres.

One sentence a published transformation could carry

Everything above reduces to a request. A published seven-parameter transformation states its parameters and an accuracy. It could state one number more: how much of that accuracy is the transformation’s own error.

It is computable from what the agency already has — the common points, the fitted parameters, and one evaluation over a clean grid — it does not require a new survey or a new standard, and it converts a figure everybody misreads into two figures that mean what they say. Six per cent and ninety-four per cent are very different instructions about what to do next.

Where the model stops

The distortion field is a stated quadratic, not a real network. The strain field is a smooth second-order function of position with an amplitude in parts per million, chosen because constant strain is absorbed by the scale parameter and linear strain by the rotations, so the leading term a fit cannot reach is the second. A real network’s distortion has power at many scales, so the six per cent above is a property of this field’s spectrum as much as of the arithmetic.

The noise control is independent between markers. Real observation errors in a triangulation are correlated along chains, which is precisely what makes them behave more like the strain case than like the noise case — so the two curves here bracket reality rather than describing it, and the real one is nearer the strain.

Height is not in it. Every marker here has zero ellipsoidal height and the distortion is horizontal. A real common point has a height whose datum is a different question again, and the vertical part of a datum shift is not small.

And the transformation error is measured over the same country. Applying the fitted parameters far outside the region the common points cover produces an error much larger than nine centimetres, because a seven-parameter fit extrapolates a rotation. That is the conditioning question rung three’s neighbour measures, and it is a different failure from this one.

The same shape, elsewhere on this site

Three other measurements on this site have the same structure and it is worth naming, because the shape recurs whenever a model is fitted to something it cannot represent.

A residual has more than one explanation finds a wrong datum and a wrong projection producing residuals that are indistinguishable below about a degree of extent — two mechanisms, one number, and no way to tell from the number which is which.

The rule scored out of sample scores a rule of thumb on a population it was not derived from, which is the same operation this rung performs on a transformation, and reaches the same conclusion: the in-sample number is a statement about the fitting rather than about the world.

And what a closed figure cannot see finds a class of error a traverse’s own closure check is blind to, which is the surveying version of a residual that cannot separate what it is made of.

Four measurements, one lesson: a small residual is evidence about the fit and not about the answer. It is the site’s own habit stated as a result rather than as a method.

What 4 parts per million of network strain leaves behind. Each arrow is where the best-fitting seven-parameter transformation leaves a marker, over 36 markers laid out as a grid across OSGB36's ground. The RMS residual is 1.64 metres and the worst is 3.00 metres. Seven parameters span the constant and linear parts of a displacement field; this one is quadratic, so no choice of the seven can reach it. The arrows are drawn 27× life size.
Fig. 8 The residual field itself, which is what all of this is about: the leftover after the best seven parameters, drawn at the markers. It is smooth, it has a pattern, and every part of that pattern is the ninety-four per cent.

Who found it, and when

The statistics are not in dispute and are not new. In-sample error understating out-of-sample error is the oldest result in model fitting; cross-validation dates from the 1930s and became routine in the 1970s; the distinction between a model’s parameter uncertainty and its approximation error is the bias–variance decomposition and is in every textbook.

None of it is standard practice in datum transformation, and the reason is worth stating without blame. A published transformation is a legal and administrative object as much as a statistical one: it has to be exactly reproducible, quotable, and identical for everyone who uses it, and a leave-one-out estimate is a second number that would have to be defined, agreed and maintained. What agencies did instead was to publish the grid shift, which solves the larger problem directly and makes the smaller one moot.

So the practice is right and the reported number means something other than it appears to. That is the same shape as two parameter sets, one transformation, where two agencies publish different parameters for the same pair of datums and both are correct — and it is the same underlying cause, which is that seven numbers are being asked to describe an object that has more than seven degrees of freedom.

Why a published number has constraints a statistic does not

The defence of current practice deserves more than a clause, because it names a genuine constraint that a purely statistical reading of the problem misses.

A published transformation has to be identical for everybody. Two parties computing coordinates from the same inputs must get the same answer, to the last digit, for as long as the transformation is current — because contracts, boundaries, planning consents and cadastral records depend on the answers agreeing. That requirement is stronger than accuracy and it is not negotiable.

A leave-one-out estimate is a different kind of object. It depends on the station set, the exclusion rule and the fitting procedure, all of which would have to be specified, agreed and frozen to be quotable — and any later addition of a station changes it. A quantity that improves as more data arrives is exactly what a legal reference cannot be.

So the agencies took the option that solves the larger problem outright. A grid-shift file carries the residual pattern the parameters cannot reach, is exactly reproducible, and makes the question of how much a further common point would buy irrelevant — because the shift file already holds what the further points would have contributed.

There is a second constraint that pulls the same way. A published transformation must also be stable in time, because a coordinate computed last year and one computed today have to agree. Any quantity that improves with the arrival of new observations forces a choice between staying current and staying consistent, and a reference system chooses consistency — which is why transformations are versioned and superseded rather than continuously updated, and why the version is part of the identifier.

That is a genuine trade rather than a compromise, and it is the reason the two constraints above are not obstacles to be engineered around.

The two constraints together explain a practice that looks conservative and is not. An agency publishing a transformation is not declining to use better statistics; it is meeting a requirement the statistics do not address, and the requirement is the reason the object exists. A transformation that was slightly more accurate and slightly different every year would be worse at its job than one that is fixed and known, because its job is to make two parties agree.

Which reframes what this rung’s number is for. It is not advice to agencies, who have a better answer. It is for anybody fitting their own transformation over their own region — a survey firm, a research project, an engineering scheme — where no published grid exists, the station set is theirs to choose, and the question is one more point worth observing is a real one with a budget attached.

Where the ladder goes next

The ladder has priced what a datum refers to, how it is fitted, where it leaves residuals, what its epoch means, how its parameters trade against each other, what a chain of them does, and what their own uncertainties are. This rung adds what more of them are worth.

The question it leaves open is the one it has just shown to matter most: how large is the ninety-four per cent, in a real network, and what does it look like? That is not answerable from a stated strain field. It is answerable from a published grid-shift file, which is exactly a map of the quantity — and reading one would be the first time this collection took geometry from a dataset rather than from a rule, which is a decision the site has declined twice before and would have to take deliberately.

Named alongside this one

Essays reaching for the same objects. Nobody chose these; they are what the concept index makes visible.

What links here

Every essay whose body links to this one.

The objects this essay names

Each one links to every other essay that touches it.

Common pointConditioningDatumError budgetGeodetic datumHelmert transformationLeast-squaresNetwork distortionNoiseResidualSamplingVerification