What another common point buys
Assumes Where a fit leaves residuals.
Where a fit leaves residuals establishes the fact this rung is about. Seven parameters carry a rigid motion and a size; a triangulation network is neither, so the best possible transformation between two datums leaves metres on the table, and it leaves them in a pattern rather than as noise.
A pattern has a consequence that essay did not draw, and it is the one every user of a published transformation is affected by. If the leftover were noise, adding common points would average it away and the transformation would get better as the reciprocal square root of the count. It is not noise. So what does another common point buy?
The answer is a number, and it is small.
Two things are being added together
The residual at a common point is a sum of two quantities with completely different behaviour, and no published accuracy separates them.
The transformation’s own error. How far the fitted seven parameters put a point from where the true seven would. This is what a bigger fit reduces, it is measured here by applying both parameter sets to a clean grid with no distortion anywhere on it, and it comes to nine centimetres.
The distortion the seven cannot represent. The part of the network’s own strain that no rigid motion and scaling can follow. It is at every point of the network, used in the fit or not, and no number of common points touches it. It comes to a metre and a half.
The check that the decomposition is real rather than a story is the control at the bottom of the figure: switch the distortion off and the whole residual falls to 2.9 × 10⁻⁴ metres, which is the linearisation of the rotations in the transformation itself and is four orders of magnitude below the distorted case.
More points, and the same answer
Now the ladder, with the two error sources given the same size and different shapes.
A factor of 3.8 against a factor of 1.53 is suggestive and is not proof, because both curves fall. The normalisation settles it.
That is the finding. A structured distortion and a random error of exactly the same size behave differently under the one operation everybody applies to reduce error, and nothing in a residual table distinguishes them. A fit given a hundred common points on a distorted network has not measured the transformation a hundred times; it has measured the same wrong thing more precisely.
The number that gets published
There is a second, smaller finding underneath, and it is about which residual is quoted.
Leave-one-out is the standard estimate of out-of-sample error in every field that has thought about the question, it takes as long as one fit per marker, and it appears in no published datum transformation. The gap here is eight per cent — small, because the fit has forty-nine points and seven parameters and is nowhere near overfitting.
The eight per cent is not the interesting part. What is interesting is that the in-sample and out-of-sample numbers are close, which says the fit is well determined — and yet the quantity both of them measure is ninety-four per cent something the fit cannot reduce. A well-determined fit to the wrong model is still the wrong model, and the two diagnoses look identical in a residual table.
Why the seven cannot follow it, in one line
The reason is worth putting plainly because it is the whole mechanism and it is short.
A seven-parameter transformation is three translations, three rotations and a scale. Applied to a set of points, it can move them bodily, turn them, and make them bigger or smaller — and that is all. Written as a displacement field over the country, those seven degrees of freedom span exactly the constant and linear parts of the field: a constant displacement is a translation, a uniform expansion is the scale, and an antisymmetric linear part is a rotation.
So a distortion field that is constant is absorbed entirely. One that is linear is absorbed to the extent that it is a rotation plus a scale — the symmetric traceless part, which is a pure shear, is not, and there are two of those. And a field with any second-order structure at all is not touched.
That is why the strain field in these measurements is a quadratic: it is the lowest order at which the question has an answer, and putting a constant or a linear field in would produce a residual of zero and prove nothing. The choice is stated in the machinery and it is the same discipline that makes an assertion that has never rejected anything worthless.
What this says about published accuracies
An agency publishing a seven-parameter transformation typically states an accuracy of a metre or two over a country. That number is real, it is honest, and it is almost entirely a statement about the old network rather than about the transformation.
Two practical consequences follow and they point in opposite directions.
More common points are worth much less than they look. Doubling the count improves the part that is nine centimetres and leaves the part that is a metre and a half. Any effort spent gathering more common points for a seven-parameter fit is buying a fraction of a fraction, and the fraction is measurable in advance from the residual’s own structure.
And a grid shift is worth much more than it looks. When a formula is not enough sets out why national agencies distribute datum shifts as grids of numbers rather than as parameters, and prices the spacing such a table needs. This rung supplies the other half of the argument: the grid is not a refinement of the seven parameters, it is an attack on the ninety-four per cent, and it is the only thing that can be.
Between them the two essays say something a table of parameters does not. The seven parameters and the grid are not two accuracies of the same kind of object. The seven are a transformation and the grid is a map of a network’s distortion, and a user who has the first and not the second has a transformation whose error is nine centimetres and an answer whose error is a metre and a half.
The rule of thumb this replaces
There is a rule everybody working with datum transformations has heard: use as many common points as are available, spread as widely as possible. Half of it is right and the essay above says which half.
Spread widely: yes, and for a reason this ladder has already measured. Two parameter sets, one transformation shows that markers confined to a small region leave whole parameter combinations undetermined — a translation and a rotation trade against each other, and the network cannot see the difference. Spreading the markers is what conditions the fit, and it acts on the nine-centimetre term through the geometry rather than through the count.
As many as are available: no, or rather, past a point it stops mattering. The count enters only through the averaging, the averaging only works on the part that is random, and the part that is random is six per cent of what gets reported. Doubling from twenty-five markers to fifty moves the transformation’s error from about 0.10 metres to about 0.09.
So the rule survives with its reason changed. Gather markers to condition the fit and to check it; do not gather them expecting the published accuracy to improve, because the published accuracy is measuring something else.
One sentence a published transformation could carry
Everything above reduces to a request. A published seven-parameter transformation states its parameters and an accuracy. It could state one number more: how much of that accuracy is the transformation’s own error.
It is computable from what the agency already has — the common points, the fitted parameters, and one evaluation over a clean grid — it does not require a new survey or a new standard, and it converts a figure everybody misreads into two figures that mean what they say. Six per cent and ninety-four per cent are very different instructions about what to do next.
Where the model stops
The distortion field is a stated quadratic, not a real network. The strain field is a smooth second-order function of position with an amplitude in parts per million, chosen because constant strain is absorbed by the scale parameter and linear strain by the rotations, so the leading term a fit cannot reach is the second. A real network’s distortion has power at many scales, so the six per cent above is a property of this field’s spectrum as much as of the arithmetic.
The noise control is independent between markers. Real observation errors in a triangulation are correlated along chains, which is precisely what makes them behave more like the strain case than like the noise case — so the two curves here bracket reality rather than describing it, and the real one is nearer the strain.
Height is not in it. Every marker here has zero ellipsoidal height and the distortion is horizontal. A real common point has a height whose datum is a different question again, and the vertical part of a datum shift is not small.
And the transformation error is measured over the same country. Applying the fitted parameters far outside the region the common points cover produces an error much larger than nine centimetres, because a seven-parameter fit extrapolates a rotation. That is the conditioning question rung three’s neighbour measures, and it is a different failure from this one.
The same shape, elsewhere on this site
Three other measurements on this site have the same structure and it is worth naming, because the shape recurs whenever a model is fitted to something it cannot represent.
A residual has more than one explanation finds a wrong datum and a wrong projection producing residuals that are indistinguishable below about a degree of extent — two mechanisms, one number, and no way to tell from the number which is which.
The rule scored out of sample scores a rule of thumb on a population it was not derived from, which is the same operation this rung performs on a transformation, and reaches the same conclusion: the in-sample number is a statement about the fitting rather than about the world.
And what a closed figure cannot see finds a class of error a traverse’s own closure check is blind to, which is the surveying version of a residual that cannot separate what it is made of.
Four measurements, one lesson: a small residual is evidence about the fit and not about the answer. It is the site’s own habit stated as a result rather than as a method.
Who found it, and when
The statistics are not in dispute and are not new. In-sample error understating out-of-sample error is the oldest result in model fitting; cross-validation dates from the 1930s and became routine in the 1970s; the distinction between a model’s parameter uncertainty and its approximation error is the bias–variance decomposition and is in every textbook.
None of it is standard practice in datum transformation, and the reason is worth stating without blame. A published transformation is a legal and administrative object as much as a statistical one: it has to be exactly reproducible, quotable, and identical for everyone who uses it, and a leave-one-out estimate is a second number that would have to be defined, agreed and maintained. What agencies did instead was to publish the grid shift, which solves the larger problem directly and makes the smaller one moot.
So the practice is right and the reported number means something other than it appears to. That is the same shape as two parameter sets, one transformation, where two agencies publish different parameters for the same pair of datums and both are correct — and it is the same underlying cause, which is that seven numbers are being asked to describe an object that has more than seven degrees of freedom.
Why a published number has constraints a statistic does not
The defence of current practice deserves more than a clause, because it names a genuine constraint that a purely statistical reading of the problem misses.
A published transformation has to be identical for everybody. Two parties computing coordinates from the same inputs must get the same answer, to the last digit, for as long as the transformation is current — because contracts, boundaries, planning consents and cadastral records depend on the answers agreeing. That requirement is stronger than accuracy and it is not negotiable.
A leave-one-out estimate is a different kind of object. It depends on the station set, the exclusion rule and the fitting procedure, all of which would have to be specified, agreed and frozen to be quotable — and any later addition of a station changes it. A quantity that improves as more data arrives is exactly what a legal reference cannot be.
So the agencies took the option that solves the larger problem outright. A grid-shift file carries the residual pattern the parameters cannot reach, is exactly reproducible, and makes the question of how much a further common point would buy irrelevant — because the shift file already holds what the further points would have contributed.
There is a second constraint that pulls the same way. A published transformation must also be stable in time, because a coordinate computed last year and one computed today have to agree. Any quantity that improves with the arrival of new observations forces a choice between staying current and staying consistent, and a reference system chooses consistency — which is why transformations are versioned and superseded rather than continuously updated, and why the version is part of the identifier.
That is a genuine trade rather than a compromise, and it is the reason the two constraints above are not obstacles to be engineered around.
The two constraints together explain a practice that looks conservative and is not. An agency publishing a transformation is not declining to use better statistics; it is meeting a requirement the statistics do not address, and the requirement is the reason the object exists. A transformation that was slightly more accurate and slightly different every year would be worse at its job than one that is fixed and known, because its job is to make two parties agree.
Which reframes what this rung’s number is for. It is not advice to agencies, who have a better answer. It is for anybody fitting their own transformation over their own region — a survey firm, a research project, an engineering scheme — where no published grid exists, the station set is theirs to choose, and the question is one more point worth observing is a real one with a budget attached.
Where the ladder goes next
The ladder has priced what a datum refers to, how it is fitted, where it leaves residuals, what its epoch means, how its parameters trade against each other, what a chain of them does, and what their own uncertainties are. This rung adds what more of them are worth.
The question it leaves open is the one it has just shown to matter most: how large is the ninety-four per cent, in a real network, and what does it look like? That is not answerable from a stated strain field. It is answerable from a published grid-shift file, which is exactly a map of the quantity — and reading one would be the first time this collection took geometry from a dataset rather than from a rule, which is a decision the site has declined twice before and would have to take deliberately.
Named alongside this one
Essays reaching for the same objects. Nobody chose these; they are what the concept index makes visible.
- The weights are a guess the solve believes conditioning · error budget · least-squares · noise · residual · verification
- Where the control points are common point · conditioning · error budget · least-squares · residual · verification
- A meridian boundary moves when its datum does datum · geodetic datum · helmert transformation · verification
- The height a coordinate does not carry datum · geodetic datum · helmert transformation · verification
- The nodes were evenly spaced least-squares · residual · sampling · verification
- A coordinate is the output of a solve datum · least-squares · residual
What links here
Every essay whose body links to this one.
The objects this essay names
Each one links to every other essay that touches it.
Common pointConditioningDatumError budgetGeodetic datumHelmert transformationLeast-squaresNetwork distortionNoiseResidualSamplingVerification