The parameters are not independent
The seven parameters have their own uncertainty gives each of the seven a standard deviation and propagates them into a position on the ground. It is the rung that turns a parameter table from seven exact numbers into seven measurements, and it is the first place this ladder treats a published transformation as a fitted object rather than as a definition.
It stops at seven numbers, and a covariance matrix has twenty-eight.
Why the off-diagonal is not small
The reason is geometric and it can be stated before any arithmetic.
A rotation of the Earth about an axis through its centre moves a point by ω × r. For a region a long way from the axis, over an extent much smaller than the Earth’s radius, that displacement is very nearly the same at every point in the region — which is a translation.
So over a national area, a rotation and a translation are almost the same motion. A fit given a set of common points cannot tell them apart: shifting one and compensating with the other leaves the residuals essentially unchanged, and the least-squares problem has a nearly flat direction.
A nearly flat direction is a nearly singular normal matrix, and a nearly singular normal matrix has enormous, strongly correlated entries in its inverse. That inverse is the covariance.
What was computed, and how
A Helmert fit to sixty-four stated common points over a nine-degree region, with the covariance from σ₀²(AᵀA)⁻¹ and the units converted back to metres, arcseconds and parts per million — so that the numbers can be read beside a published parameter table.
| pair | correlation |
|---|---|
| tᵧ against rₓ | +0.940 |
| tₓ against rᵧ | −0.849 |
| tᵧ against r_z | −0.817 |
| t_z against s | −0.768 |
| t_z against rᵧ | +0.636 |
| rₓ against r_z | −0.573 |
| tₓ against s | −0.524 |
Seven of the twenty-one entries are above a half. The three largest are translation against rotation, exactly as the geometry predicts, and the fourth is translation against scale — which is the same argument in the radial direction, since a scale change moves a region radially and so does the radial component of a translation.
The condition number of the normal matrix is 4.5 × 10¹⁶. In double precision that leaves under one significant figure of the smallest eigendirection, which is the numerical statement of the same fact.
What a published table lets a reader do
Here is where it stops being bookkeeping.
A published transformation gives seven parameters and, at best, seven standard deviations. The only propagation that supports is the one that treats them as independent: add the seven contributions in quadrature and take the root.
The diagonal-only figure is larger everywhere, by a factor between 20 and 36.
That is the opposite of the expected direction and it is worth being clear about why. The errors in the seven parameters are not independent errors that happen to be correlated; they are the same error, expressed in whatever mixture of parameters the fit happened to choose. Over the region the fit was made on, they cancel, because that cancellation is what the fit optimised. Take the parameters apart and add their contributions in quadrature and the cancellation is thrown away.
So a reader with a published table and no covariance computes an error bar that is right in units, right in shape and wrong by a factor of thirty. And it is wrong in the safe direction over the fit’s own region — which is the direction that is hardest to notice, because the position is well inside it.
Outside the region the cancellation weakens, and a full treatment would show the two converging and eventually crossing. Nothing here measures that, and it is the one place the diagonal-only figure could be optimistic rather than pessimistic.
The correlation is about the region
| extent | condition number | strongest correlation |
|---|---|---|
| 1° | 3.7 × 10¹⁸ | 0.9407 |
| 4° | 2.3 × 10¹⁷ | 0.9405 |
| 9° | 4.5 × 10¹⁶ | 0.9401 |
| 20° | 8.9 × 10¹⁵ | 0.9376 |
| 45° | 1.6 × 10¹⁵ | 0.9245 |
| 90° | 2.9 × 10¹⁴ | 0.7738 |
The conditioning improves monotonically and by four orders of magnitude. The ill-conditioning is a fact about how much of the Earth was observed rather than about how well, and no amount of extra points inside a small region improves it — which is what another common point buys reaching the same conclusion by a different route, and the network’s answer is decided before it is measured stating it as a general property of a design.
The correlation itself is remarkably flat below twenty degrees: 0.9407 at one degree and 0.9401 at nine. That is the geometric argument being exactly right — a rotation and a translation are indistinguishable over any patch small compared with the Earth, and how small hardly matters until the patch stops being small.
The check that the near-degeneracy is real
A condition number of 4 × 10¹⁶ is close enough to the limits of double precision that it deserves a control rather than a footnote, because a matrix that ill-conditioned can produce a correlation matrix that is arithmetic rather than geometry.
Two things separate them. The correlation is stable across the region sizes — 0.9407 at one degree and 0.9401 at nine, while the condition number moves by nearly two orders of magnitude over the same range. A number that came from rounding would move with the conditioning; a number that comes from the geometry does not.
And the pair is the predicted pair. The geometric argument names which translation should correlate with which rotation: a rotation about the x axis moves a region in the y direction, so tᵧ against rₓ. That is the top row of the table, and it is asserted rather than observed — a fit whose strongest correlation was between two translations, or between a rotation and the scale, would be reporting something other than the near-degeneracy this rung is about.
What follows for practice
A published parameter table is not a complete statement of a transformation’s uncertainty, and cannot be made into one by adding standard deviations to it. The missing object is a matrix, it has twenty-one more numbers in it, and those numbers are large.
The recommendation that follows is unusually cheap: publish the covariance. It is twenty-eight numbers rather than seven, it comes out of the fit at no cost, and every serious transformation is fitted by software that has it in memory and discards it.
Failing that, there are two partial repairs and both are worse.
Publish the parameters in a decorrelated form. The eigenvectors of the covariance are combinations of translations, rotations and scale that are independent, and quoting those seven numbers with seven standard deviations would be a complete statement. It is also unusable: nobody’s software takes a transformation in that form, and the two sign conventions already show what happens when a parameter set does not carry its own interpretation.
Publish a positional accuracy instead. Quote the transformation’s uncertainty as a figure on the ground over a stated region, which is what a national agency usually does. That is honest, it is what a user needs, and it throws away the ability to propagate anywhere else.
Why the seven are the seven
It is worth asking why the transformation is written in these parameters at all, given that they are so badly separated.
The answer is that the seven are the ones with physical names. A translation is a shift of the coordinate origin, a rotation is a misalignment of the axes, a scale is a difference in the unit — each has a cause somebody can point at, and a fitted value that can be argued about. The decorrelated combinations have none: they are directions in a seven-dimensional space with no interpretation.
So the parameterisation is chosen for interpretability and pays for it in conditioning, which is a trade this collection has met before with a different object. Report the map, not the parameters says the same thing about an aspect search: the parameters are how the answer is written down and the map is what the answer is, and the two have different stabilities.
The difference here is that nobody can report the map. A transformation is its parameters as far as any software is concerned, so the interpretable-but-correlated form is the only one that gets used, and the correlations have to travel with it or be lost.
Where the model stops
The common points are stated. Sixty-four markers on a grid with stated noise, rather than a real network with real geometry. The correlation structure is a property of the positions rather than of the noise, so it would survive any realistic point set; the absolute standard deviations would not.
σ₀ comes from the fit’s residual. A real fit’s residual contains network distortion as well as observation noise — which is where a fit leaves residuals — so a real σ₀ is larger and every standard deviation above scales with it. The correlations do not, because they are ratios.
And the fit is unweighted. Weighting the common points changes AᵀPA and therefore the covariance. It does not change the near-degeneracy, which is in A alone.
What a user should actually do
The recommendations above are for whoever publishes a transformation. A user has a table of seven numbers, no covariance, and a job to do, and the practical guidance is short.
Inside the transformation’s stated region, trust the agency’s positional accuracy figure and not a propagation. The agency has the covariance; its accuracy figure is the honest summary of it; a propagation from the diagonal is thirty times too large and is worse than no number.
Outside it, do not propagate at all. The correlations that make the diagonal pessimistic inside are a property of the fit’s own region, and they weaken with distance from it. A propagation outside the region has no known sign of error, and the transformation itself has no claim to apply there — which is what the region is for.
And when two transformations both claim a place, the difference between them is the honest error bar. That is a measurement rather than a propagation, it needs no covariance, and it is available to anybody with two published parameter sets: two parameter sets, one transformation is the rung that prices it.
One clarification, because the direction of the error is easy to misremember. The diagonal-only propagation is too LARGE over the region the transformation was fitted on — it throws away a cancellation the fit created — so a user who computes it is being conservative rather than reckless. That is the comfortable case. Whether it stays conservative outside the region is not measured here and has no obvious answer, which is a reason to treat the stated validity extent as a hard edge rather than as a suggestion.
The generalisation
A parameter is not a measurement; a parameter set is.
This is the same statement two parameter sets, one transformation makes about the values, arriving now about their uncertainties. There the point was that two different seven-tuples describe the same map, because the parameters are not separately identifiable. Here the point is that the errors in the seven are not separately meaningful either, and for the same reason: the quantity that exists is the transformation, and the seven numbers are a coordinate system for it.
The collection meets the shape elsewhere. The weights are a guess the solve believes is about a parameter whose value is not identifiable from the data; and the datum hides inside the projection’s parameters is the same failure between two different kinds of parameter.
The transferable form: whenever a quantity is reported as a list of parameters, ask what the fit could not separate, because the answer is where the correlations are and they are the part that gets dropped.
One number worth carrying
If a single figure survives this rung it should be the condition number, because it is the one that converts a qualitative statement into an engineering one.
At 4 × 10¹⁶, a double-precision solve of the normal equations retains under one significant figure in the weakest direction. That is not a warning about the arithmetic — a modern implementation uses a QR or an SVD and never forms the normal matrix at all — but it is an exact statement of how much information the data contains about that direction, and the answer is nearly none.
So the seventh parameter of a regional seven-parameter fit is, in a precise sense, not determined by the data. It is determined by whichever combination the solver happened to pick, and a different solver picking a different combination gets a different seven-tuple describing the same transformation to the same accuracy. That is the mechanism behind the rung two below this one, and this is its numerical form.
Who found it, and when
The near-degeneracy of a seven-parameter transformation over a small region is standard geodesy and is why national agencies fit over their whole territory rather than over a district. The full covariance is produced by every least-squares implementation, and the recommendation to publish it appears in the technical literature regularly.
What is unusual is how rarely it happens. The registries that carry transformation parameters — the EPSG dataset foremost — have fields for the parameters and for an accuracy figure, and no field for a covariance. So the information exists in the office that fitted the transformation, is discarded at publication, and cannot be recovered by anybody downstream — which is the same omission a map does not say what it is records about a drawn map, in a register rather than on a sheet.
Parameters are not comparable; transformations are
The near-degeneracy has an immediate consequence for a thing users do constantly, and it is worth stating as a prohibition because the alternative is easy.
Two published parameter sets for the same transformation cannot be compared parameter by parameter. The seventh is not determined by the data, so two offices fitting the same region against the same points can produce seven-tuples that differ substantially and describe the same transformation to within its accuracy. A table putting them side by side and highlighting the differences is highlighting the solvers.
The comparison that means something is between the transformations. Take a set of test points spanning the region, push them through both, and look at the differences in position. That is the quantity a user experiences, it is in metres rather than in parts per million and arcseconds, and it is insensitive to which combination each solver happened to land on.
The test costs nothing and needs no covariance, which is what makes it the right advice for a world in which the covariance is not published. A user holding two candidate transformations does not need to know why the parameters differ; they need to know whether the answers do.
And the outcome is usually reassuring, which is the point. Parameter sets that look alarmingly different mostly agree to a few centimetres on the ground, and that is the fact a user needs before deciding whether a discrepancy is worth investigating. The alarming appearance is an artefact of reporting a transformation by a parameterisation that has a nearly flat direction in it.
The one case where the parameters do matter is when somebody intends to interpolate or extrapolate them — to blend two neighbouring transformations, or to apply one outside the region it was fitted over. Both operate on the parameters directly, both are sensitive to exactly the combination that is undetermined, and both are the reason the covariance would have been worth publishing.
Both are also the operations a user is most likely to attempt without realising they have left what the fit supports.
Where the ladder goes next
Ten rungs price a transformation’s parameters, their fit, their residuals, their uncertainty, their conventions and now their dependence on each other. What none of them prices is the transformation’s own validity region: every fit above is stated over a box, every published transformation carries one, and what happens at its edge — where two neighbouring transformations both claim to apply and disagree — is a discontinuity in a quantity that has no business having one.
Named alongside this one
Essays reaching for the same objects. Nobody chose these; they are what the concept index makes visible.
- The difference of two coordinates correlation · covariance · datum · least-squares · precision
- The answer is a set degeneracy · identifiability · least-squares · precision
- Where the control points are conditioning · error budget · identifiability · least-squares
- A coordinate is the output of a solve covariance · datum · least-squares
- The error ellipse is not an ellipse covariance · precision · propagation
- The error that does not average down correlation · error budget · precision
The objects this essay names
Each one links to every other essay that touches it.
ConditioningCorrelationCovarianceDatumDegeneracyError budgetHelmert transformationIdentifiabilityLeast-squaresPrecisionPropagationSeven parameters