When a formula is not enough
Assumes Where a fit leaves residuals.
Ask a national mapping agency how to convert a coordinate from its own datum to a global one and the answer is not seven numbers. It is a file: a rectangular table of displacements at grid nodes, with a stated interpolation rule, distributed under a licence and versioned like software.
That is a strange artefact for a subject with as much closed-form mathematics in it as geodesy, and the reason is in where a fit leaves residuals: the relationship between two realisations is a rigid motion plus a strain field, and seven parameters cannot carry the second half. So the table is not a convenience wrapped around a formula. The table is the definition, and the formula is the compression of it that captures most of the variance.
Which raises the engineering question this essay answers. A table has a spacing. What does a stated tolerance cost in nodes?
The claim
The spacing a table needs is set by the roughness of what it tabulates, not by the size. A shift of a hundred metres can be represented on a coarse grid to millimetres. A ripple of thirty centimetres riding on that shift can defeat the same grid entirely.
That is not intuitive. The instinct is that a bigger thing needs more resolution to capture, and it is exactly backwards: what interpolation is bad at is curvature, and a large smooth field has less curvature per unit of amplitude than a small rough one.
What was computed, and how
The measurement needs a field with known content, so it is built rather than downloaded.
The smooth part is the real shift: the published OSGB36 transformation, applied through the exact Cartesian route at every point, converted to a local east and north displacement in metres. Over Britain it runs to about 99 metres on the ground, and it is smooth because it is the projection of a rigid motion onto a smooth surface — its variation across the country comes only from the changing orientation of the local frame.
The rough part is stated: a sinusoid of stated amplitude and stated wavelength added to both components. That stands for the network strain a real file exists to carry, and it says so.
Then the tabulation is simulated honestly. Sample the field at nodes of a given spacing; build the table; interpolate bilinearly back at 1,681 probe points spread over the region; and report the worst error, not the mean, because a table’s specification is a guarantee rather than an average.
The smooth field alone:
| spacing | nodes | worst error |
|---|---|---|
| 4° | 16 | 238 mm |
| 2° | 36 | 57.5 mm |
| 1° | 100 | 15.6 mm |
| 0.5° | 361 | 3.9 mm |
| 0.25° | 1,369 | 0.98 mm |
| 0.125° | 5,329 | 0.24 mm |
Every halving divides the error by almost exactly four. That is the signature of bilinear interpolation, whose error is second order in the spacing, and the fact that it comes out at 4.00 rather than at 3.6 or 4.5 is the check that the tabulation is being simulated correctly rather than something else being measured.
Why the ripple costs so much
Add the ripple — 30 centimetres, two degrees across — and the trajectory changes shape.
At 2° spacing, the error is 326 millimetres: the table samples the ripple at exactly its own wavelength and reproduces essentially none of it, so the error is the ripple’s own amplitude. At 1° it is 309 millimetres, barely better, because a sinusoid sampled at half its wavelength is still nearly invisible to linear interpolation. Only at a quarter of a degree does the table begin to resolve it, and only at an eighth does the error fall under twenty millimetres.
The controlling quantity is nodes per wavelength, and bilinear interpolation needs something like sixteen of them before its error is small. The size of the field does not enter at all: doubling the ripple’s amplitude doubles the error at every spacing and changes the required spacing not at all.
So the comparison at a fixed spacing of one degree is stark. The 99-metre shift is reproduced to 15.6 millimetres. The 0.3-metre ripple, which is three parts in a thousand of it, contributes 293 millimetres of the 309. The part that costs the table almost everything is the part that is almost none of the answer.
The same trajectory, without the ripple
Setting that beside the hero figure isolates the variable. The second ripple has a sixth of the amplitude and three times the wavelength, and the wavelength is doing almost all of the work: a field that varies gently is cheap to tabulate at any amplitude, and a field that varies quickly is expensive at any amplitude.
This is why the generator takes both the amplitude and the wavelength rather than one number called “roughness”. A single figure would have conflated the two quantities that the whole argument is about separating.
What the seven parameters would have cost instead
It is worth pricing the alternative. Applying the seven-parameter transformation alone, over the same ground, leaves the residual measured in where a fit leaves residuals: about 1.6 metres RMS and 3.0 metres at worst, for four parts per million of strain.
Put on the same axis as the table:
- seven parameters: 3,000 mm worst, 7 numbers;
- a 1° table: 309 mm worst, 100 nodes;
- a 0.125° table: 11 mm worst, 5,329 nodes.
Three orders of magnitude of accuracy for three orders of magnitude of nodes, which is the trade being made and is a good one at modern storage costs. It was not a good one on a floppy disc, and the first generation of these files was coarser than their authors wanted for exactly that reason.
What real files look like
The numbers above line up with what the agencies actually ship, which is the check that the model is about the right thing.
Every one of them is a response to the same measurement, and the shape of that measurement is where a fit leaves residuals: a rigid motion taken out, metres left behind, in a pattern. Britain’s national transformation is a grid at one kilometre, quoted as good to about a centimetre. Canada’s and the United States’ NTv2-format files are typically at a few minutes of arc, with denser sub-grids over regions whose networks are worse. Australia’s are at a minute. None of them is at a degree, and the reason is exactly the one measured here: the strain they carry has structure at the scale of individual survey chains — tens of kilometres — and sixteen nodes per wavelength of that is a kilometre-scale grid.
Two design features follow from the same arithmetic and are worth naming because they look like implementation detail and are not.
Sub-grids. NTv2 allows a coarse parent grid with finer children over selected areas, because roughness is not uniform: a network is worse where it was observed badly, and refining everywhere to suit the worst region multiplies the file by the square of the ratio.
A stated interpolation rule. The file specifies bilinear, and that is part of the definition rather than advice. A user who interpolates bicubically gets different numbers — smoother, and not the ones the agency computed its accuracy statement for. The rule and the nodes together are the transformation; either alone is not.
That ordering is the reason a table exists at all, and it is worth restating as three numbers on one scale. Applying the shift by the exact route rather than by Molodensky’s shortcut buys seven millimetres. Applying a table rather than the seven parameters buys three metres. The effort went where the error was.
Where the nodes come from
The arithmetic above treats the table’s values as available at any spacing, which is the one place the model is friendlier than reality.
A node’s value is the difference between two realisations at that node, and neither realisation is a formula. The old one is a set of published marker coordinates; the new one is a set of observed positions. So a node’s value is only known where there is a mark that has been coordinated both ways, and everywhere else it is itself an interpolation — of a different kind, from a network adjustment, with its own error.
This is why refining a grid is not free even though storage is. The limit is the density of doubly-coordinated marks, and past that density a finer grid interpolates its own interpolation — smoother output, no more information, and an accuracy statement that is no longer supported by anything. An agency that publishes a kilometre grid where its marks are ten kilometres apart is making a claim about the smoothness of the strain field, not about its measurements.
Where the model stops
The rough component is synthetic. Its amplitude and wavelength are stated inputs standing in for network distortion, and the demonstration is about how interpolation responds to roughness rather than about any particular country’s error. Real distortion is not a single sinusoid; it has a spectrum, and the spacing a real file needs is set by the short-wavelength end of that spectrum, which is what makes the honest version of this calculation an agency-specific one.
Worst error over a finite probe set. The 1,681 probe points are dense compared with the coarser spacings and comparable to the finest, so the reported worst at 0.125° is a slight underestimate. It is reported anyway because the trajectory’s shape is the argument and the shape is unaffected.
Bilinear only. Higher-order interpolation converges faster on a smooth field and does not help on an under-sampled rough one — a bicubic scheme still cannot reconstruct a sinusoid sampled twice per wavelength, because the information is not in the samples. The Nyquist limit is a statement about the data rather than about the method, which is why “use a better interpolator” is the wrong answer to an under-sampled table.
The Nyquist limit applies to the survey too. If the strain field has structure at ten kilometres and the marks are twenty kilometres apart, no table at any spacing can carry it, because the observations never contained it. That is a limit on the data and it cannot be repaired downstream — the same shape of limit as the one distortion over a region records for a distortion measure computed from too few samples.
Nothing here is about file size. A 5,329-node table is a few hundred kilobytes and nobody cares. What the node count actually costs is observation: every node’s value has to come from somewhere, and in a real file it comes from a network adjustment over marks that had to be visited. The grid’s density is limited by the survey, not by the disc.
The generalisation
Three portable statements, and the third is the one that gets forgotten.
Interpolation error is set by the second derivative, not the value. A field’s amplitude scales the error linearly and its curvature scales it quadratically in the spacing, so a small rough component beats a large smooth one whenever the roughness is short enough. The design question for any lookup table is therefore what is the shortest wavelength in this field, and never how big is it.
The convergence rate is the check that the machinery is right. Bilinear interpolation must fall by four when the spacing halves. Measuring 4.00 across five halvings is not a result about datums; it is the evidence that the simulated tabulation is a tabulation. A trajectory that fell by three would mean something else was being measured — and this is the same instrument used in the ellipsoid is a level surface, where a quadrature’s residual is required to fall by four when the sample count doubles before the agreement it reports is believed.
A table is a model with as many parameters as nodes. That sounds like a criticism and is the point. The progression from three parameters to seven to a grid is a progression in how much of the truth the model can express, and each step is taken when the previous model’s residual exceeds the tolerance of the work. That is the same discipline every projection minimises something applies to the other half of this subject: a model is chosen against a stated objective, and saying which objective is most of the work. There is nothing unprincipled about the end of it. What would be unprincipled is a table with no stated interpolation rule, or a parameter set with no stated residual — in both cases a model presented without the thing that says how much of the truth it holds.
The right way to hold the two artefacts together is that the parameters are a summary and the table is the record. A summary is portable, quotable and approximately right; a record is large, licensed and exact to its own stated accuracy. Every field that has both eventually stops arguing about which is correct and starts labelling which is which — which is all this essay is asking of a published transformation, and is what the seven parameters asks in the other direction.
A rougher field, and a coarser one
Three trajectories across this essay, at three roughnesses, and the smooth curve is identical in all of them. That is the claim stated as a picture: the shift being tabulated does not enter the spacing decision at all.
Interpolating a grid, and the order it converges at
A shift grid is interpolated bilinearly and the essay establishes that the error falls by four for every halving of the spacing — second order, which is the signature of the method and is asserted rather than assumed.
The applied field measures the same property of the same family of methods in a different setting: resampling a raster when it is warped into another projection. Fitted over five grid resolutions, the orders come out at 1.004 for nearest-neighbour, 1.983 for bilinear and 2.930 for a Catmull–Rom cubic — the same second order for the same method, arrived at through a picture rather than a shift table.
The same family of methods has a second cost there that a shift grid does not pay. Warping a raster into another projection and back moves no coordinate and loses contrast anyway, because a target cell’s centre does not fall on a source cell’s — whereas a shift grid is interpolated once, at the point asked for, and never resampled onto a lattice of its own — reprojecting a raster invents values is the measurement of the difference.
Who found it, and when
Grid-shift files are a product of the 1980s and 1990s, and each country arrived at one by the same route: a new space-based realisation, a comparison against the existing network, a seven-parameter fit, and a residual too large to publish as an accuracy figure.
The United States was first at scale, with NADCON in 1990 — a grid relating NAD27 to NAD83, at 15 minutes, distributed on floppy disc. Canada’s NTv2 followed and became the de facto interchange format, largely because it solved the sub-grid problem cleanly. Britain’s OSTN02 arrived in 2002 and its successors since; Australia’s grids came with GDA94.
The intellectual content is older than any of them and belongs to signal processing rather than to geodesy: how finely a field must be sampled to be reconstructed is Nyquist’s question, answered in 1928, and the answer is that the sampling rate is set by the highest frequency present and by nothing else. Geodesy met it in the form of a question about why a table has to be a kilometre when the shift it carries is a hundred metres, and the answer had been in another discipline for sixty years.
The name has drifted in an unhelpful direction. These are usually called transformation grids, which suggests a correction applied to a transformation. What they are is the transformation, in its full form, with the parameter set as the approximation.
The drift matters because the name decides what a user thinks they may skip. A transformation grid sounds optional in a way a transformation does not, and a pipeline that applies the parameters and omits the table has, under the honest name, simply not performed the transformation.
Where this goes next
Three of a datum’s four commitments have now been taken apart: the shape, the placement, and the realisation. The fourth is the one that was invisible until measurement caught up with it. The epoch is part of the coordinate puts a velocity on the ground under every marker, and finds that a coordinate without a date goes out of tolerance within about five years.
Named alongside this one
Essays reaching for the same objects. Nobody chose these; they are what the concept index makes visible.
- The third coordinate moves too datum · osgb36 · realisation · tolerance · verification
- A coordinate without its system is not a location datum · osgb36 · realisation · verification
- An area on the grid is not an area on the ground national grid · osgb36 · tolerance · verification
- One pair of numbers, a hundred and twenty places national grid · realisation · tolerance · verification
- The answer is a set realisation · residual · tolerance · verification
- The plumb line is not the normal datum · realisation · tolerance · verification
What links here
The 8 essays that link to this one and share the most of its objects, of 12 that link here.
The objects this essay names
Each one links to every other essay that touches it.
DatumGrid shiftInterpolationNational GridNetwork strainOSGB36Quadratic lawRaster gridRealisationResidualToleranceVerification