Grids, and what a survey does

A coordinate is the output of a solve

Six essays measure a tape, close a traverse, spread a misclosure and reduce a chain. The coordinate that comes out of the far end is the solution of a least-squares problem, and the problem has a decision in it that is not a measurement: what to hold fixed. Change it and every coordinate moves by centimetres while not one residual moves at all.

The reduction ladder has followed a measurement from the tape to the page: what a tape actually measures, why a traverse must close, how a misclosure is spread, what a tolerance decides, what a closed figure cannot see, and what chain a satellite does not have. At the end of all of that a coordinate is written down and the ladder stops.

It should not. The coordinate is not the end of a chain of arithmetic; it is the solution of an overdetermined system, chosen by a criterion, subject to constraints that are not observations. Everything a surveyor is told about the quality of the work — the residuals, the variance factor, the error ellipses — comes out of that solve, and so does one thing that is not about quality at all and is reported as though it were.

Five stations, ten distances, three spare. A braced quadrilateral with a centre point. Every distance between the corners and every distance to the centre is observed, 10 in all, each with a standard deviation of 8 mm. Holding one station and one bearing leaves 8 unknown coordinates, so the network has three degrees of freedom: three independent statements the observations make that could be contradicted. Everything the adjustment can tell anybody about the quality of the work comes out of those three.
Fig. 1 A braced quadrilateral with a centre point: five stations, all six distances between the corners and all four to the centre, each observed with a standard deviation of 8 mm. Holding one station and one bearing leaves eight unknown coordinates, so the network has three degrees of freedom — three independent statements the observations make that could be contradicted. Everything the adjustment can say about the quality of the work comes out of those three.

More observations than unknowns, and no solution

Ten distances, ten coordinates, and a datum defect of three: a network of distances in the plane is free to translate in two directions and to rotate, because none of those changes any distance. Hold a station and a bearing and eight unknowns remain, against eleven equations. There is no set of coordinates that satisfies all of them, because the observations disagree — they always disagree, which is why the network was braced.

The criterion that picks one answer is least squares: minimise the weighted sum of squared residuals, with the weight of each observation the reciprocal of its variance. That is a choice, and it is the choice that makes the answer the maximum-likelihood estimate if the errors are independent and Gaussian — a hypothesis about the field work that the adjustment assumes and then, in one number, tests.

The mechanics are three lines. Linearise the observation equations about approximate coordinates to get a design matrix A; form the normal matrix N = AᵀPA and the right-hand side AᵀPl; factor and solve. The factorisation is Cholesky, and its failure is informative rather than annoying: a normal matrix that is not positive definite means the constraints did not remove the datum defect, so the network has a direction it can move in for free. That is a mistake in the problem, and it is worth having the arithmetic say so rather than returning a number.

The redundancy numbers, which are the answer to a question nobody asks

The residuals are the visible output. Underneath them is a quantity that says how much each observation was actually checked, and it is the most useful number in an adjustment report that most reports do not print.

How much of each observation the network can see. An observation's redundancy number is the share of its own error that shows up in its residual; the rest goes into the coordinates. They run from 0.144 to 0.467 here and sum to 3.000000, which is the network's three degrees of freedom — not approximately, identically. The diagonals are the best checked because they are the only observations with two independent routes; the four sides of the quadrilateral are the worst, and a gross error in one of them shows barely a seventh of itself.
Fig. 2 Each observation’s redundancy number: the share of its own error that shows up in its residual. They run from 0.144 to 0.467 here and sum to 3.000000 — the network’s three degrees of freedom, not approximately, identically. The diagonals are the best checked because they are the only observations with two independent routes; the four sides of the quadrilateral are the worst, and a gross error in one of them shows barely a seventh of itself.

The identity is what makes them interpretable. Each rᵢ is that observation’s share of the network’s total checking power, and the shares add to the number of independent checks the network contains. An observation with r near one is fully checked; one with r near zero is not checked at all, and its value goes straight into the coordinates.

The consequence is the reason to print them:

A gross error splits in the redundancy number's proportion. Each of the ten observations was given a 150 mm gross error in turn, and the change in its own residual measured. It is −r∇ every time, to 1.2 micrometres, which is the nonlinearity of the distance equation over that displacement and not a fitting error. The rest, (1 − r)∇, is absorbed by the coordinates and reported as nothing at all. The worst-checked observation here hides 128 mm of 150.
Fig. 3 Each of the ten observations given a 150 mm gross error in turn, and the change in its own residual measured against the prediction −r∇. It is exact to 1.2 micrometres, which is the nonlinearity of the distance equation over that displacement rather than a fitting error. The rest, (1 − r)∇, is absorbed by the coordinates and reported as nothing at all.

The identity v = −r∇ is the whole content of a redundancy number. The worst-checked observation here, D–A at r = 0.144, shows 21.6 mm of a 150 mm error and hides 128.4 mm of it. The best-checked, B–D at r = 0.467, shows 70.0 mm and hides 80.0.

A 150 mm error in the least-checked observation. D–A has the lowest redundancy number in the network, 0.144. Adding 150 mm to it changes its own residual by -22 mm — which is inside what a careless reader would call ordinary — and moves the coordinates by up to 149 mm. The displacement is drawn 900× so it can be seen at all. Nothing in the adjustment report says which station moved, because from inside the solve nothing did: the network simply found the coordinates that best fit the observations it was given.
Fig. 4 The same 150 mm error in the least-checked observation, with the resulting displacement of every station drawn 900× so it can be seen. The residual it produces is 22 mm — inside what a careless reader would call ordinary field work — and the coordinates move by up to 149 mm. Nothing in the adjustment report says which station moved, because from inside the solve nothing did: the network found the coordinates that best fit the observations it was given.

Turn the identity round and it gives a minimal detectable blunder: the smallest gross error that would produce a standardised residual large enough to be flagged, which is δ₀σ/√r for whatever detection threshold δ₀ the practice uses. At the conventional 4.13 that is 48 mm on the diagonals and 87 mm on the sides of this quadrilateral — a factor of 1.8 between two observations made with the same instrument on the same afternoon, decided entirely by where they sit in the figure.

The datum is not a measurement, and here is the arithmetic that says so

Now the part the ladder came for. Three of the eleven equations are not observations at all: they are the statement that station A is where it is and that the bearing to B is what it is. Those numbers came from somewhere else.

Four datums, one set of residuals. The same 10 observations solved four times, holding a different station and bearing each time. Every coordinate in the network moves, by up to 11.0 mm. Not one residual moves: the four sets agree to 4.5e-10 mm, which is arithmetic noise. That is the arithmetical content of the datum is not a measurement — the observations decide the shape and the datum decides where the shape is put, and a reader can tell the two apart without seeing the constraint list, because an over-constrained network fails this test.
Fig. 5 The same ten observations solved four times, holding a different station and bearing each time. Every coordinate in the network moves, by up to 11.0 mm. Not one residual moves: the four sets agree to 5 × 10⁻¹³ mm, which is arithmetic noise. That is the arithmetical content of the datum is not a measurement.

This is the test that separates a minimal constraint from an over-constraint without looking at the constraint list, and it is worth stating as a rule because it is checkable from an adjustment report:

Under minimal constraints the residuals do not depend on which minimal constraints they were. The observations decide the network’s shape, entirely; the datum decides where the shape is put, entirely; and the two do not interact. A report whose residuals change when the fixed station changes is a report of an over-constrained network, whatever its constraint list says.

Over-constraining, and where it gets reported

The reason anyone over-constrains is respectable. The extra stations have published coordinates, the new work has to join the existing framework, and holding them is how the join is made. What happens then is that the network is required to reproduce numbers that came out of somebody else’s adjustment, with somebody else’s residuals in them.

The same observations, two datums. Hollow: the network on minimal constraints — one station and one bearing, the least that can be held. Filled: the same observations required also to reproduce C's published coordinates, which sit 72 mm from where this network would put them. Every station moves, by up to 83 mm at C, drawn 900× — and the variance factor rises from 1.159 to 1.619, which is the only place the report mentions it, and it mentions it as though the field work were worse.
Fig. 6 Hollow: the network on minimal constraints. Filled: the same observations required also to reproduce C’s published coordinates, which sit 72 mm from where this network would put them. Every station moves, by up to 82.7 mm at C itself, drawn 900×. The variance factor rises from 1.159 to 1.619 — and that is the only place the report mentions any of it, and it mentions it as though the field work were worse.

The residuals move too, by up to 14.9 mm on one diagonal, and they move in a pattern that has nothing to do with which observation was good. A reviewer reading the report sees a network with a variance factor of 1.6 and residuals up to twice what the instrument’s specification suggests, and the natural conclusion — that somebody’s field work was sloppy — is wrong. The field work was fine. The framework it was tied to disagreed with it by 72 mm — which is what a chain of transformations that does not close looks like when somebody else has already absorbed it, and the adjustment had nowhere to put that except into the observations.

The distinction has a name in practice — a free or minimally constrained adjustment against a constrained one — and the recommended procedure is to run both: the first to judge the observations, the second to deliver coordinates in the required frame. What the two-run procedure amounts to is refusing to let one number answer two questions.

The variance factor, and the four things it can mean

One number in the report claims to summarise the whole adjustment, and it is worth being explicit about what it can and cannot distinguish.

The variance factor is the weighted sum of squared residuals divided by the degrees of freedom, and its expectation is one if the observations really do have the standard deviations they were given. Here it is 1.159 on the minimal adjustment — three degrees of freedom, so a value between about 0.2 and 2.6 is unremarkable — and 1.619 once C is held.

What a value above one can mean, in the order a reviewer should consider them:

  • the instrument’s stated precision is optimistic, so every weight is too large;
  • one observation is a blunder, and its squared residual is carrying the sum on its own;
  • the network is over-constrained, and the residuals are absorbing a disagreement between frameworks;
  • the functional model is missing a term — an unmodelled scale error, an atmospheric correction — so the residuals are systematic rather than random.

The four have different remedies and the variance factor cannot tell them apart. What separates them is the pattern of the residuals, which is why the standardised residuals and the redundancy numbers belong in the report beside it, and why the free adjustment has to be run first: three of the four diagnoses are unavailable once the constraints have had a chance to redistribute the evidence.

What an undetected blunder does to the coordinates

The minimal detectable blunder says how large an error has to be before the report flags it. The question a client actually has is the other one: if a blunder that large slipped through, how far would it have moved the answer? Baarda’s pair of names for these is internal and external reliability, and the second follows from the same identity in one line.

An error ∇ shows −r∇ in its own residual and puts (1 − r)∇ into the coordinates. The largest error that escapes detection is ∇₀ = δ₀σ/√r, so the largest coordinate displacement an undetected blunder can produce is

(1r)0=δ0σ1rr.(1-r)\,\nabla_0 = \delta_0\sigma\,\frac{1-r}{\sqrt{r}}.

On this quadrilateral, with σ = 8 mm and a detection threshold of 4.13, that is 74.6 mm for the least-checked observation at r = 0.144 and 25.7 mm for the best-checked at r = 0.467. So the network’s exposure is about three times worse on its sides than on its diagonals — a ratio of 2.9, against a ratio of 1.8 in the detectable blunders themselves.

The two ratios pointing the same way but not by the same amount is the part worth carrying. The function (1 − r)/√r falls faster than 1/√r does, so a poorly checked observation is doubly penalised: the blunder in it has to be larger before anybody notices, and a larger share of whatever is there goes into the coordinates rather than into the residual. Internal reliability alone understates the difference between the two classes of observation in this figure, and it is the number more usually quoted.

It also says where the effort belongs. Improving the instrument lowers σ and scales both quantities together, so it buys a proportional improvement everywhere and changes no ratio. Adding an observation changes the redundancy numbers — every r rises, since they must sum to the new degrees of freedom — and it raises them most where they were lowest, because a new tie gives a previously unchecked line a second route. A better instrument improves the network uniformly; a better figure improves it where it is worst, and only the second changes which observations the report can be trusted about.

What the network knows about where it is

The last output of the solve is the one that connects this ladder to the precision one.

What the network knows about where each station is. The inverse of the normal matrix has a 2 × 2 block for every unknown station, and each block is an ellipse — in metres already, because the design matrix of a distance is a pair of direction cosines and carries no units. Drawn 2600×, the semi-major axes run from 7.36 mm to 11.12 mm against observations of 8 mm, with axis ratios up to 2.08. B is the exception and is not a measurement: its bearing from the held station is a datum constraint, so it has no freedom at all across that line and its ellipse is a segment. They are the same object as an error ellipse pushed through a projection, arriving from the other end — there a known covariance is mapped, here an unknown one is inferred from the geometry of what was observed.
Fig. 7 The inverse of the normal matrix has a 2 × 2 block for every unknown station, and each block is an ellipse — in metres already, because a distance’s design row is a pair of direction cosines and carries no units. The semi-major axes run from 7.8 mm to 11.1 mm against observations of 8 mm, with axis ratios up to 2.08. B is the exception and is not a measurement: its bearing from A is a datum constraint, so it has no freedom across that line at all and its ellipse is a segment.

These are the same object as an error ellipse pushed through a projection, arriving from the other end. There, a known covariance is mapped and the projection’s own anisotropy changes its shape. Here, the covariance is not known: it is inferred from the geometry of what was observed, and its anisotropy is a statement about which directions the observations constrained well — the same reading the indicatrix gets on a map.

The B ellipse is the useful oddity. Its axis ratio is 7,356, which is not a measurement of anything — it is the datum constraint appearing as an infinitely certain direction. Every constrained adjustment does this at every held station, and a plot of error ellipses that includes them is showing three real ellipses and two artefacts of the constraint list.

One caution goes with the arithmetic. Both reliability figures are computed at the redundancy numbers of the adjusted network, which are a property of the geometry and the weights rather than of the observations, so they are available before any field work is done — which is what makes them a design tool rather than a report. What they cannot say is whether more than one blunder is present, because the identity is linear in a single ∇ and two simultaneous errors interact through the same normal matrix.

What was computed, and how

The observations are generated from stated coordinates with stated noise from a stated seed, and every caption says so. This is not squeamishness: what is being measured here is the adjustment, and for that the observations have to be ones whose errors are known, which no field measurement ever is. A real network, whose coordinates would have to be reduced to the grid before any of this, would let none of the identities above be checked, because the true coordinates would be unavailable and the blunder would be unlabelled.

The solve is Cholesky on a dense 8 × 8, which is small enough to print and large enough that the redundancy numbers are not all equal. The bearing constraint is applied as an observation of the across-track displacement with a weight of 10¹², which is what a datum constraint is: an observation the network is not allowed to argue with.

The residual convention is the surveying one — a positive residual means the adjustment lengthened the observation — and it is checked rather than assumed: the adjusted distance between the converged coordinates less the observed distance agrees with the reported residual to 7 × 10⁻¹³ m.

Where the model stops

Distances only. A real network mixes distances, directions and angles, whose design rows have different units and whose relative weighting is a decision with consequences. Nothing here has to choose between a millimetre and an arcsecond, which is the hardest weighting question in practice and is entirely absent.

Two dimensions. A three-dimensional network has a datum defect of six, or seven if the scale is not observed, and holding it minimally is correspondingly more awkward. The identities all survive; the arithmetic is longer.

And the noise is Gaussian by construction. The variance factor tests that hypothesis and here it passes because the hypothesis is true. On real work it is the one assumption most likely to be false, and a variance factor that fails can mean bad observations, an optimistic instrument specification, an over-constraint, or a model that is missing a systematic term — four diagnoses from one number.

The generalisation

A fitted quantity is not a measurement, and the difference is a gauge freedom. Everything the observations determine is invariant under the constraint choice; everything they do not determine is decided by it, and the two must be kept apart in the report or the second will be read as the first.

The pattern is everywhere a model has more parameters than the data constrains. A regression through data with a collinear pair has a direction in parameter space the data says nothing about, and the reported coefficients depend on the regularisation rather than on the world. A phylogeny has an unrooted topology the data supports and a root it does not. A structural analysis has rigid-body modes.

What surveying contributes is the test — residuals that must be invariant, and are — and the vocabulary for the two runs. It is a discipline that had to build the distinction into its procedure because the consequences were legal, and it is worth borrowing wherever a fit is reported.

Who found it, and when

Least squares arrived in geodesy before it arrived anywhere else and was invented for it. Legendre published the method in 1805 in an appendix on determining the orbits of comets; Gauss claimed to have used it since 1795 and gave it its probabilistic justification in 1809, having applied it to the Ceres orbit. The application that occupied him for the rest of his life was the triangulation of the Kingdom of Hanover, which is where the normal equations, the Gaussian elimination that solves them, and the least-squares criterion all met the problem they were built for.

Redundancy numbers are much younger. Willem Baarda’s 1968 monograph on testing in geodetic networks introduced the reliability theory that gives them their meaning — internal reliability as the minimal detectable blunder, external reliability as the effect of an undetected one on the coordinates — and the data snooping procedure that reads standardised residuals against a threshold. It is one of the few places in applied statistics where the question “what could this data set fail to notice” was answered before the data was collected rather than after.

The insistence on running a free adjustment first is procedural rather than theoretical and is written into national survey specifications rather than into papers, which is why it is often the part a new practitioner meets last.

Where the ladder goes next

This rung takes the coordinate as far as the solve. What it has not asked is what the published coordinate does afterwards — how it is transformed into other frames, and what the transformation’s own parameters, which are themselves the output of a fit like this one, do to the width of the answer. That question belongs to the datum ladder, which has been printing seven parameters as exact numbers for nine essays.

What this makes readable

Essays that name this one as a prerequisite.

Named alongside this one

Essays reaching for the same objects. Nobody chose these; they are what the concept index makes visible.

What links here

The 8 essays that link to this one and share the most of its objects, of 13 that link here.

The objects this essay names

Each one links to every other essay that touches it.

AdjustmentBlunderConstraintCovarianceDatumDegrees of freedomError ellipseLeast-squaresNetworkRedundancyResidualSurvey