The average of noisy positions moves
The first two rungs are about one position: how wide it is written, and what shape its uncertainty is drawn as. Both are questions about a single number and its interval.
Something different happens when several noisy positions are combined. A mean of many observations is the one calculation everybody trusts, because averaging is what makes noise go away — and on a map it acquires an error that averaging does not touch.
Why a mean does not commute with a map
For a linear function, the average of the images is the image of the average. That is the property averaging relies on and it is the property a projection does not have.
Write the projection near a point as a Taylor expansion in a ground displacement x measured east and north:
f(x) = f(0) + Ax + ½ xᵀHx + …
Take the expectation. The mean of x is zero — the observations are unbiased on the ground, which is the assumption doing all the work. The linear term therefore vanishes, and what survives is
E[f(x)] − f(0) = ½ tr(HΣ)
where Σ is the covariance. The displacement depends on the second derivative of the projection and on the width of the scatter, and on nothing else. It contains no n. Averaging more observations does not reduce it, because it was never an average of anything that could cancel.
This is the second derivative the flexion ladder is about, met in a place with no geodesics in it. Tissot’s indicatrix is the A in that expansion; flexion and skewness are the H; and the displacement of a mean is the H contracted with an instrument’s covariance instead of with a direction.
The two errors, drawn together
The crossover is the number worth carrying away. With a sixty-kilometre scatter at 55° north on a Mercator map it is 44,066 observations. Below that the map’s contribution is buried in the noise and nobody would ever see it; above it, the reported precision of the answer is a statement about the wrong quantity entirely, and the tighter it gets the more confidently wrong it is.
Datasets of that size are ordinary. A national record of lightning strikes, a decade of species observations, a year of vehicle telemetry, an archive of geotagged photographs — all of them are hundreds of thousands to hundreds of millions of points, all of them are averaged to a centre, and the averaging is done in whatever coordinate system the file happened to arrive in.
Which projections do it, and by how much
Three things in that figure are worth separating.
The plate carrée does not move the answer at all. Its measured displacement is two metres against Mercator’s four hundred and four, and two metres is this measurement’s own noise: the exact value is zero. For a normal cylindrical projection with y = f(φ) the ground displacement works out to ½ (σ²/R) f″/f′, and the plate carrée’s f is φ, whose second derivative is zero. The projection nobody chooses is the only one in the library that averages correctly.
That exemption is fragile. It holds only while the two components of the error are uncorrelated: give the same scatter a correlation of 0.7 and the plate carrée acquires a displacement of 282 m, due east, because the cross-derivative it does have is now being contracted with something that is not zero.
Mercator and the equal-area cylindrical are exact opposites. f″/f′ is tan φ for Mercator, whose f is the isometric latitude, and −tan φ for the Lambert cylindrical, whose f is sin φ. So the displacement is ½ (σ²/R) tan φ and its negative: 403.494 m and −403.494 m at 55°, agreeing with the closed form to a part in ten thousand and with each other to a part in ten million.
The two projections at the centre of the loudest argument in the subject displace an average by the same distance in opposite directions. Mercator pulls it towards the pole and the equal-area cylindrical pulls it towards the equator, and neither of the properties they are argued about — angles, areas — has anything to do with it.
The gnomonic is the worst by a factor of three, at 1,269 m, with its crossover at 4,472 observations. That is the projection whose whole purpose is that every great circle is straight, bought at a second derivative larger than anything else in the library.
And the projection this actually happens on is Web Mercator. Almost every collection of geotagged records is stored, tiled, clustered and averaged in the pyramid the web is built on, whose formulae are Mercator’s applied to spherical latitudes. Its displacement is 403.5 m at 55° north, identical to Mercator’s to every figure computed here, because on a sphere the two are the same map. Clustering a million points into a marker, computing the mean of a cluster’s members, snapping a label to a centroid — each of those is this arithmetic, done in tile coordinates, at a latitude the software never looks at.
The shape of the curve deserves a sentence of its own. tan φ is not a gentle function: the displacement is 283 m at 45°, 489 m at 60°, 821 m at 71° and passes a kilometre just past 74°. A dataset covering a mid-latitude country is somewhere on the steep part of it, and the displacement is different at the two ends of the country — so the error is not even a constant offset that could be subtracted once. It is a field.
It is quadratic in the scatter, and the model has an end
A second-order term must grow as the square of the width of what it acts on, and that is the test that separates it from every other candidate explanation.
The departure at the wide end is not a failure of the measurement, which is exact at every width — it is a Monte Carlo average of the projection itself. It is the failure of the prediction, and it says where the local model stops: a quadratic is a statement about a neighbourhood, and 640 km is not a neighbourhood.
The same defect, arriving from three directions
This site has now met this displacement three times without recognising it.
A centroid belongs to a plane computes the mean of a polygon’s vertices in a projected plane and finds it is not the mean on the sphere. That is this calculation with a covariance built from a set of vertices instead of from an instrument.
The indicatrix is a limit measures the image of a finite circle against the ellipse that claims to describe it, and reports a centroid offset — the image of a circle is not merely the wrong shape, it is in the wrong place. That offset is ½ tr(HΣ) with Σ the covariance of a uniform disc.
And reprojecting a raster invents values averages cell values over ground that the projection has stretched unequally, which is the same contraction with the same H.
Three separate essays, three separate machineries, one term. Naming it is the whole content of this rung: a projection’s second derivative displaces every average taken on the page, and the displacement is the Hessian contracted with the second moment of whatever is being averaged.
What was computed, and how
The measured displacement is a Monte Carlo expectation with antithetic pairs: every draw x is used together with −x, so the odd terms of the expansion cancel identically and the estimator’s variance comes only from the quadratic term. That is what makes a number good to a few centimetres out of a scatter sixty kilometres wide, from twenty thousand draws rather than the billions a plain sample would need.
The predicted displacement is ½ tr(HΣ), with H obtained by second differences of the forward map along the local east and north directions at a step of two kilometres, and the result carried back to the ground through the same inverse linear map the previous rung uses.
The two agree to 0.14 per cent at σ = 10 km. They are meant to: one is a simulation of the projection and the other is a two-term description of it, and the point of computing both is that the second can be checked rather than trusted.
The one place they violently disagree is a finding. For the Robinson projection the prediction is 26,961 m against a measurement of 718 m — a factor of thirty-seven. Robinson is defined by a table of parallel spacings with an interpolant through it, so its second derivative is a property of the interpolation rather than of the projection, and a finite difference of it returns the interpolant’s own kinks. One projection in the library has no second derivative at all, and this is the third time that has produced a wrong number in a different calculation. The Monte Carlo measurement is unaffected, because it never differentiates anything.
The assertions have to refuse as well as confirm. They require the spread of the mean to fall by a factor of at least 32 between n = 1 and n = 4,096 — the control that says the two quantities are different kinds of number — the displacement to be identical at every count, the measured value to sit within five per cent of the closed form, and the displacement to quadruple when σ doubles.
What to do instead, which is short
The correction is not the interesting part of this rung and it is worth giving anyway, because it is one line.
Average on the sphere. Convert each position to a unit vector, take the mean of the vectors, normalise it, and convert back. That is exact for any projection, costs three multiplications per point more than averaging the projected coordinates, and has no second-order term at all — the vector mean’s bias is a fact about the sphere’s own curvature and is smaller than this one by the ratio of the scatter to the Earth’s radius squared.
Where that is impossible — because the software averages in whatever plane it is handed, which is the usual case — the displacement is computable in advance from the projection, the latitude and the scatter’s covariance, and can be subtracted. The figures above are that calculation, and its inputs are three numbers anybody averaging positions already has.
Where the model stops
The observations are assumed unbiased on the ground. If they are not, everything here is added to whatever bias they already had, and no measurement on the map can separate the two.
The covariance is assumed the same at every observation. A real population’s scatter varies from place to place, and the displacement is then an average of ½ tr(HΣ) over the population’s own distribution — which is a heavier calculation and not a different phenomenon.
And the whole treatment is second order. Beyond about a hundred kilometres of scatter at these latitudes the fourth-order terms are measurable, as the quadratic figure shows, and beyond a few hundred the expansion is not the right tool: the honest calculation there is the Monte Carlo one, or an average taken on the sphere in the first place.
The generalisation
Averaging is a linear operation, and it commutes only with linear maps. Every field that averages transformed data meets this, and the correction is always the same shape — half the trace of the Hessian against the covariance. It is Jensen’s inequality with a sign attached, it is the Itô correction in stochastic calculus, and it is the reason a log-transformed mean is not the mean of the logs.
The cartographic case is unusual only in that the map is chosen for reasons that have nothing to do with the average — for looking right, or for being the format the data arrived in — and that the same map is used for both the picture and the arithmetic. A projection selected because countries look the right size is deciding, silently, which way a continent’s worth of averaged observations drifts.
Who found it, and when
Jensen published his inequality in 1906, and the second-order correction to a transformed mean is standard in every treatment of propagation of error since Gauss. The specific consequence for coordinates was noticed by geodesists computing network adjustments in projected coordinates, which is why textbooks on adjustment say — without much explanation — that the adjustment should be done in the ellipsoidal or Cartesian frame and only then projected.
That instruction is usually read as fastidiousness about rounding. It is not: it is this term, and the reason it was possible to ignore for a century is that surveying networks span kilometres, where a scatter of a few centimetres makes ½ tr(HΣ) a quantity of nanometres. The term became visible when the data changed — when the things being averaged stopped being control points and started being millions of records spread across a continent.
More data makes it worse, not better
There is one consequence of the term’s structure that inverts the usual expectation, and it is the reason the defect became visible when it did.
The bias depends on the population’s spread, not on how many points were averaged. The correction is , and there is the scatter of the points being combined — a property of how far apart they are on the ground. Averaging ten thousand records rather than a hundred does nothing to it, because the ten thousand are spread over the same continent.
What more data does shrink is the random part, as , in the ordinary way. So the two terms move in opposite directions with sample size: the noise falls away and the systematic displacement sits exactly where it was.
The consequence is that a large dataset is the dangerous case. With a hundred records the bias is buried in the scatter of the estimate and nobody could see it if they looked. With ten million the estimate is reproducible to a metre, it is stable across subsamples, it survives every bootstrap — and it is in the wrong place, by an amount that no amount of further collection will reduce.
That is the classical bias-and-variance distinction arriving in coordinates, and it explains the historical timing the essay’s own account gives. A century of surveyors averaging control points never met the term, not because their instruments were crude but because their scatter was kilometres and their was small enough that nothing looked suspiciously precise. A modern pipeline averaging a continent’s worth of records has a scatter of thousands of kilometres and an large enough that the answer looks settled.
It also explains why the defect resists being found by comparing two pipelines. Two teams working in the same plane on the same records get the same displaced answer, agree to the metre, and conclude that the number is settled — because they share the term rather than differing in it.
And every reassurance available to the analyst points the wrong way. Reproducibility, tight confidence intervals, agreement between subsamples, stability under resampling — all of them measure the term that is shrinking, and none of them touches the one that is not. The only thing that detects it is doing the arithmetic in a frame where it does not arise, which is what the short recommendation above says and is why it is short.
The term is shared rather than disputed, which is the hardest kind of error to find.
Where the ladder goes next
The three rungs so far take a coordinate’s width from the way it is written, from the instrument that produced it, and from the population it belongs to. What none of them asks is the question underneath: whether the transformation between two coordinate systems is itself known well enough for the difference between two candidates to be worth arguing about. That is a covariance on the parameters rather than on the position, and it is where this ladder goes.
Named alongside this one
Essays reaching for the same objects. Nobody chose these; they are what the concept index makes visible.
- A local model has an order flexion · numerical differentiation · quadratic law · second-order
- The size at which the second derivative arrives flexion · numerical differentiation · second-order · taylor expansion
- A dot map's density is partly the projection's bias · equal-area · thematic mapping
- The area is unbiased and the perimeter is not bias · precision · quadratic law
- The class breaks were computed on the page bias · equal-area · thematic mapping
- The sample was drawn on the page bias · covariance · equal-area
What links here
Every essay whose body links to this one.
The objects this essay names
Each one links to every other essay that touches it.
BiasCentroidCovarianceEqual-areaFlexionNumerical differentiationPlate carréePrecisionQuadratic lawSecond-orderTaylor expansionThematic mapping