Measuring distortion

The error that does not average down

Five rungs give a coordinate a width, and all five assume the observations it was averaged from are independent. They are not. Under the noise spectrum every published analysis of a position time series reports, the standard error falls as the eleventh root rather than the square root — so a factor of ten costs a hundred observations if the noise is white and a thousand million if it is not.

Assumes The error ellipse is not an ellipse.

Every rung of this ladder gives a coordinate a width and pushes it somewhere: through a projection, into a mean, into a difference, past the first order. All five share an assumption that none of them states: that the observations a coordinate was averaged from are independent of one another.

They are not, and the consequence is not a correction. It is a different exponent.

Waiting longer buys almost nothing under the spectrum that is actually there. The scatter of block means against the block length, for the three spectra. White noise falls with a slope of -0.484 — the inverse square root everybody assumes. Flicker falls with a slope of -0.111, so a hundredfold longer average buys a factor of 1.67 rather than ten. A random walk is very nearly flat.
Fig. 1 The scatter of block means against the length of the block, for three stated noise spectra. White noise falls with a slope of −0.484, which is the inverse square root everybody assumes. Flicker noise falls with a slope of −0.111.

Where the square root comes from

The standard error of the mean of N observations is σ/√N, and the derivation is one line: the variance of a sum of independent variables is the sum of their variances, so the variance of their mean is σ²/N.

The word doing all the work is independent. A position time series from any continuously observing instrument is not: it has power at every frequency, with more of it at the low ones, and a mean over N observations removes the high frequencies and leaves the low ones exactly where they were.

The spectrum, stated rather than fitted

The noise here is a power law, S(f) ∝ f^−α. α = 0 is white noise; α = 1 is flicker, which is what the geodetic literature reports for position time series; α = 2 is a random walk.

Stating the exponent rather than fitting one to real observations is this collection’s usual decision. A real series carries an instrument, a site, a processing chain and a filter, and a figure computed from one is partly a measurement of somebody’s software — the coastline decision, applied to a third kind of input.

Three noise spectra, at the same variance. Five hundred and twelve observations of a coordinate under each of three stated spectra, all scaled to the same standard deviation. The top trace is white noise, where every observation is independent; the middle is flicker, which every published analysis of a geodetic position series reports; the bottom is a random walk. Any one of them would pass the same test for its scatter, and they buy a survey entirely different amounts of accuracy per hour.
Fig. 2 Five hundred and twelve observations under each of the three spectra, all scaled to the same standard deviation. Any one of them would pass the same test for its scatter. The difference between them is entirely in how the values are arranged in time, which no summary of the scatter can see.

That is the trap in one picture. Three series with identical standard deviations and completely different value per hour of observing, and the quantity every instrument’s specification quotes is the standard deviation.

The measurement

Cut the series into blocks of each length, take the mean of each block, and measure the scatter of those means. That is the empirical version of the quantity, and it makes no assumption at all about what the observations are doing — which is the point, since the assumption is what is on trial. That is the standard error the length actually buys, without any assumption about independence.

spectrum exponent
white, α = 0 −0.484 0.999
α = 0.25 −0.371 0.999
α = 0.5 −0.267 0.999
α = 0.75 −0.178 0.993
flicker, α = 1 −0.111 0.966
α = 1.5 −0.040 0.822
random walk, α = 2 −0.015 0.649
The half everybody assumes is the value at one end of an axis. The fitted exponent against the spectral index, with the leading-order prediction (α − 1)/2 drawn through it. At α = 0 the exponent is -0.484, which is the inverse square root; by α = 1 it is -0.111. The measurement sits above the prediction at the correlated end because a finite series cannot carry the lowest frequencies the spectrum asks for, which makes every figure here an optimistic one.
Fig. 3 The fitted exponent against the spectral index, with the leading-order prediction (α − 1)/2 drawn through it. The half everybody assumes is the value at one end of an axis, and the axis runs to zero.

The prediction is (α − 1)/2 and the measurement follows it closely up to about α = 0.5 and then sits above it. That departure is honest and it is in the optimistic direction: a finite series cannot carry the lowest frequencies the spectrum asks for, so the simulation understates the correlation and every exponent here is closer to a half than the truth is.

What that costs a schedule

The exponent is not an academic quantity. It is the conversion between observing time and accuracy, and it is what a survey specification is written against.

How many observations a factor of ten costs. The number of observations each spectrum needs before the standard error of their mean is a tenth of a single observation's, on a logarithmic scale. Under white noise it is a hundred, which is a morning. Under flicker it is 1.0e+9, which at one observation a day is 2.8e+6 years. The bars are logarithms because the column cannot be drawn any other way.
Fig. 4 The number of observations each spectrum needs before the mean’s standard error is a tenth of a single observation’s, on a logarithmic scale. Under white noise it is a hundred. Under flicker it is 1.0 × 10⁹.

A hundred observations against a thousand million. At one observation a day that is four months against two and a half million years — and the second figure is not a rhetorical flourish, it is the arithmetic of the fitted exponent, which is itself an optimistic reading of the spectrum.

The shape of the failure is worth naming. Under white noise, halving the error costs four times the observing. Under flicker it costs the observing raised to the power of nine, so every halving is more expensive than the last by a factor of five hundred. There is no schedule that reaches a target that way; there is only a point past which the schedule stops being the variable.

The right reading of that is not that a factor of ten is unobtainable — it is that it cannot be bought by waiting. What buys it is a second instrument, a second site, a different technique: anything that is independent of the first, because independence is the thing the square root needed and the thing a longer average does not supply.

The crossover this moves

The previous rungs of this ladder are built around a comparison: a bias that does not fall with the sample count, against a noise that does. Where the two cross is the sample size at which it stops being worth observing more.

The error that averages away, and the one that does not. Falling line: the standard error of the mean of n observations, which is σ over the root of n and is what everybody expects averaging to do. Flat line: the displacement the projection itself puts into the answer, which is 404.2 m at every n because the expectation of a sample mean does not depend on the sample size. They cross at n = 44,066, and past that point the map is the larger of the two errors — the one place where collecting more data makes the answer worse rather than better, since it removes the noise that was hiding the bias.
Fig. 5 The average of noisy positions moves draws the comparison: one term that averages away and one that does not, against the number of observations. Everything in it rests on the falling curve falling as the inverse square root, which is the assumption this rung is about.

Under a correlated spectrum the falling curve barely falls. At an exponent of −0.111 the scatter after ten thousand observations is 0.36 of a single observation’s, where the square root would give 0.01 — so the crossover arrives at a sample size thirty-six times smaller in that quantity, and in practice arrives almost immediately.

That is a correction to a published result of this collection rather than a new claim, and it goes in the direction that makes the earlier rung’s point stronger: the map’s own bias overtakes the instrument’s noise sooner than it said.

Where the distinction stops being useful

The uncomfortable version of the same statement is that under a spectrum with enough memory, every part of the error stops averaging — and the distinction between a bias and a noise stops meaning anything.

A random walk’s mean over a finite window is a random variable whose scatter does not shrink with the window at all: the measured exponent is −0.015, which is zero to the resolution of the fit. So an observer who runs their instrument twice as long gets a second answer that differs from the first by as much as it ever would, and calling the difference random is true and useless.

Which term is systematic and which is random is a statement about the observing window rather than about the physics. A drift that looks like a bias over an afternoon looks like noise over a decade, and the same series supplies both descriptions. That is the reason a specification cannot be written as one number and the reason an exponent belongs beside it.

The practical form of this is old and is usually stated as advice rather than as arithmetic: do not average past the point where the noise stops being white. The measurement above says where that point is and what it costs to ignore it.

The excess spread is second order in the accuracy, which is why it hides. How much longer the propagated cloud is than the first-order ellipse, against the accuracy, on log axes. The fitted slopes are 1.96 for Mercator, 2.01 for Lambert cylindrical, 2.01 for Plate carrée — a square law, as the second-order term requires. A quantity that grows as the square of the input is one that stays invisible over the range anybody tests at and then arrives quickly, which is the same shape as every other second-order finding on this ladder.
Fig. 6 And the other half of the earlier rung: the excess spread that a projection’s second derivative adds to a scattered population, against the size of the scatter. It is second order in the accuracy, so it is invisible at survey scales and is not invisible at the scales a global average covers — and it is a term the square root never applied to in the first place.

Why the specification does not say so

An instrument’s accuracy is quoted as a standard deviation, and three series with wildly different value per hour have the same one. Nothing in the number records the arrangement.

This is the same failure the collection meets in the tolerance that decides the verdict and in a coordinate is a number with a width: a single summary statistic is being asked a question it cannot answer, and the answer it gives instead is plausible. A scatter is a second moment; correlation lives in the second moment between observations, and no report of the first tells anybody anything about it.

The honest quantity is the exponent, and reporting it costs an extra number in a table.

Where this bites in the rest of the collection

Three places, and the third is uncomfortable.

A datum’s realisation. Every national frame is a set of coordinates averaged over years of observation, and what a fit leaves as residuals is judged against the precision that averaging is supposed to have bought. A published coordinate is a result of a solve over many epochs of observation, and the formal precision that solve reports assumes the epochs are independent. At α = 1 the reported precision of a coordinate averaged over a year is too good by a factor of about six.

The clock a network runs on. A grid stops fitting the ground it was laid on computes the years until a tolerance is broken from a strain rate treated as exact. The strain rate is estimated from a position series, and a rate estimated from correlated noise has an uncertainty that the usual formula understates by the same factor.

And this ladder’s own earlier rungs. The average of noisy positions moves measures a bias that does not fall with the sample count, and contrasts it with a noise that does. Under a correlated spectrum the noise very nearly does not fall either, so the crossover that rung computes — where the map’s bias overtakes the scatter — arrives much earlier than it says.

What the check refuses

A measurement of an exponent is worth exactly as much as its controls, and this one has two.

White noise has to come back at a half. The same series generator, the same block-mean estimator, the same fit, on an uncorrelated series, returns −0.484 at a coefficient of determination of 0.999. If it did not, everything below would be a measurement of the estimator rather than of the correlation. The 3 per cent shortfall from a clean −0.500 is the finite length of the series, and it shrinks as the series grows.

And the ordering has to hold. A random walk cannot average down faster than flicker, whatever the arithmetic does, because it has strictly more power at every low frequency. The measurement returns −0.015 against −0.111, in the right order, at every seed tried.

Between them those two say the axis is real. What they do not say is where any particular instrument sits on it, and that is a measurement nobody can make from a specification sheet.

What the exponent does to a specification

A specification is written as a precision and a method, and the exponent is what connects them — so a specification that omits it is not enforceable.

Occupy each mark for four hours is a method. Five millimetres horizontal is a precision. Whether the first delivers the second depends entirely on how the error falls with the occupation, and that is the exponent: under white noise four hours is four times better than fifteen minutes, and under flicker it is 1.3 times better.

The consequence is that two surveyors following the same specification with different instruments — or the same instrument at different sites, since the noise spectrum is site-dependent — deliver different accuracies, and neither has departed from the specification. What varies is a number nobody wrote down.

The repair is to specify the exponent rather than the hours, or, more practically, to specify the result and let the observer measure their own series. That is what a modern processing chain does when it reports a realistic error rather than a formal one, and the difference between the two is exactly the factor this rung is about.

There is a diagnostic in this that costs nothing. Split an existing series in half and average each half; the ratio of the two half-averages’ scatter to the whole’s is 2^α rather than √2 whenever the noise is not white, so a single existing dataset reports its own exponent with no extra observation at all. The number nobody writes down is one already in hand.

Where the model stops

The series here are 4,096 observations long and the block lengths run to 256, so the fit sees three decades of the spectrum and not the four or five a decade of daily positions would. The exponents at the correlated end are the ones this limits, and it limits them in the direction of a half.

The spectrum is a pure power law. A real series is a sum of white and flicker with a station-dependent ratio, and its exponent is somewhere between the two ends of the axis above — closer to white at short averaging times and to flicker at long ones, which is a curve rather than a line and which this measurement cannot produce.

And nothing here is a time series in any physical sense: the observations are a sequence with a stated correlation structure and no clock attached. What that omits is everything periodic — a daily cycle, a seasonal one, the draconitic period every satellite constellation puts into a position series — and a periodic term does not average down at any rate at all until a whole number of cycles have gone by.

The generalisation

The square root in σ/√N is a property of the observations and not of the estimator, and it is the assumption in the derivation that nobody checks.

Stated that way it is not about geodesy. It is the same statement that makes an effective sample size smaller than a sample size, that makes an autocorrelated Monte Carlo chain worth less than its length, and that makes a time-averaged measurement of anything with long-memory noise disappointing. The quantity everybody wants is independent observations, and the quantity everybody has is observations.

What makes the geodetic case sharp is the size of the gap. A factor of a thousand million between two exponents on the same axis is not a correction to a schedule; it is the difference between a plan that works and one that cannot.

Who found it, and when

Flicker noise in geodetic position series was established in the 1990s, by analyses of the early continuous GPS networks that found the white-noise assumption gave rate uncertainties several times too small. The consequence — that formal errors from a least-squares solve over correlated observations are optimistic by a factor that depends on the spectrum — is now standard practice in the field and is routinely ignored outside it.

The mathematics is older and belongs to the study of long-memory processes: that a process with a spectral density diverging at zero frequency has a sample mean whose variance falls more slowly than 1/N is a classical result, and the exponent (α − 1)/2 is its statement.

Telling which regime a series is in, without a spectrum

The distinction is worth acting on and the usual route to it — fitting a spectral index to a periodogram — is enough work that most people will not do it. There is a cheaper test that answers the question the schedule actually asks.

Split the series and compare the halves. Take the estimate from the first half and from the second half, and look at how far apart they are. Under white noise, each half-mean has 2\sqrt2 times the scatter of the whole-series mean, and the pair should disagree by about that much. Under a flicker or random-walk spectrum they disagree by much more, because the low-frequency content that separates the two halves has not been averaged away by either.

Repeat it at several splits. Quarters against quarters, eighths against eighths: under white noise the disagreement grows as N\sqrt{N} with the number of pieces, and under a correlated spectrum it barely grows at all. The slope of that curve is the exponent, obtained from means rather than from a transform, and it needs no window, no detrending choice and no taper.

That is the Allan variance in its plainest form, which is how the frequency-standards community answered the same question for clocks decades ago, and it transfers unchanged because the question is identical: how does the uncertainty of an average fall as the average gets longer.

The verdict is one-sided in the useful direction. A series whose half-means disagree far more than white noise predicts is proved to be correlated, and the formal error from the solve is proved optimistic. A series whose halves agree has not been shown to be white — the correlation may simply be at frequencies below what the record length can see, which is the standing hazard with any long-memory process and is why a longer record so often makes the uncertainty grow.

And it costs one pass over data already in hand. Which is the argument for it: the spectral analysis is better and it is not being done, while this can be done by anybody who can compute a mean.

Where the ladder goes next

This rung is about the spread of an estimate. The next is about its centre: a length computed from two noisy points is not merely uncertain, it is systematically too long, and the amount does not fall when more baselines are measured either.

What this makes readable

Essays that name this one as a prerequisite.

Named alongside this one

Essays reaching for the same objects. Nobody chose these; they are what the concept index makes visible.

What links here

Every essay whose body links to this one.

The objects this essay names

Each one links to every other essay that touches it.

AveragingCoordinate widthCorrelationError budgetExponentFlicker noiseNoise spectrumPrecisionRandom walkStandard errorSurveyVerification