The error that does not average down
Assumes The error ellipse is not an ellipse.
Every rung of this ladder gives a coordinate a width and pushes it somewhere: through a projection, into a mean, into a difference, past the first order. All five share an assumption that none of them states: that the observations a coordinate was averaged from are independent of one another.
They are not, and the consequence is not a correction. It is a different exponent.
Where the square root comes from
The standard error of the mean of N observations is σ/√N, and the derivation is one line: the variance of a sum of independent variables is the sum of their variances, so the variance of their mean is σ²/N.
The word doing all the work is independent. A position time series from any continuously observing instrument is not: it has power at every frequency, with more of it at the low ones, and a mean over N observations removes the high frequencies and leaves the low ones exactly where they were.
The spectrum, stated rather than fitted
The noise here is a power law, S(f) ∝ f^−α. α = 0 is white noise; α = 1 is flicker, which is what the geodetic literature reports for position time series; α = 2 is a random walk.
Stating the exponent rather than fitting one to real observations is this collection’s usual decision. A real series carries an instrument, a site, a processing chain and a filter, and a figure computed from one is partly a measurement of somebody’s software — the coastline decision, applied to a third kind of input.
That is the trap in one picture. Three series with identical standard deviations and completely different value per hour of observing, and the quantity every instrument’s specification quotes is the standard deviation.
The measurement
Cut the series into blocks of each length, take the mean of each block, and measure the scatter of those means. That is the empirical version of the quantity, and it makes no assumption at all about what the observations are doing — which is the point, since the assumption is what is on trial. That is the standard error the length actually buys, without any assumption about independence.
| spectrum | exponent | R² |
|---|---|---|
| white, α = 0 | −0.484 | 0.999 |
| α = 0.25 | −0.371 | 0.999 |
| α = 0.5 | −0.267 | 0.999 |
| α = 0.75 | −0.178 | 0.993 |
| flicker, α = 1 | −0.111 | 0.966 |
| α = 1.5 | −0.040 | 0.822 |
| random walk, α = 2 | −0.015 | 0.649 |
The prediction is (α − 1)/2 and the measurement follows it closely up to about α = 0.5 and then sits above it. That departure is honest and it is in the optimistic direction: a finite series cannot carry the lowest frequencies the spectrum asks for, so the simulation understates the correlation and every exponent here is closer to a half than the truth is.
What that costs a schedule
The exponent is not an academic quantity. It is the conversion between observing time and accuracy, and it is what a survey specification is written against.
A hundred observations against a thousand million. At one observation a day that is four months against two and a half million years — and the second figure is not a rhetorical flourish, it is the arithmetic of the fitted exponent, which is itself an optimistic reading of the spectrum.
The shape of the failure is worth naming. Under white noise, halving the error costs four times the observing. Under flicker it costs the observing raised to the power of nine, so every halving is more expensive than the last by a factor of five hundred. There is no schedule that reaches a target that way; there is only a point past which the schedule stops being the variable.
The right reading of that is not that a factor of ten is unobtainable — it is that it cannot be bought by waiting. What buys it is a second instrument, a second site, a different technique: anything that is independent of the first, because independence is the thing the square root needed and the thing a longer average does not supply.
The crossover this moves
The previous rungs of this ladder are built around a comparison: a bias that does not fall with the sample count, against a noise that does. Where the two cross is the sample size at which it stops being worth observing more.
Under a correlated spectrum the falling curve barely falls. At an exponent of −0.111 the scatter after ten thousand observations is 0.36 of a single observation’s, where the square root would give 0.01 — so the crossover arrives at a sample size thirty-six times smaller in that quantity, and in practice arrives almost immediately.
That is a correction to a published result of this collection rather than a new claim, and it goes in the direction that makes the earlier rung’s point stronger: the map’s own bias overtakes the instrument’s noise sooner than it said.
Where the distinction stops being useful
The uncomfortable version of the same statement is that under a spectrum with enough memory, every part of the error stops averaging — and the distinction between a bias and a noise stops meaning anything.
A random walk’s mean over a finite window is a random variable whose scatter does not shrink with the window at all: the measured exponent is −0.015, which is zero to the resolution of the fit. So an observer who runs their instrument twice as long gets a second answer that differs from the first by as much as it ever would, and calling the difference random is true and useless.
Which term is systematic and which is random is a statement about the observing window rather than about the physics. A drift that looks like a bias over an afternoon looks like noise over a decade, and the same series supplies both descriptions. That is the reason a specification cannot be written as one number and the reason an exponent belongs beside it.
The practical form of this is old and is usually stated as advice rather than as arithmetic: do not average past the point where the noise stops being white. The measurement above says where that point is and what it costs to ignore it.
Why the specification does not say so
An instrument’s accuracy is quoted as a standard deviation, and three series with wildly different value per hour have the same one. Nothing in the number records the arrangement.
This is the same failure the collection meets in the tolerance that decides the verdict and in a coordinate is a number with a width: a single summary statistic is being asked a question it cannot answer, and the answer it gives instead is plausible. A scatter is a second moment; correlation lives in the second moment between observations, and no report of the first tells anybody anything about it.
The honest quantity is the exponent, and reporting it costs an extra number in a table.
Where this bites in the rest of the collection
Three places, and the third is uncomfortable.
A datum’s realisation. Every national frame is a set of coordinates averaged over years of observation, and what a fit leaves as residuals is judged against the precision that averaging is supposed to have bought. A published coordinate is a result of a solve over many epochs of observation, and the formal precision that solve reports assumes the epochs are independent. At α = 1 the reported precision of a coordinate averaged over a year is too good by a factor of about six.
The clock a network runs on. A grid stops fitting the ground it was laid on computes the years until a tolerance is broken from a strain rate treated as exact. The strain rate is estimated from a position series, and a rate estimated from correlated noise has an uncertainty that the usual formula understates by the same factor.
And this ladder’s own earlier rungs. The average of noisy positions moves measures a bias that does not fall with the sample count, and contrasts it with a noise that does. Under a correlated spectrum the noise very nearly does not fall either, so the crossover that rung computes — where the map’s bias overtakes the scatter — arrives much earlier than it says.
What the check refuses
A measurement of an exponent is worth exactly as much as its controls, and this one has two.
White noise has to come back at a half. The same series generator, the same block-mean estimator, the same fit, on an uncorrelated series, returns −0.484 at a coefficient of determination of 0.999. If it did not, everything below would be a measurement of the estimator rather than of the correlation. The 3 per cent shortfall from a clean −0.500 is the finite length of the series, and it shrinks as the series grows.
And the ordering has to hold. A random walk cannot average down faster than flicker, whatever the arithmetic does, because it has strictly more power at every low frequency. The measurement returns −0.015 against −0.111, in the right order, at every seed tried.
Between them those two say the axis is real. What they do not say is where any particular instrument sits on it, and that is a measurement nobody can make from a specification sheet.
What the exponent does to a specification
A specification is written as a precision and a method, and the exponent is what connects them — so a specification that omits it is not enforceable.
Occupy each mark for four hours is a method. Five millimetres horizontal is a precision. Whether the first delivers the second depends entirely on how the error falls with the occupation, and that is the exponent: under white noise four hours is four times better than fifteen minutes, and under flicker it is 1.3 times better.
The consequence is that two surveyors following the same specification with different instruments — or the same instrument at different sites, since the noise spectrum is site-dependent — deliver different accuracies, and neither has departed from the specification. What varies is a number nobody wrote down.
The repair is to specify the exponent rather than the hours, or, more practically, to specify the result and let the observer measure their own series. That is what a modern processing chain does when it reports a realistic error rather than a formal one, and the difference between the two is exactly the factor this rung is about.
There is a diagnostic in this that costs nothing. Split an existing series in half and average each half; the ratio of the two half-averages’ scatter to the whole’s is 2^α rather than √2 whenever the noise is not white, so a single existing dataset reports its own exponent with no extra observation at all. The number nobody writes down is one already in hand.
Where the model stops
The series here are 4,096 observations long and the block lengths run to 256, so the fit sees three decades of the spectrum and not the four or five a decade of daily positions would. The exponents at the correlated end are the ones this limits, and it limits them in the direction of a half.
The spectrum is a pure power law. A real series is a sum of white and flicker with a station-dependent ratio, and its exponent is somewhere between the two ends of the axis above — closer to white at short averaging times and to flicker at long ones, which is a curve rather than a line and which this measurement cannot produce.
And nothing here is a time series in any physical sense: the observations are a sequence with a stated correlation structure and no clock attached. What that omits is everything periodic — a daily cycle, a seasonal one, the draconitic period every satellite constellation puts into a position series — and a periodic term does not average down at any rate at all until a whole number of cycles have gone by.
The generalisation
The square root in σ/√N is a property of the observations and not of the estimator, and it is the assumption in the derivation that nobody checks.
Stated that way it is not about geodesy. It is the same statement that makes an effective sample size smaller than a sample size, that makes an autocorrelated Monte Carlo chain worth less than its length, and that makes a time-averaged measurement of anything with long-memory noise disappointing. The quantity everybody wants is independent observations, and the quantity everybody has is observations.
What makes the geodetic case sharp is the size of the gap. A factor of a thousand million between two exponents on the same axis is not a correction to a schedule; it is the difference between a plan that works and one that cannot.
Who found it, and when
Flicker noise in geodetic position series was established in the 1990s, by analyses of the early continuous GPS networks that found the white-noise assumption gave rate uncertainties several times too small. The consequence — that formal errors from a least-squares solve over correlated observations are optimistic by a factor that depends on the spectrum — is now standard practice in the field and is routinely ignored outside it.
The mathematics is older and belongs to the study of long-memory processes: that a process with a spectral density diverging at zero frequency has a sample mean whose variance falls more slowly than 1/N is a classical result, and the exponent (α − 1)/2 is its statement.
Telling which regime a series is in, without a spectrum
The distinction is worth acting on and the usual route to it — fitting a spectral index to a periodogram — is enough work that most people will not do it. There is a cheaper test that answers the question the schedule actually asks.
Split the series and compare the halves. Take the estimate from the first half and from the second half, and look at how far apart they are. Under white noise, each half-mean has times the scatter of the whole-series mean, and the pair should disagree by about that much. Under a flicker or random-walk spectrum they disagree by much more, because the low-frequency content that separates the two halves has not been averaged away by either.
Repeat it at several splits. Quarters against quarters, eighths against eighths: under white noise the disagreement grows as with the number of pieces, and under a correlated spectrum it barely grows at all. The slope of that curve is the exponent, obtained from means rather than from a transform, and it needs no window, no detrending choice and no taper.
That is the Allan variance in its plainest form, which is how the frequency-standards community answered the same question for clocks decades ago, and it transfers unchanged because the question is identical: how does the uncertainty of an average fall as the average gets longer.
The verdict is one-sided in the useful direction. A series whose half-means disagree far more than white noise predicts is proved to be correlated, and the formal error from the solve is proved optimistic. A series whose halves agree has not been shown to be white — the correlation may simply be at frequencies below what the record length can see, which is the standing hazard with any long-memory process and is why a longer record so often makes the uncertainty grow.
And it costs one pass over data already in hand. Which is the argument for it: the spectral analysis is better and it is not being done, while this can be done by anybody who can compute a mean.
Where the ladder goes next
This rung is about the spread of an estimate. The next is about its centre: a length computed from two noisy points is not merely uncertain, it is systematically too long, and the amount does not fall when more baselines are measured either.
What this makes readable
Essays that name this one as a prerequisite.
Named alongside this one
Essays reaching for the same objects. Nobody chose these; they are what the concept index makes visible.
- A length measured from noisy points is too long error budget · precision · standard error · survey · verification
- How many triangles it takes averaging · error budget · standard error · survey · verification
- The parameters are not independent correlation · error budget · precision
- A long window and a square one error budget · verification
- How big a triangle it takes precision · verification
- The answer is a set precision · verification
What links here
Every essay whose body links to this one.
The objects this essay names
Each one links to every other essay that touches it.
AveragingCoordinate widthCorrelationError budgetExponentFlicker noiseNoise spectrumPrecisionRandom walkStandard errorSurveyVerification