What is taught wrongly

A second tracing tells a drifting hand from a copied stretch

A hand tracing a coast drifts rather than jitters, and a drifting error is a smooth departure no projection explains — so the criterion that counted copied stretches correctly under independent noise counts four or five seams where two were made, and three or four on a coast with no copy at all. Undo the drift before fitting and every count comes back, for no copy, one and two, in every tracing. That needs the drift's correlation length from outside the coast, known to within a factor of three on the short side and much less on the long, and a second tracing of the same coast supplies it: the difference between two tracings holds nothing but the two errors, and read that way the count is right in 22 tracings of 24, and never too high.

Assumes A residual cannot count the pieces of a map.

A residual cannot count the pieces of a map counted the copied stretches in a compiled coast. It cut the coast exactly into one to seven pieces, each explained by a candidate projection placed by a similarity, and charged each piece for the numbers it spent. The Bayesian criterion’s charge, the logarithm of the number of observations, counted one copied stretch as two seams and two as four in every trial, until the digitising error reached a thousandth of the map. Its failure was the eased copy: a stretch blended into place is no projection of anything, and the criterion paid for it in pieces — three seams, then four, then six.

It said in passing what else would fail the same way. It had assumed the digitising error was independent from point to point, and a hand drawing a coast does not jitter; it drifts. Its error at one point is mostly its error at the last, and an error that wanders smoothly along the outline is, to a criterion that only knows projections and similarities, one more smooth departure no projection explains. Whether easing and drift could be told apart, it said, was the question the rigid-copy model could not ask.

Before easing and drift can be told apart, drift has to be handled on its own — and it turns out it can be, completely, by information the coast itself does not contain.

With the drift's length known, the count comes back. A compiled coast of 360 points traced with a drifting error — each point's error mostly the last one's, over a correlation length of 18 points, three ten-thousandths of the map in size — and the number of seams BIC chooses once the error is whitened at an assumed length, eight tracings each. No copy (0 seams made): 0/8 right at none, 1/8 right at 2, 8/8 right at 6, 8/8 right at 12, 8/8 right at 18, 8/8 right at 30, 0/8 right at 60; one copy (2 seams made): 0/8 right at none, 6/8 right at 2, 8/8 right at 6, 8/8 right at 12, 8/8 right at 18, 7/8 right at 30, 0/8 right at 60; two copies (4 seams made): 1/8 right at none, 7/8 right at 2, 8/8 right at 6, 8/8 right at 12, 8/8 right at 18, 5/8 right at 30, 0/8 right at 60. Unwhitened the criterion buys pieces for the drift in every compilation; whitened at 6, 12, 18 points it counts every compilation right; assumed 3.3 times too long, it takes the most it is offered.
Fig. 1 A compiled coast of 360 points traced with a drifting error — each point’s error mostly the last one’s, over a correlation length of 18 points, three ten-thousandths of the map in size — and the number of seams the criterion chooses once the error is whitened at an assumed length, eight tracings each. Unwhitened it buys pieces for the drift whether the coast holds no copy, one or two; whitened at 6, 12 or 18 points it counts every compilation right; assumed 60 points long, it takes the most it is offered.

A drifting hand is a smooth departure

The coast is the earlier measurements’ Japan, an outline with no graticule, which a map with no graticule showed still names its projection. The drift is stated by two numbers. Its size is its standard deviation, here three ten-thousandths of the map’s width — three times the independent error the earlier measurement handled perfectly, and on a sheet sixty centimetres wide about a fifth of a millimetre. Its correlation length is how many points along the outline it takes for the error to forget itself: here eighteen points of 360, five per cent of the way round. Between those, the error is a first-order autoregression, each point’s error the last one’s times ρ=e−1/18=0.946\rho = e^{-1/18} = 0.946, plus a small fresh step.

A drifting hand draws departures no projection explains, a sixteenth the size of a copy's. The same coast two ways, as distance from the host sheet's own coast in thousandths of the map's width. Solid: a stretch from 55 to 85 per cent of the way round copied from a sinusoidal sheet, drawn exactly — a departure that rises from each seam to 13.6 thousandths. Shaded: the host's coast with nothing copied, traced with a drifting error of three ten-thousandths over a correlation length of 18 points — a departure that wanders smoothly up to 0.84 thousandths and back, and that no projection placed by a similarity explains either.
Fig. 2 The same coast two ways, as distance from the host sheet’s own coast in thousandths of the map’s width. Solid: a stretch from 55 to 85 per cent of the way round copied from a sinusoidal sheet, drawn exactly, rising from each seam to 13.6 thousandths. Shaded: the host’s coast with nothing copied, traced with the stated drift, wandering smoothly up to 0.84 thousandths and back.

The figure shows what that looks like against a copy. The copied stretch departs from the host sheet’s coast by up to 13.6 thousandths of the map, rising from each seam — the profile a compiled map agrees with its graticule except where it was copied found. The drift alone, on a coast with nothing copied, wanders up to 0.84 thousandths and back, a sixteenth of the copy’s size and of the same smooth kind. No candidate projection placed by a similarity explains a wander, just as none explained an eased copy, and the criterion has the same one response to both — the response when the answer is not in the library found to a whole map drawn in a projection nobody offered.

Left alone, it overcounts. On the coast with one copied stretch, eight tracings with this drift give four or five seams, once three, where two were made. On a coast with two copies, four to six where four were made. And on a coast with nothing copied at all, three to five, where none were. A reader following the criterion would report a compiled sheet with several sources, from a coast drawn in one projection by a hand that wandered by a fifth of a millimetre.

Undoing the drift before fitting

The earlier essay’s charge assumed each point was fresh evidence, and a drifting error makes that false: eighteen neighbouring points share most of one error between them. The standard remedy is to undo the correlation before fitting. If each point’s error is ρ\rho times the last one’s plus a fresh step, then the difference wt−ρ wt−1w_t - \rho\,w_{t-1} between each drawn point and ρ\rho times its predecessor carries only the fresh step — independent from point to point, of a known and smaller size.

The same operation applied to each candidate projection’s coast keeps the fit’s closed form. A similarity w≈az+bw \approx a z + b becomes wt−ρwt−1≈a (zt−ρzt−1)+b(1−ρ)w_t - \rho w_{t-1} \approx a\,(z_t - \rho z_{t-1}) + b(1 - \rho) — still a complex scale and a complex offset, so each stretch’s best residual is still a few subtractions of running totals, and the search over every segmentation is still exact. A stretch uses the pairs of neighbouring points inside it, and the one pair that straddles a seam belongs to neither piece, which costs one pair of 359 per seam.

With the drift’s correlation length and size known, the first figure’s middle columns are the result. Whitened at the true length of eighteen points and charged for the fresh step’s true size, the criterion counts no seams on the coast with no copy, two on the coast with one, four on the coast with two, in all eight tracings of each. Nothing else about the method changed. The overcounting was entirely the criterion being told that correlated evidence was independent.

How well the length has to be known

The first figure also answers how much the method depends on that knowledge. Whitened at a length assumed three times too short — six points — every count is still right. At two points, a ninth of the truth, the one-copy coast is counted right six times in eight and the no-copy coast once. At thirty points, two thirds too long, one tracing in eight of the one-copy coast gains a seam and three of the two-copy coast do. At sixty, every count on every coast jumps to six, the most offered.

The asymmetry is worth taking apart, because an over-long assumption changes two things at once. It whitens too hard, subtracting more of each predecessor than the error really shares. And it charges too little, because a longer drift of the same size would have smaller fresh steps.

Assumed too long, the drift fails by what it charges, not by how it whitens. The coast with one copied stretch and drift of correlation length 18 points, eight tracings. Rows: the length the error is whitened at. Columns: the length whose innovation size the criterion charges. Whitened at 18, charged at 18: 2, 2, 2, 2, 2, 2, 2, 2 seams; whitened at 18, charged at 60: 6, 6, 6, 6, 6, 6, 6, 6 seams; whitened at 60, charged at 18: 1, 1, 1, 1, 1, 2, 1, 1 seams; whitened at 60, charged at 60: 6, 6, 6, 6, 6, 6, 6, 6 seams. Charged for the innovation a sixty-point drift would have — about half the true one — the criterion overcounts whatever it whitens at; charged correctly, whitening too long costs it one of the copy's two seams instead.
Fig. 3 The coast with one copied stretch, eight tracings, whitened at the true length or at 60 points and charged for the fresh step a drift of either length would have. Whitened and charged at 18: two seams in all eight. Whitened at 18, charged at 60: six in all eight. Whitened at 60, charged at 18: one in seven tracings and two in one. Whitened and charged at 60: six in all eight.

Separating the two settles it. Whitened at the true length but charged for a sixty-point drift’s fresh step — about half the true one — the criterion takes six seams in every tracing. Whitened at sixty points but charged for the true step, it takes one seam in seven tracings of eight: it undercounts. So the jump to six is the charge. A criterion charged for less noise than there is buys pieces to explain the difference, whatever else is done; over-whitening on its own blunts the copy’s fainter seam, which is the safe direction to fail.

That makes the practical requirement one-sided. The size of the fresh step must not be underestimated. The length can be wrong by a factor of three if it errs short, and much less if it errs long, mostly because a long length implies a small step.

The count holds until the copy is buried

Unwhitened, drift of any size buys pieces; whitened, the count holds until the copy is buried. The coast with one copied stretch, traced with drift of a stated size over a correlation length of 18 points, eight tracings each. Hollow: the count with the drift's size known but its correlation ignored — median 5, 5, 5, 4, 4 seams. Filled: whitened at the true length — 2 in all eight; 2 in all eight; 2 in all eight; 1 in all eight; 1 0 1 0 0 1 0 0 at 0.3/10,000, 1/10,000, 3/10,000, 1/1000, 3/1000 of the map. The ignored correlation overcounts at every size, because the criterion's charge assumes each point is fresh evidence; whitened, the count is right until the drift reaches a thousandth of the map, and then the copy itself is lost in it.
Fig. 4 The coast with one copied stretch, traced with drift of a stated size over a correlation length of 18 points, eight tracings each. With the drift’s size known but its correlation ignored, the median count is 5, 5, 5, 4 and 4 seams at three hundred-thousandths, one, three and ten ten-thousandths, and three thousandths of the map. Whitened at the true length: two in all eight tracings at the three smallest sizes, one in all eight at a thousandth, and none or one at three thousandths.

The overcount does not depend on the drift being large. With its correlation ignored and its size known exactly, drift of three hundred-thousandths of the map — ten times smaller than the independent error the earlier essay counted perfectly through — still produces five seams where two were made. The criterion’s error is not about how much noise there is but about how many independent pieces of evidence it thinks it has, and a correlation of 0.946 from one point to the next divides that by about thirty-six at every size — (1+ρ)/(1−ρ)(1+\rho)/(1-\rho), twice the correlation length.

Whitened, the count is right at every size up to three ten-thousandths, then falls to one seam at a thousandth and to none or one at three thousandths. That is the copy being lost in the error rather than the error being mistaken for copies. A drift of a thousandth of the map, over eighteen points, carries departures as large as a fifth of the copy’s own, and the criterion, correctly, declines to pay for a seam it cannot see — the same undercount the earlier essay found for independent error at three thousandths.

The count comes back and the seams move

A count is one output of the method and not the only one. The earlier essay found three quantities with three tolerances to noise — the count most robust, the seams’ positions next, the pieces’ names most fragile — and whitening changes what each is made from.

Whitening brings the count back and loses some of where the seams are. The worst seam's distance from where it was made, as a share of the coast, for eight tracings with drift of length 18 points. One copy, whitened fit: median 4.2%, worst 13.6%; plain fit at that count: median 3.6%, worst 6.9%; two copies, whitened fit: median 4.9%, worst 13.3%; plain fit at that count: median 3.1%, worst 5.6%. The bar marks each row's median. The whitened fit finds a seam by the corner in the copy's departure alone; given its count, the plain fit also uses the departure's level and places the seams closer. The copied pieces are named for their true sources in 15 of the 16 tracings.
Fig. 5 The worst seam’s distance from where it was made, as a share of the coast, for eight tracings with drift of length 18 points. One copy, whitened fit: median 4.2 per cent, worst 13.6. The plain fit at that count: median 3.6, worst 6.9. Two copies, whitened fit: median 4.9, worst 13.3. The plain fit at that count: median 3.1, worst 5.6. The copied pieces are named for their true sources in 15 of the 16 tracings.

The whitened fit places its seams worse than one might expect from a fit that counts them perfectly. Its worst seam lies a median 4.2 per cent of the coast from where it was made on the one-copy coast, and 13.6 per cent in the worst tracing; on the two-copy coast, 4.9 and 13.3. The reason is what whitening keeps. Subtracting ho ho times each predecessor removes almost all of a slowly varying signal and leaves its steps, so on the whitened coast a copied stretch is visible mainly as a corner — the place its departure starts to grow. A corner softened by drift can be placed a dozen points either way at little cost.

The plain fit sees more. Forced to the count the whitened fit chose, and run on the tracing as drawn, it uses the departure’s whole level as well as its corners, and places the seams closer: a median 3.6 per cent of the coast on the one-copy coast and 3.1 on the two, with the worst tracing at 6.9 and 5.6. Its fault was never where it put the seams; it was how many it paid for. Given the right number, the drift that misled it about the count costs it little in position, because drift moves a seam’s evidence sideways by a few points and does not manufacture a corner where none was made.

So the two fits divide the work between them. The count comes from the whitened fit, and the seams from the plain fit at that count. The first is right because it knows what the evidence is worth; the second is better at locating what the first has decided is there, because it throws nothing away. A residual has more than one explanation is the general form of the lesson: which model is fitted and which is believed are two different decisions.

The names follow the earlier essay’s pattern. The copied stretches are named for their true sources — sinusoidal, and on the second coast sinusoidal and Mercator — in fifteen tracings of sixteen; the one miss names the sinusoidal copy a Mollweide. The host stretches are mostly named as whichever of the conformal conic, the equal-area conic and the polyconic the noise favours, which is the set the answer is a set found indistinguishable over a region the size of Japan. In two of the one-copy tracings the last host stretch, from the copy’s end round to the start of the walk, is named a sinusoidal or a Mercator instead. That stretch is the shortest host piece, about fifty points, and a wander of the hand over fifty points is enough to tilt a choice among candidates that differ by less than the drift there. The count does not notice, because a wrong name costs no seam; a reader using the names would be told the sheet had a third source it never had.

A second tracing says how the hand drifts

All of this has assumed the drift’s length and size were known, and the earlier attempt to estimate them from the coast itself had failed circularly. Estimated from the residuals of each candidate segmentation, the correlation depends on the segmentation — extra pieces soak up the drift and make the residuals look independent — so the criterion picks the segmentation whose residuals best justify it. The coast alone cannot say how much of a smooth departure is drift and how much is copying, because both are smooth departures.

Two tracings of the same coast can. Traced twice, by the same hand or the same process, a coast gives two outlines that share everything real — every projection, every copied stretch, every eased seam — and differ only by their two errors. The difference between them contains no projection and no copy. Its correlation from one point to the next is the drift’s ρ\rho, and its variance is twice the drift’s.

A second tracing of the same coast says how the hand drifts, and that is enough. Each compilation traced twice with independent drift of true length 18 points; the difference between the tracings holds only the two errors, and its lag-one correlation gives the length — 12.5, 20.5, 16.0, 20.1, 19.4, 10.9, 21.5, 11.6 points over eight pairs. The first tracing is then whitened at that length and charged at the size the difference implies. Seams chosen: no copy, 0, 0, 0, 0, 0, 0, 0, 0 (8 of 8 right); one copy, 2, 1, 2, 1, 2, 2, 2, 2 (6 of 8 right); two copies, 4, 4, 4, 4, 4, 4, 4, 4 (8 of 8 right).
Fig. 6 Each compilation traced twice with independent drift of true length 18 points. The lag-one correlation of the difference between the two tracings gives lengths of 12.5, 20.5, 16.0, 20.1, 19.4, 10.9, 21.5 and 11.6 points over eight pairs. Whitened and charged at what the difference says, the first tracing is counted right in all eight pairs for no copy and for two copies, and in six of eight for one, the other two counting one seam.

Read that way, eight pairs of tracings give correlation lengths between 10.9 and 21.5 points against the true 18. Whitening the first tracing of each pair at what its difference says, and charging for the step it implies, counts the coast with no copy right in all eight, the coast with two copies right in all eight, and the coast with one copy right in six. The two misses undercount, finding one seam of the two: an estimated length a little long, a step a little small, and the copy’s fainter end lost. Nowhere does the count go above the truth.

Twenty-two right of twenty-four, with no oracle anywhere and every error in the conservative direction, is the practical result. A compiled sheet that is to be counted should be traced twice.

What was checked

With nothing to whiten, the whitened count must be the plain one. At a correlation of zero, charged for the independent error’s known size, the whitened criterion is the earlier one charged the same way, and must choose the same count on the same tracing. It does.

The finding must be there to fail. Unwhitened, drift must be paid for in pieces: on each of the three compilations at least three tracings in four must be counted above the truth. They are — all eight, for each. Whitened at the true length every count must be right, and every one is.

An over-long assumption must fail by its charge. Whitened at sixty points but charged for the true step, the count must not exceed the truth in any tracing. It gives one seam or two.

Where the tracing is a model

A first-order autoregression. It is one error model among many, and — as with the datum and the projection’s parameters in the datum hides inside the projection’s parameters — a model of the error can trade against the pieces it is fitted beside. A real hand’s drift has more structure: it may correlate over more than one scale, depend on the curvature of the coast, or reset at the end of a stroke. A drift that is not first-order is not exactly undone by one subtraction, and the whitened residuals would keep some correlation; how much the count would suffer for it is not measured.

Two independent tracings. The twin estimate needs the two tracings’ errors to be independent of each other. The same hand tracing twice may repeat its own habits — always overshooting a headland — and a shared habit cancels in the difference and stays in both tracings, where the criterion would read it as part of the coast. A scan and a hand tracing, or two different digitisers, would be closer to independent.

The same coast, the same pieces. Everything is on the Japan coast of the earlier measurements, with its candidate projections and its twelve-point floor on a piece.

Drift, not easing. An eased copy is still paid for in pieces by the whitened criterion, because easing is in both tracings and cancels in their difference: the twin estimate sees none of it. This essay removes one of the two smooth departures the earlier one named, and leaves the other exactly where it was.

Still open: whether easing survives the drift being gone

With the drift undone, the one smooth departure left on a traced coast is the compiler’s own — the easing that blends a copy into its neighbours, which a copy eased into place loses its seams before its source measured and the earlier count paid for in pieces. That is a real property of the sheet, not of its tracing, and no second tracing will remove it.

What would separate it is a model that contains it: a piece that is a projection placed by a similarity and blended into its neighbours over a stated length, charged for that length as one more number. Whether such a model, on a whitened coast, can count an eased copy as one copy and report how far it was eased — and whether the easing length it reports is recognisable as a draughtsman’s rather than as whatever the criterion needed — are questions a rigid piece cannot ask. The sheet moved before it was measured adds the third departure, the paper’s own shrinkage, which is in both tracings too.

Named alongside this one

Essays reaching for the same objects. Nobody chose these; they are what the concept index makes visible.

The objects this essay names

Each one links to every other essay that touches it.

CompilationDegrees of freedomEstimatorLeast-squaresProjection identificationResidualSeamSearchSimilarityVerification