A residual cannot count the pieces of a map
Assumes A copy eased into place loses its seams before its source.
A compiled map agrees with its graticule except where it was copied found the seams in a compiled coast by assuming there were two of them. That was the whole of its model: one stretch of coast lifted from another sheet, two ends where it joins, and a search for the longest run of coast that agrees with the graticule. The model was right because the coast it was tested on had been built to match it.
A compiled map in an archive carries no such promise. It may have one borrowed stretch or four, or none. Some sheets were assembled from a dozen surveys by a draughtsman who never recorded which, and the first question anybody reading one has to answer is how many pieces there are. Where they join is the second question.
The obvious answer does not work, and the reason is structural rather than practical. Try every count, fit each as well as possible, keep the best. The best is always the largest count, because a coast cut into seven pieces contains every way of cutting it into two: set five of the seven pieces to agree with their neighbours and the two-piece answer is recovered exactly. More pieces cannot fit worse. So the residual goes down at every step whatever the truth is, and the count it points to is simply the largest one tried.
The coast here is the one the earlier measurements used. It is an outline of Japan on a conformal conic sheet, with the stretch from 55 to 85 per cent of the way round copied from a sinusoidal sheet of the same ground and fitted in by a two-point similarity. No graticule is needed to read it, since a map with no graticule showed an outline alone names its projection. It is digitised now at 360 points with noise of a ten-thousandth of the map’s width in each coordinate. With no seam, the best single projection fits it to 1.9 thousandths of the width. With the two seams that were made, it fits to one ten-thousandth, which is the noise, and nothing more should be possible. With six seams it fits to 0.97 ten-thousandths. The residual has gone below the noise, and the extra seams are fitting it.
What a piece is, and what it costs
To count pieces, a piece has to be defined in a way that makes it cost something.
A piece here is a stretch of coast explained as some candidate projection of that ground, placed on the page by a similarity. Nine candidates are offered: the conformal and equal-area conics, the polyconic, the sinusoidal, the Mercator, the Miller, the plate carrée, the Mollweide and Lambert’s cylindrical equal-area. The similarity is a scale, a rotation and a shift, which is four numbers, and each seam is one more, its position along the coast. So a model with pieces spends numbers on placing them and on placing the seams, and makes a choice among nine for each piece besides.
Given a count, the best segmentation is found exactly and not by guessing. There are only so many places a seam can go, one per digitised point. The best way to cover the first points with pieces is the best way to cover some earlier point with pieces, plus the best single piece from to . That recursion, run over every and , finds the optimum with no search heuristic in it.
It is fast because each stretch’s residual has a closed form. Writing points as complex numbers, the best similarity between a candidate’s stretch and the drawn stretch has
where and are the stretches with their means removed. Every sum in it is a difference of two running totals, so the residual of any stretch under any candidate costs a handful of subtractions. The whole search over seven counts, nine candidates and every pair of seam positions takes under a third of a second.
One guard is needed, and it is the same trap in miniature. A similarity has four numbers and fits any two points exactly, so a coast cut into pairs of points would have zero residual. Each piece is therefore required to be at least twelve points long, a thirtieth of the coast. That floor is a stated choice, and a much shorter copied stretch would be invisible to this whole method.
A charge for each number
The standard way to choose a count when every larger count fits better is to add to the misfit a charge for every number spent. Two charges are in common use, and they differ by a factor that turns out to decide everything.
AIC, Akaike’s criterion, charges two per number. BIC, the Bayesian or Schwarz criterion, charges the logarithm of the number of observations. Here that is , since each of 360 points gives two coordinates. Both measure the misfit the same way, by times the log of the mean squared residual, so a model whose extra pieces lower the residual by a fraction gains about . An extra piece with its seam costs five numbers. So AIC buys it if it lowers the residual by more than about 1.4 per cent, and BIC only if it lowers it by more than about 4.6.
On this coast, after the true two seams, each further seam lowers the mean squared residual by between 1.4 and 2.2 per cent. That is at or above AIC’s threshold and well below BIC’s. BIC stops at two. AIC does not: it rises to the largest count offered, six seams, and offered up to eleven it takes ten.
The drops are larger than they should be, and that is the whole story. A piece whose five numbers were merely estimated from pure noise would lower the mean squared residual by about five parts in 720, 0.7 per cent. These pieces lower it by two or three times that, because each extra seam goes in the best place out of hundreds and its piece takes the best projection out of nine. A search over that many options finds a placement that happens to fit the noise better than typical, just as the best of a thousand coin-tossers finds a run of heads. AIC’s charge of two per number is calibrated for a number that is estimated, not for one chosen by search, and the search is worth more than the charge. BIC’s charge of 6.58 is large enough to pay for the search as well, on this coast. That is a property of the charge’s size, not of its derivation.
Drawn out, AIC’s extra seams are visibly arbitrary. They split a conformal conic stretch into two conformal conic stretches, or hand a few dozen points to a neighbouring projection that fits the noise a little better over that span. None of them corresponds to anything that happened when the map was made. A reader following AIC would report a sheet compiled from six sources, most of them in the same projection as the host, with seams where the draughtsman never paused.
How far it can be trusted
A result on one noisy draw is an anecdote, so each count below is taken from five independent digitisings at each noise level. It is run for two compilations: the single copied stretch above, and a coast with two stretches copied from two different sheets, a sinusoidal between 12 and 30 per cent of the way round and a Mercator between 55 and 85.
BIC recovers the true count in every trial from a noise of a hundred-thousandth of the map’s width up to a thousandth: two seams when two were made, four when four were made. At three thousandths it undercounts, to one seam for the single copy and two for the pair. By then the copied stretches depart from the host’s coast by little more than the noise, and a criterion that charges for pieces correctly declines to pay for ones it cannot see.
AIC is wrong in all sixty trials, and always in the same direction. It never undercounts, even where the copy is buried in noise, and in fifty-seven of the sixty it takes the maximum it was offered. So its failure is not a matter of noise; it is built in.
On a real sheet a thousandth of the width is the practical boundary. On a sheet 60 centimetres wide that is 0.6 millimetres, about a line width, and a coast digitised from a scan is often good to a third of that. Compilations whose borrowed stretches depart by several thousandths of the width are the ones this method can count, and the Japan compilation’s largest departure is 9.4 thousandths.
The count survives when the names do not
At a thousandth of the map’s width, with two copied stretches, the departure profile is barely readable by eye, and BIC still counts four seams, the right number. They are located less well than at low noise: the worst is 6.4 per cent of the coast from its true place, against about half a per cent at a ten-thousandth.
The pieces it names are worse still, and the way they are wrong is instructive. The two copied stretches are named correctly, sinusoidal and Mercator. The three host stretches, all in fact conformal conic, are named equal-area conic, polyconic and polyconic. The answer is a set established why. Over a mid-latitude region the size of Japan, the three conics and the polyconic cannot be told apart at a thousandth of the map, and any of them is an admissible reading of the host.
That gives three quantities with three different tolerances to noise. The count is the most robust, because it only has to see that something changes. The seam positions come next, because they need the change located. The host’s identity is the most fragile, because it needs to tell apart projections that differ by less than the noise over the stretch in hand. The copied stretches’ identities are protected by the fact that they were chosen as the misfits, so they differ from their neighbours by definition.
Where the right criterion goes wrong
Everything above assumed the copy was moved rigidly. A copy eased into place loses its seams before its source showed what happens when it is not. A compiler pins the copied stretch at both ends and eases it along its length until it blends into the coast either side. The seams stop being corners, the method that found them loses them first, and the source’s identity survives longer.
The piece count fails in a different way, and a more dangerous one.
A rigid copy is counted correctly as two seams, and so is one eased over a couple of per cent of the coast either side. From 3.5 per cent of easing the count rises: three seams, then four at eight per cent, then six from ten per cent on. It does not fall. The earlier failure was a seam that could not be found. This one is seams that were never made.
The mechanism is the one the criterion is designed around. BIC chooses the count that best pays for the misfit, among models that are all of the form “stretches, each some projection placed by a similarity”. An eased stretch is not of that form. The blend between the host’s coast and the copy’s, graded along a tenth of the coast, is no projection of the ground, so no single piece can explain it. BIC sees a misfit it cannot pay off with the pieces it has, and the only currency it has is more pieces. It cuts the eased region into short stretches, each explained by whichever projection happens to follow that part of the blend. At twenty per cent of easing the pieces named include a Mollweide and an equal-area conic, neither of which was ever on the draughtsman’s table.
That is the general hazard of model selection, and a compiled map is an unusually clear instance of it. When the answer is not in the library met the same situation for a whole map drawn in a projection nobody had offered the fit. A criterion chooses among the models offered, and when the truth is not among them, it chooses the one that best imitates the truth, which may be much larger than anything real. Every extra piece BIC places on an eased coast is honestly earned against the models it was given, and every one of them is wrong about the map.
What a reader of a real sheet should take from this
Three practical rules follow, and each is a claim about which output to believe.
The count, from BIC, on a rigid compilation, is the most trustworthy thing the method produces. It is right in every trial here, for one copy and for two, down to a noise of a thousandth of the map. That is above the precision of a careful scan.
A count larger than the evidence of the drawing is a warning, not a result. If BIC reports five seams on a coast where the departure profile shows two obvious excursions, the likelier reading is one or two eased copies, not five sources. The test is whether the extra seams cluster at the ends of the obvious excursions, which is where easing lives.
AIC should not be used for this at all. Its charge is too small for a model whose numbers are chosen by search. It will always report the largest count offered, whatever the sheet.
The first rule is also a correction to the method the earlier seam-finding essays built. That method assumed two seams and was accurate given the assumption. With BIC in front of it the assumption becomes a measurement, and a residual has more than one explanation is the reminder that a measured count is still a statement about the models it was chosen from.
The controls on these numbers
The residual must never rise with the count. A model with pieces contains every model with , so a segmentation whose best residual went up with an extra piece would mean the search had missed its optimum. It never rises, on any trial.
At a ten-thousandth of the map the count must come out right and AIC’s must not. That is the claim, so it is required of the result and not merely read off it. BIC gives two seams, AIC six.
A rigid copy must be counted as two, or the eased copies’ larger counts could be the criterion’s own error rather than the easing’s. It is, at zero easing.
The seams chosen at low noise must fall where they were made. At a ten-thousandth they are within about half a per cent of the coast, one or two points of 360.
And each piece’s residual must agree with a fit made the long way. The closed form above is the least-squares similarity. On any stretch it must give the same residual as the four-by-four normal equations the earlier measurements solved, and it does, to four parts in ten million on the stretches tried.
What the count does not cover
Nine candidates, all ordinary. A piece copied from a projection not among them would be explained, badly, as several pieces of ones that are, which is the eased-copy failure arriving by another route. The library in the earlier essays holds twenty; the nine here are the ones whose sheets centred on Japan are plausible, which leaves out the azimuthals and several world maps.
Linear from the north point. The walk round the coast starts at its northernmost point, so a copied stretch that straddled that point would be counted as two stretches with a seam between them that does not exist. A circular walk removes that at the cost of trying every starting point, 360 times the work.
Rigid copies at a stated noise. The noise is independent from point to point. A real digitising error is correlated along the line, because a hand drawing a coast drifts rather than jitters, and correlated error looks like a smooth departure, which is to say like an eased copy. How much of a real scan’s error BIC would count as pieces is not measured here, and it could be a great deal.
Twelve points is the shortest piece. A copied stretch of fewer than twelve points, a thirtieth of this coast, cannot be counted however large its departure.
Still open: whether easing and drift can be told apart
The eased-copy result leaves a question that decides whether any of this can be used on an archival sheet. BIC over-counts an eased copy because easing is a smooth departure no projection explains. A digitiser’s slow drift along a coastline is also a smooth departure that no projection explains. So is the paper’s own uneven shrinkage, which the sheet moved before it was measured found to be real and directional. On a real sheet all three are present at once, and all three will be paid for in pieces.
What would separate them is a model that contains them. That means a smooth deformation term shared across the whole sheet for shrinkage, one per stretch for easing, and a correlated noise model for drift, each charged for honestly. A criterion choosing among those could say “two sources, eased, on paper that shrank”. Whether the terms can be told apart by one coastline, or whether they trade off against each other as a datum and a projection’s parameters were found to do, is a question the rigid-copy model cannot ask.
Named alongside this one
Essays reaching for the same objects. Nobody chose these; they are what the concept index makes visible.
- The residual reports the error the fix was immune to degrees of freedom · estimator · least-squares · residual
- What the extra unknown costs where nothing can see it degrees of freedom · estimator · least-squares · residual
- A coordinate is the output of a solve degrees of freedom · least-squares · residual
- A map does not say what it is least-squares · projection identification · residual
- The nearest map to an impossible request degrees of freedom · least-squares · residual
- The span ladder, run on all five estimator · least-squares · seam
The objects this essay names
Each one links to every other essay that touches it.
CompilationDegrees of freedomEstimatorLeast-squaresProjection identificationResidualSeamSearchSimilarity