What a machine does with it

When the edges do not line up

Rung eight held the cell counts equal so that shape could be compared without the count ratio drowning it, and recorded a doubt: a longitude–latitude source shares its boundaries with a longitude–latitude target wherever their counts share a factor. The mechanism is real and worth a factor of two. It was not what the published number was made of.

Assumes The same number of cells, in two shapes.

The same number of cells, in two shapes set out to answer a question the ladder had been unable to ask. Rebinning a field from one cell scheme to another loses accuracy, and how much it loses depends on two things at once: how much coarser the target is, and what shape the cells are. Holding the counts equal removes the first, and what was left was a separation of about twenty-five points between a gnomonic cube source and a longitude–latitude one, in the cube’s favour.

That essay recorded a doubt about its own answer, in these words: the matched-count rebinning still favours the longitude–latitude source. Its column edges coincide with the target’s wherever the counts share a factor, so the separation is a lower bound on the shape effect rather than an estimate of it.

The doubt was well founded as a mechanism and wrong as a diagnosis, and the difference between those two is the essay.

What a rebinning loses depends on where the target's edges are. Two grids of fixed counts, fixed shapes and fixed resolution, with the target slid across the source from perfect alignment to a full cell. Nothing about either grid changes except where its boundaries fall. The loss runs from 27.5 per cent at zero to 56.3 at half a cell — a factor of 2.04 — and the longitude-only curve returns to its starting value at a full cell to six decimal places, which is the periodicity check. A cell boundary that coincides with a target boundary loses nothing, and a grid comparison that does not say where its boundaries are has left that out.
Fig. 1 Two grids of fixed counts, fixed shapes and fixed resolution, with the target slid across the source from perfect alignment to a whole cell. Nothing changes except where the boundaries fall. The loss runs from 27.5 per cent at zero to 56.3 at half a cell — a factor of 2.04 — and the longitude-only curve returns to its starting value at a full cell to six decimal places.

The mechanism, drawn

Rebinning is an area-weighted average: each target cell takes the mean of the source values over the ground it covers. If a target cell covers exactly one source cell, that mean is the source value and nothing is lost at all. If it straddles two, the two are blended, and blending is the whole of the error.

What sharing a factor looks like. The column boundaries of a source grid and a target grid, in longitude, for three cases. On the left the two counts are equal and every boundary coincides, so no target cell straddles a source cell and the rebinning is a relabelling. In the middle the same two grids are slid half a cell apart and every target cell draws on two sources. On the right the counts are coprime and only the seam is shared. This is the whole mechanism, and it has nothing to do with the shape of a cell or with how much coarser the target is.
Fig. 2 The column boundaries of a source and a target in longitude, for three cases. Equal counts and no offset means every boundary coincides and the rebinning is a relabelling. The same two grids slid half a cell apart means every target cell draws on two sources. Coprime counts mean only the seam is shared. This is the whole mechanism, and it has nothing to do with the shape of a cell or with how much coarser the target is.

So a grid comparison that does not say where the boundaries are has left out a variable, and the variable is worth a factor of two on the experiment above. That is larger than the shape effect the whole previous rung exists to measure.

Sliding is the clean experiment

The offset sweep is the manipulation this rung is built on because it changes exactly one thing. The source keeps its counts, its cell shapes, its aspect ratio, its resolution and its field. The target keeps all of the same. Only the phase of the target’s boundaries moves, and it moves continuously from zero to one cell.

Two things in that curve are worth reading beyond the factor of two.

It is exactly periodic in longitude. Slid by a whole target column, the loss returns to 27.542 per cent from 27.542 — six decimal places, which is arithmetic rather than agreement. That is the check that the quantity being measured is the phase and not something else that happens to vary with the offset.

It is not periodic in latitude, and that is honest rather than a bug. A longitude banding is a circle and shifting it by a whole cell returns it to itself. A latitude banding has ends: shifting the interior boundaries by a whole row thickens the polar row at one end and thins it at the other, because the outer boundaries have to stay at the poles. So the two-coordinate curve comes back to 33.1 rather than to 27.5, and the periodicity check is run on the coordinate that is genuinely a circle. This is the same asymmetry the antimeridian essay is about, arriving in a place nobody would look for it.

And the counts comb

The offset sweep is clean and it is also artificial: nobody slides a grid. What people do is choose counts, and the counts decide the alignment for them.

The loss combs with the arithmetic of the counts. A source of sixteen columns rebinned to targets of eleven to twenty-one, and back. The underlying trend is downward, because a target with more columns is a finer target. Sitting on it is a comb: every count sharing a factor with sixteen dips below the trend, and the sixteen-column target — where every source boundary is a target boundary — loses 27.5 per cent against 50.1 and 47.0 on either side. A count ratio cannot produce a dip; only shared edges can.
Fig. 3 A source of sixteen columns rebinned to targets of eleven to twenty-one columns and back. The underlying trend is downward, because a target with more columns is a finer target. Sitting on it is a comb: every count sharing a factor with sixteen dips below the trend, and the sixteen-column target loses 27.5 per cent against 50.1 and 47.0 on either side of it.

The dip is the argument. A count ratio is a smooth, monotone function of the column count, so no amount of count-ratio effect can produce a low point at sixteen with higher values at fifteen and seventeen. Only shared boundaries can, and they do — a 45 per cent reduction in the loss, from choosing a number rather than from choosing a geometry.

The lesser dips are the same effect at lower order. Twelve and twenty share a four with sixteen, eighteen and fourteen share a two, and each sits below the line through its coprime neighbours by an amount that falls with the size of the shared factor.

Eight sources against one target, ordered by what they share with it. Each row is a longitude–latitude source rebinned to the same 16 × 8 target and back, labelled with the product of the two greatest common divisors. The sources sharing a large factor average 32.9 per cent and those sharing almost nothing average 48.4. This is the discrete corroboration of the sliding experiment and it has confounds the sliding experiment does not: these grids differ in cell count and in aspect ratio as well as in what they share, which is why it is reported beside the slide rather than instead of it.
Fig. 4 Eight longitude–latitude sources rebinned to the same 16 × 8 target and back, ordered by the product of the two greatest common divisors. The sources sharing a large factor average 32.9 per cent and those sharing almost nothing average 47.6. This is the discrete corroboration, and it carries confounds the sliding experiment does not — these grids differ in cell count and in aspect ratio as well as in what they share.

That last sentence is why there are two experiments rather than one. Integer grids cannot be made to vary in commensurability while holding everything else fixed: changing a count changes the resolution and the cell shape too. The slide holds everything and is artificial; the counts are natural and hold nothing. They agree, which is the only reason either is worth reporting.

What it does to the published number

Now the correction, and it is not the one the shortfall expected.

The correction to the published separation, which is small. The recorded shortfall said the rectangular source's edges coincide with the target's wherever the counts share a factor, and therefore that the published separation understated the cube's advantage. The mechanism is real and is worth a factor of two on grids that do share factors. It was not what the published number was made of: a source of ninety-six cells comes out at 14 columns by 7 rows, 14 shares only a 2 with 16 and 7 shares nothing at all with 8, so the run was already nearly incommensurate. Sliding the target half a cell moves the separation from 24.3 points to 24.9.
Fig. 5 The cube source, the published rectangular source, and the same rectangular source with the target slid half a cell. The published separation was 24.3 points in the cube’s favour and the corrected one is 24.9 — a movement of six tenths of a point, on a mechanism worth a factor of two elsewhere.

The reason is arithmetic that nobody did. A source of ninety-six cells, laid out as near square as the count allows, comes out at fourteen columns by seven rows. Fourteen and sixteen share a two. Seven and eight share nothing at all. So the published run was already close to incommensurate, most of its boundaries already fell inside target cells, and there was very little alignment left to remove.

The shortfall said wherever the counts share a factor, which is true, and then assumed that this experiment’s counts did. They did not, and the way to find out was to compute two greatest common divisors — a step that takes a second and that neither the original run nor its recorded doubt took.

The mechanism was real, the magnitude was real, and the diagnosis was wrong. The separation between the two geometries was already an estimate rather than a bound, and it is now known to be one for a measured reason rather than assumed to be either.

What this changes about the earlier rungs

Two earlier measurements on this ladder used grids whose commensurability was never stated, and both need re-reading.

The same data on two grids compares an equal-angle scheme with an equal-area one at matched counts. Their boundaries in longitude coincide whenever the column counts share a factor, exactly as here — but their boundaries in latitude never coincide at all, because one bands in latitude and the other in the sine of latitude, and those two agree only at the equator and the poles. So that comparison is protected in the coordinate that matters and exposed in the one that does not, which is a piece of luck rather than a design.

Cells that are rectangles in no coordinate compares a clipped cube with a longitude–latitude grid, and a cube face’s boundaries are not curves of constant longitude, so no coincidence is possible anywhere except at the four points where a face corner happens to sit on the equator. That measurement is immune, and immune by construction rather than by choice.

The vulnerable comparison is the one between two schemes of the same family, which is the comparison practitioners actually make: two longitude–latitude products at different resolutions, which is what almost every climate and population dataset is rebinned between. Halving a grid shares every boundary. Going from a one-degree grid to a 0.75-degree grid shares one boundary in four. Going from one degree to 1.1 degrees shares one in eleven, and loses measurably more.

Two source geometries of 96 cells each, rebinned to the same three targets. Both curves start from a source of 96 cells and rebin to targets of 32, 128, 512 cells, so the count ratio is identical along them and the only difference is the shape of the source cells: gnomonic squares on a cube against rectangles in longitude and latitude. The ratio dominates — both curves fall by more than half across the range — and the shapes still separate by 25 points at the middle target. The cube loses less, because its cells are all much the same size and the lon/lat source's collapse towards the poles.
Fig. 6 The measurement this rung is a correction to. Two source geometries of ninety-six cells each, rebinned to three targets of increasing fineness, with the count ratio identical along both curves. The separation between them is what the alignment question was about, and it survives the correction essentially unchanged.

The number a practitioner needs

Everything above can be reduced to one rule, and the rule is unwelcome.

Rebinning between two grids of the same family, choose counts that share a factor. A one-degree grid onto a two-degree grid loses the least any rebinning can lose, because every target boundary is a source boundary and no target cell straddles anything. A one-degree grid onto a 1.1-degree grid shares one boundary in eleven and loses close to the maximum. The choice is free — it costs nothing to pick 0.5 instead of 0.45 — and it is worth more than any amount of care about the interpolation.

The unwelcome part is that the rule cuts against the other one. A cell system trades area for shape and an address is an area between them establish that the grid should be chosen for what the data is for: a resolution that matches the phenomenon, cells whose areas are comparable, an addressing scheme whose ordering suits the queries. None of those considerations has anything to say about divisibility, and a resolution chosen well for the phenomenon has no reason to share a factor with the one it will be compared against.

So the rule is really a rule about pairs, and a dataset does not know in advance what it will be rebinned onto. What a publisher can do is state the boundaries — not the cell size, which is what every specification gives, but the phase — because two one-degree grids offset by half a cell are not the same grid and their difference is worth a factor of two.

Where a rebinning does its damage. The round-trip error against latitude, as a root mean square over each row of cells and as the worst cell in it. The error is 0.276 at -83° and 0.096 near the equator — a factor of 2.9. The two schemes agree best where their cells are most alike, and an equal-angle cell and an equal-area cell are least alike where the equal-angle one has collapsed.
Fig. 7 Where a rebinning loses, by latitude. The loss is not spread evenly: it concentrates where the source cells are largest relative to the target’s, which on a longitude–latitude grid is near the equator in longitude and near the poles in area. The alignment effect rides on top of this, so a grid pair that is commensurate in longitude is protected exactly where the loss would otherwise be worst.

Why the doubt was recorded, and why it was wrong

The shortfall this essay pays was written by the run that produced the number, at the moment it produced it, in one sentence. That practice is the reason this rung exists at all: nothing else would have brought anyone back to a measurement that passed every gate and looked finished.

It is also the reason the diagnosis went wrong. A shortfall written at the end of a run is written from the mechanism rather than from the data — the author knows that shared factors flatter a rectangular source, notices that both grids are rectangular, and writes it down. Computing the two greatest common divisors would have taken a second and would have said not this time.

The lesson is not to write fewer doubts. It is that a recorded doubt is a hypothesis rather than a result, and this collection has now met the distinction twice: the aspect valley was described as a valley for two rungs before anybody sampled its level set and found fourteen disconnected basins. Both times the description was reasonable, both times it was written by whoever knew the mechanism best, and both times the measurement disagreed.

What survives repeated rebinning, and what does not. The field's variance after each round trip, as a share of what it started with: 61 per cent after one and 23 after 6. The total is preserved to 1e-15 at every one of them. Rebinning is an averaging operator, and repeated averaging is a diffusion — the data becomes smoother every time it is moved, and the number that would reveal it is the one that never moves.
Fig. 8 What repeated round trips do, which is the reason any of this is worth a factor. Each pass through a coarser grid and back blends a little more, and the loss accumulates rather than settling — so a pipeline that rebins six times has spent six times the alignment penalty, and a pipeline that happened to choose commensurate grids has spent almost none of it.

The variable nobody publishes

A dataset’s documentation states its resolution. It states the cell size, the projection, the datum, the extent and often the nominal accuracy. It does not state the phase — where the cell boundaries fall — because a grid described as “one degree” is assumed to have its boundaries at whole degrees, and very often it does not.

That assumption is exactly what this rung shows to be worth a factor of two. Two one-degree grids offset by half a cell are not the same grid, rebinning between them loses twice what rebinning between aligned ones loses, and nothing in either dataset’s documentation distinguishes the two cases.

The cheapest possible remedy is a sentence: state the coordinate of one cell corner. It is one number, it is already known, and it turns an unmeasurable variable into a stated one — which is the same move an address is an area makes for a cell identifier and a coordinate is a number with a width makes for a written coordinate.

Where the model stops

The field is a single stated function. MATCHED_FIELD is a smooth analytic surface with structure at about the scale of the coarser grid, chosen so that a rebinning has something to lose. A field with structure much finer than either grid would lose almost everything at any offset and the alignment effect would vanish into the floor; a field much smoother than both would lose almost nothing and it would vanish into the ceiling. The factor of two is a property of the field’s spectrum relative to the grid as much as of the grids.

Only two round trips are measured. Everything here is source → target → source, which is the operation whose error can be compared against the original values. A one-way rebinning has no such reference, and the two are not the same: a round trip squares the blending, so the numbers above are pessimistic for a pipeline that rebins once and stops.

Latitude alignment is only half-explored. The slide moves both coordinates together and the periodicity check runs on longitude alone. What a latitude-only slide costs, and whether the polar-row artefact contaminates it, is not measured here — and it is the coordinate in which real grids differ most, because the choice between banding in latitude and banding in its sine is exactly a choice about where the row boundaries fall.

And nothing here is about the sphere. Every result in this essay would hold on a plane, between two rectangular grids, with no projection anywhere. That is unusual on this site and worth saying: the alignment effect is a property of piecewise-constant representation, and it is included on this ladder because the ladder’s own measurements were contaminated by it, not because it is cartographic.

An effect that is not cartographic, on a cartographic ladder

The admission that nothing here is about the sphere deserves more than a paragraph, because it says something about how a measurement programme finds its subjects.

The effect was not sought; it was met. The ladder was measuring what rebinning costs between two cell schemes, and the numbers moved for a reason that had nothing to do with the sphere — the alignment of the two grids’ boundaries, which is a property of piecewise-constant representation and would be present between two rulers.

Including it is a decision about honesty rather than about scope. The alternative was to control for it silently and report the corrected numbers, which would have produced a cleaner ladder and hidden a variable that every other published rebinning comparison also has and none of them states. The measurements on this ladder were contaminated by it, so the ladder owes the reader the contamination.

And it is the more transferable of the two findings. The spherical rebinning numbers apply to two particular cell schemes; the alignment effect applies to any pair of piecewise-constant representations anywhere, which includes histogram rebinning, image resampling to a shifted grid, time-series re-bucketing and every raster reprojection in this collection’s applied field.

It also changes what the earlier rungs’ numbers mean. Any comparison on this ladder that did not control for alignment was reporting a mixture of the effect it intended to measure and this one, in proportions set by an accident of where the two grids’ boundaries happened to fall — which is why the doubt was recorded before it was understood, and why the record turned out to be pointing at a real variable rather than at a suspicion.

Which is a reminder about where a measurement’s value ends up. A programme aimed at one subject finds its sharpest results in the machinery it had to build to get there, and the general result arrives as a by-product of a specific one rather than from having set out to be general.

Recording a doubt one cannot yet explain is the practice that made the difference, and it is worth saying so plainly: the variable was found because somebody wrote down that a number looked wrong.

It could as easily have been controlled away and forgotten.

Where the ladder goes next

The alignment variable is now measured and can be controlled for. Two questions it opens are not answered here.

A rebinning between commensurate grids is a different operation. When every source boundary is a target boundary the area-weighted mean is a plain sum, no interpolation happens, and the round trip is a projection in the linear-algebra sense — idempotent, with a kernel that is exactly the within-target variation. That suggests the loss should be computable in closed form for the commensurate case, against a measured 27.5 per cent here, and that has not been checked.

And the comb should have a shape. The dip depth falls with the size of the shared factor, and the four points measured here are consistent with several laws. Whether it goes as the reciprocal of the shared factor, or as the fraction of boundaries that coincide, or as something else, is a question with an exact answer and one figure’s worth of work.

Named alongside this one

Essays reaching for the same objects. Nobody chose these; they are what the concept index makes visible.

The objects this essay names

Each one links to every other essay that touches it.

AggregationAliasingCellCell systemCombinatoricsConservationError budgetRebinningResamplingResolutionSampling latticeVerification