What a machine does with it

The same number of cells, in two shapes

Moving a field between two cell schemes loses 18 per cent of it per cell in one geometry and 39 in another, and the earlier measurement could not say whether that was the shape of the cells or the ratio of their sizes, because changing the schemes changed both. Holding the counts settles it: the count ratio decides most of the loss, and the shape is still worth a quarter of the field.

Moving a field between two cell schemes conserves the total exactly and loses the per-cell values, and the loss was measured at 39 per cent of the field’s own standard deviation after one round trip between two longitude–latitude schemes. Doing it between a gnomonic cube and a longitude–latitude grid, where every overlap has to be clipped on the sphere rather than read off a closed form, gave 18.4 per cent.

Eighteen against thirty-nine is the wrong way round from what the harder geometry ought to give, and that essay said so and recorded why it could not be read as a result: changing the schemes changed the ratio of cell sizes as well as their shape. A rebinning’s per-cell loss is dominated by how much coarser the target is than the source. The shape is a second-order effect on top of that, and the two were confounded.

What settles it is holding the counts.

Two source geometries of 96 cells each, rebinned to the same three targets. Both curves start from a source of 96 cells and rebin to targets of 32, 128, 512 cells, so the count ratio is identical along them and the only difference is the shape of the source cells: gnomonic squares on a cube against rectangles in longitude and latitude. The ratio dominates — both curves fall by more than half across the range — and the shapes still separate by 25 points at the middle target. The cube loses less, because its cells are all much the same size and the lon/lat source's collapse towards the poles.
Fig. 1 Two source geometries of 96 cells each, rebinned to targets of 32, 128 and 512 cells and back. The count ratio is identical along both curves, so the only difference between them is whether the source cells are gnomonic squares on a cube or rectangles in longitude and latitude. The ratio dominates — both curves fall by more than half across the range — and the geometries still separate by 25 points at the middle target.

The experiment, and why it is the one that was owed

Both runs start from a source of 96 cells. In the clipped case that is a gnomonic cube at level two: six faces of sixteen cells, whose boundaries are great-circle arcs and whose overlaps with anything must be computed by clipping spherical polygons against planes. In the rectangular case it is a longitude–latitude scheme of sixteen columns and six rows, whose overlaps with the target are rectangles with closed-form areas.

Both rebin to the same three targets — longitude–latitude grids of 4 × 8, 8 × 16 and 16 × 32 cells — and back. Same field, same counts, same operation, two geometries.

The results:

target cells ratio to source cube source lon/lat source
32 0.33 75.9 % 76.3 %
128 1.33 28.2 % 53.0 %
512 5.33 17.3 % 32.6 %

Two things fall out and they answer two different questions.

The count ratio is the dominant variable, in both geometries. Going from a target three times coarser than the source to one five times finer takes the loss from 76 per cent to 17 in one geometry and from 76 to 33 in the other. That is the effect the earlier measurement was confounded by, and it is large.

The shape is worth a quarter of the field, once the counts are held. At 128 target cells the cube source loses 28.2 per cent and the longitude–latitude source loses 53.0 — a separation of 24.8 points that has nothing left to be a size effect.

Which way round, and why

The cube loses less, which is the opposite of the intuition that a harder geometry should cost more, and the reason is not the clipping at all.

A longitude–latitude scheme’s cells collapse towards the poles: at sixteen columns and six rows, a polar cell is a fraction of the area of an equatorial one, and the scheme’s cell areas span a factor of several. A gnomonic cube’s cells vary by a factor of about 1.5 across a face and are the same from face to face — that is what the cube scheme is for, and it is the trade the whole cell ladder is about.

A rebinning is an area-weighted average, so its loss per cell is governed by how much averaging happens inside each target cell. A source whose cells vary wildly in size does more averaging where its cells are large and less where they are small, and the round trip’s per-cell error is dominated by the places where the coarse source cell had to be spread back over many fine ones. The cube’s uniformity means the amount of averaging is nearly the same everywhere, and the error is correspondingly smaller.

So the 18.4 against 39 recorded earlier was not an artefact of the count ratio after all. It was an underestimate of a real effect that the count ratio happened to be masking — which is the opposite of what the shortfall predicted, and is the reason the shortfall was recorded rather than the number quoted.

The piece two schemes share, clipped rather than assumed. A cell of a gnomonic cube, whose four edges are great-circle arcs because a straight line on a gnomonic face is one, against a cell of a longitude–latitude grid, whose north and south edges are parallels and are not. Their overlap is neither a rectangle nor a spherical polygon of any standard kind, and it is 5965687 km² of the cube cell's 5965687 km² — 100.0 per cent. Computing it needs the arc of one boundary intersected with the plane of the other, which is three equations and two roots, and it is exact.
Fig. 2 The two geometries whose overlaps are being computed: a gnomonic cube and a longitude–latitude grid, drawn on the same map. Every overlap between them is a spherical polygon bounded by great-circle arcs and by parallels, and the clipper handles all three kinds of boundary as planes through the centre of the sphere.

Conservation, at two different precisions

Both round trips conserve the total, and they do it to different precisions, and the difference is worth stating because it is not a defect.

The rectangular round trip conserves to between zero and 7 × 10⁻¹⁶ relative — machine precision — because every overlap area is a closed form and the sums telescope exactly. The clipped round trip conserves to between 2 × 10⁻¹⁶ and 3 × 10⁻¹³, because every overlap area is the output of a clip and a spherical excess, and the conservation is as good as that arithmetic and no better.

Three parts in a thousand billion is not a limitation of anything. It is quoted because a check that reports “conserved” without a number cannot tell the difference between a formula that is exact and one that is nearly exact, and the two behave differently when the schemes get finer.

The conservation survives the hard geometry and the field does not. A field carried from a 96-cell gnomonic cube onto a 128-cell longitude–latitude grid and back, with every one of the 12288 overlaps clipped exactly. The total is conserved to 1.7e-12 relative — not the 2 × 10⁻¹⁶ of the rectangular case, because there every overlap area is a closed form and here every one is the output of a clip and a spherical excess. The per-cell error is 18.4 per cent of the field's own standard deviation. The check a validation suite runs is the total, and the total is exactly the quantity that cannot see any of it.
Fig. 3 The clipper’s own audit: every overlap it computes is a piece of the sphere, and the pieces must add to the sphere. If a spherical polygon were being clipped wrongly — an arc–plane root taken on the wrong side, a fan triangle with the wrong sign — the pieces would not add up, and every other number in this essay would be a measurement of that bug rather than of the geometry.
Equal area or steady shape, and not both. Four cell schemes plotted by how much their cells vary in area and how far from square the worst of them is. The bottom-left corner is the scheme that has both, and it is empty: the equal-area cube holds area to 1.003 and has the most elongated cells, the tangent-warped cube has the tightest shapes and lets area vary by 1.28, and the lon/lat scheme is off the scale on both. Neither axis can be driven to one while the other stays there.
Fig. 4 The trade the cube scheme makes, at the resolution the source above uses: area against shape, across the warps available. The uniformity of area that makes the cube the better rebinning source is bought with a shape distortion that a longitude–latitude grid does not have — so the result of this essay is not that one scheme is better, it is that the trade has a term nobody was counting.

What the count ratio actually does

The dominance of the count ratio is worth one more paragraph, because it is the practical result even though it is the boring one.

Rebinning to a coarser target and back is a low-pass filter applied twice. What survives is whatever the coarse target could represent, and what is lost is everything finer — so the loss is set by the target’s resolution against the field’s own structure, and the source’s geometry only decides how cleanly the filtering is done.

At a target three times coarser than the source, both geometries lose about 76 per cent, and the two curves are within half a point of each other. That is the regime where the filtering swamps everything: the target is so coarse that the shape of the source cells cannot matter, because almost all the information is being discarded regardless.

The geometries separate in the middle of the range, where the target is comparable to or finer than the source and the loss is dominated by how the source’s own averaging was arranged rather than by the target’s coarseness. That is the regime where a discrete global grid’s design pays, and it is also the regime almost all real reprojection of gridded data operates in.

Why the middle target is the one to quote

Three targets give three separations — 0.4, 24.8 and 15.3 points — and it is worth saying which of them is the number and why.

The coarsest target is in the regime where the filtering swamps the geometry, so its separation is a measurement of nothing. The finest is in the regime where the target can represent almost everything the source held, so the loss is small in both geometries and the separation shrinks with it. The middle target — where the target is about as fine as the source — is where the two effects are comparable and the geometry has the most room to matter.

That is not a convenient choice, it is the structural one: any comparison of two operators has a regime where a nuisance variable dominates, a regime where the operator barely acts, and a band between them where the difference is visible. Quoting the number from the band and saying which band it is beats quoting the largest of three and calling it the effect.

It also means the honest form of the result is a curve rather than a number, which is why the figure at the head of this essay is a curve and the table above has three rows.

What was computed, and how

The field is a stated sum of sinusoids with structure at several scales, evaluated at cell centres. It is not data and the caption of every figure says so: what is being measured is the operator, and for that the input has to be one whose spectrum is known.

The rebinning is area-weighted: each target cell’s value is the average of the source values over the overlap areas. That is the operator that conserves the total, and conservation is the property the whole ladder is built on — an address is an area, so a value attached to a cell is a value attached to a region, and moving it must preserve the integral.

The source in the rectangular case is built to the same count as the cube by choosing the number of rows nearest to the square root of half the count, which gives cells as near square as 96 allows. That is a choice and it matters: a 96-cell scheme of 96 columns and one row would have given a different answer, and the honest version of “the same count” is “the same count, arranged as squarely as the count allows”.

The check that separates the two claims is deliberately in two parts. The count ratio must dominate — both geometries’ losses must fall by more than half across the range — and the geometries must still differ once the counts are matched. If only the first held, the shape would be shown to be irrelevant; if only the second, the earlier confound would still be live.

The same address length, a tenth of the area. Every cell of a lon/lat quadtree at level 4 carries an identifier of the same length. The heavy curve is each cell's area as a fraction of the largest, against its latitude: a polar cell is 10.2 times smaller than an equatorial one. The light curve is the inverse of the cell's aspect ratio, which falls from 1.00 near the equator to 0.10 at the top — the cells stop being anything like square long before they stop being usable.
Fig. 5 The cube scheme at four levels, with the cell areas measured rather than assumed. The uniformity that makes the cube the better source is visible here as the narrowness of the area distribution at every level, and it is the property a longitude–latitude scheme does not have at any resolution.
Tangent-warped cube cells, shaded by area. The tangent-warped cube at level 2, drawn on Mollweide with each cell shaded by its own measured area. The largest cell is 1.21 times the smallest. Every area is computed with the spherical polygon formula from the cell's own boundary, not from the scheme's intentions, and the shading is what the numbers say rather than what the mesh looks like.
Fig. 6 The source geometry itself: a gnomonic cube at level two, ninety-six cells, drawn with their own boundaries. Every cell here is bounded by great-circle arcs, which is why its overlap with a longitude–latitude rectangle has no closed form and has to be clipped.

What a matched comparison could not fix

One thing about this experiment stays confounded and it is worth naming rather than leaving as an implication.

Matching the counts holds the number of cells and the arrangement of the targets. It does not hold the alignment between the source and target cells, and it cannot: a cube’s faces meet the graticule at angles that vary across the sphere, while a longitude–latitude source’s cell edges lie along the target’s own edges wherever the row and column counts share a factor.

That alignment is worth something. A source cell whose boundary coincides with a target cell’s boundary contributes to exactly one target cell and loses nothing to splitting; a source cell straddling four target cells is spread across all four. The longitude–latitude source at 16 × 6 against a target at 16 × 8 shares its column edges exactly, so it is favoured by the alignment — and it still loses nearly twice as much.

So the 24.8 points is a lower bound on the shape effect rather than a clean estimate of it. Removing the alignment advantage would need a lon/lat source whose column count shares no factor with any target’s, which is a further experiment and one this rung does not run.

Where the model stops

Two geometries, not a survey. The comparison is between a gnomonic cube and a longitude–latitude grid. HEALPix, H3 and the equal-area cube would each give a different number, and the ordering among them is a measurement nobody has run here.

One count, three ratios. The source is 96 cells in both cases, which is small. Whether the 25-point separation holds at 10,000 cells is not established; the mechanism suggests it should, because it depends on the ratio of cell areas within a scheme rather than on their number, and that ratio is a property of the scheme’s construction.

One round trip. The loss compounds with repeated trips, and it compounds towards the mean because rebinning is an averaging operator. Nothing here measures the second trip.

The field’s spectrum decides the numbers. A smoother field loses less and a rougher one more, at every count and in both geometries. What transfers is the comparison, because both geometries see the same field.

The generalisation

The result is a pair, and the pair is the useful form:

Most of what a rebinning costs is set by the resolution ratio, and a user choosing a target grid should think about that before anything else. A target twice as coarse as the source throws away most of the field however elegantly the cells are shaped.

The rest is set by how uniform the source’s cells are, and it is worth a quarter of the field. A scheme whose cells vary in area by a factor of several does its averaging unevenly, and the unevenness does not average out.

That is a second argument for the discrete global grids, and it is not the argument they are usually sold on. The published case for them is about indexing, neighbour-finding and the absence of a pole singularity. This is about the arithmetic of moving data between grids, which is the operation those systems are most often used for, and it says the uniformity pays there too.

Who found it, and when

Area-weighted rebinning under the name conservative remapping is standard in climate modelling, where fields are moved between an atmosphere grid and an ocean grid on every coupled timestep. Jones’s 1999 algorithm — clip the source cells against the target cells on the sphere, weight by the overlap areas — is the one every coupler still uses, and its conservation property is exactly the one measured above.

The per-cell error is discussed in that literature as a smoothing, and higher-order conservative schemes exist that reconstruct a gradient inside each source cell and lose less. What is rarely separated is the contribution of the grid’s own uniformity from the contribution of the resolution ratio, because in a real coupling both change at once and neither is under the modeller’s control.

Why separating the two contributions needed an experiment

The rung’s design is the part worth carrying, because the confound it removes is the reason the question had not been answered.

In any real coupling, both variables move at once. A modeller replacing one grid with another changes the cell shape and the resolution ratio, because the two grids were built by different groups for different purposes and neither is adjustable. So every comparison available in practice reports the combined effect, and there is no way to attribute it.

Which makes the literature’s silence reasonable rather than negligent. The per-cell error is discussed as a smoothing, and the discussion is correct; what is missing is an attribution, and an attribution requires holding one variable fixed, which is something only a synthetic experiment can do.

Matching the cell counts is what buys the attribution here, and it is exactly the move that is unavailable to somebody with two real grids. Two schemes with the same number of cells and different cell shapes differ in one thing, so the difference in loss is that thing.

And the result is worth having precisely because it is not what a practitioner can measure. A modeller who knows how much of their loss is shape can decide whether a different tessellation would help; one who only has the combined figure cannot tell whether they are looking at a grid problem or a resolution problem, and the two have completely different remedies.

The corresponding obligation is to say what the constructed case leaves out, which here is that neither scheme is one anybody couples in production, so the sizes of the two contributions are established and their relevance to a particular real pairing is not.

The general form is the collection’s standing method. When two causes are confounded in every naturally occurring case, the way to separate them is to construct a case that does not occur — and the constructed case is not less real for having been built, since the quantity it isolates is present in all the natural ones.

That limitation is a reason to run the experiment again on a real pairing, not a reason to distrust the attribution.

Where the ladder goes next

Eight rungs have taken an address, made it an area, priced the trade between area and shape, indexed it with a curve, moved a field between two schemes, done it where the overlaps must be clipped, and now separated the shape’s contribution from the size’s.

What none of them has done is put an error bar on a cell value. A gridded product usually carries one, the rebinning has to propagate it, and an area-weighted average of correlated cells does not combine uncertainties the way an average of independent ones does — which is the same correlation a baseline’s uncertainty turns on, arriving from the other side of the collection.

What this makes readable

Essays that name this one as a prerequisite.

Named alongside this one

Essays reaching for the same objects. Nobody chose these; they are what the concept index makes visible.

What links here

Every essay whose body links to this one.

The objects this essay names

Each one links to every other essay that touches it.

AggregationAreaCell systemClippingConfoundingConservationDiscrete global gridInterpolationRebinningResolutionSpherical polygonValidation