Paths and directions

An ensemble carries its own margin, if its spread is honest

A fixed margin of about eight kilometres for every hundred of forecast error covers the diversion floor in nineteen crossings of twenty, but it needs the error's size and it is spent on every crossing alike. Nineteen forecasts of the same crossing need neither. If the truth is one more draw from the spread they come from, the highest of them covers the floor met in nineteen crossings in twenty by rank alone — whatever the error, the streak or the drift — and it holds a fifth to two fifths less in hand than the fixed margin, because it spends nothing on the crossings the streak cannot reach. What it cannot do is notice that its own spread is too small: at half the true spread it covers 88 in a hundred.

Assumes A forecast's error in where the streak is becomes a margin.

A forecast’s error in where the streak is becomes a margin priced a jet streak’s forecast in kilometres of diversion allowance. A streak drifting along the Goose Bay–Narsarsuaq gap moves the floor for the New York–London chain by up to forty kilometres during a crossing, and an operator knows where the streak is only from a forecast. Planned to the worst floor over the forecast’s own crossing window, the plan needs a margin of about eight kilometres for every hundred of error in the streak’s position to cover the floor met in nineteen crossings of twenty, and an error in the drift speed costs a margin of its own.

That answer has two costs the essay stated and did not remove. The margin needs the forecast’s error to be known — a number no single forecast reports about itself. And it is spent on every crossing alike, including the many in which the streak cannot reach the gap while the aircraft is over water, where no margin is needed at all.

It ended by asking whether an ensemble could do better: many forecasts of the same crossing, each started from slightly different conditions, whose spread of window floors might set the margin case by case. That can be measured on the same table of exact floors, and the answer turns on one property of the ensemble that it cannot check for itself.

Nineteen forecasts of one crossing, and the floor each says must be met. One crossing, with the forecast's error in where the streak starts at 150 km and in its drift speed at 50 km/h. The line is the worst floor over a four-hour crossing against where the streak starts, at the assumed drift of 250 km/h. Dots: nineteen ensemble members, each drawn with its own start and drift speed and placed at the worst floor its own crossing window meets. Open circle: the truth, 660.3 km. Square: the ensemble's centre, 640.3 km. Dashed: the centre's floor plus the fixed margin that covers 95 per cent of crossings at these errors, 653.3 km. Solid: the highest member, 663.2 km — the plan the ensemble gives with no margin added.
Fig. 1 One crossing, with the forecast’s error at 150 km in where the streak starts and 50 km/h in its drift speed. Line: the worst floor over a four-hour crossing against the streak’s starting position, at the assumed drift of 250 km/h. Dots: nineteen members, each at the worst floor its own start and drift speed give. Open circle: the truth, 660.3 km. Square: the ensemble’s centre, 640.3 km. Dashed: the centre’s floor plus the fixed margin that covers 95 per cent of crossings at these errors, 653.3 km. Solid: the highest member, 663.2 km, which covers the truth; here the fixed margin does not.

Nineteen forecasts of one crossing

The table is the one the earlier essay built: the floor computed exactly at 35 positions of the streak’s centre along the jet, 125 kilometres apart, and read between them, so that a crossing costs a few table reads rather than a shortest-path search. Each floor is the least still-air radius at which the chain of diversion circles, each moved upwind by the wind, keeps the route within reach of an airfield. The streak is 200 km/h faster than a 100 km/h core and a thousand kilometres long; the crossing lasts four hours; the floor to be met is the worst over the stretch of positions the streak drifts through in that time.

An ensemble is modelled in the simplest way that is honest about what one is. Each crossing has a centre — the ensemble’s best estimate of where the streak starts, drawn uniformly over the 3,100 kilometres of jet from which the drift can bring the streak over the gap — and the truth lies away from it by a normal error in starting position and another in drift speed. Each of the nineteen members is drawn the same way, independently. That is what it means for an ensemble to be calibrated: the truth is statistically one more member.

Each member then gives a window floor of its own, and the plan is the highest of them. No margin is added and no error is stated. Against it stands the earlier essay’s plan, in the form that serves it best: the centre’s own window floor plus the fixed margin that covers exactly 95 per cent of these same crossings — a margin sized with the benefit of hindsight, which no operator has.

The crossing drawn above shows the difference in one case. The centre sits 2,300 kilometres short of the gap, where the floor is climbing steeply, and its window floor is 640.3 kilometres. The truth started further on and meets 660.3. One of the nineteen members meets more, 663.2, so the highest member covers the truth; the fixed margin, 13.0 kilometres at these errors and the same for every crossing, takes the plan to 653.3 and does not.

Nineteen members cover nineteen crossings in twenty

The highest of M calibrated members covers M in M + 1 crossings, whatever the error. The share of crossings in which the highest member's floor is at least the floor met, for ensembles of 4, 9, 19, 39 and 99 members, 12,000 seeded crossings each, with a position error of 100 km. Measured: 4 members 81.6 per cent, 9 members 91.1 per cent, 19 members 95.3 per cent, 39 members 98.0 per cent, 99 members 99.3 per cent. The line is M / (M + 1), the share a calibrated ensemble's rank argument gives; the measured shares sit on or just above it, because in a crossing the streak cannot reach, every member and the truth meet the same floor and a tie counts as covered. The dotted line is the 95 per cent the fixed margin is sized for.
Fig. 2 The share of crossings in which the highest member’s floor is at least the floor met, for ensembles of 4, 9, 19, 39 and 99 members, 12,000 seeded crossings each, with a position error of 100 km: 81.6, 91.1, 95.3, 98.0 and 99.3 per cent. The line is M / (M + 1); the measured shares sit on or just above it, because in a crossing the streak cannot reach, every member and the truth meet the same floor and a tie counts as covered. Dotted: the 95 per cent the fixed margin is sized for.

The single case is an anecdote. The rule behind it is not, and it is short enough to state whole. If the truth and the members are independent draws from the same spread, then among the twenty window floors — nineteen members and the truth — every order is equally likely, and the truth is the highest of the twenty with probability one in twenty. In every other case at least one member is at or above it. So the highest of nineteen members covers the floor met in nineteen crossings in twenty.

Nothing in that argument mentions the size of the error, the shape of the floor against position, the speed of the drift or the length of the crossing. It needs only that the truth is exchangeable with the members. The trials bear it out at every size of ensemble tried: four members cover 81.6 per cent against the 80 the rule gives, nine cover 91.1 against 90, nineteen 95.3 against 95, thirty-nine 98.0 against 97.5 and ninety-nine 99.3 against 99. The measured shares run slightly above the rule for a reason the figure’s caption gives: in a crossing where the streak cannot reach the gap in time, every member and the truth meet exactly the same floor, and a tie is covered.

With more members than the coverage needs, the plan need not be the highest. At thirty-nine members the member ranked ninety-fifth in a hundred covers 95.3 per cent and holds 5.4 kilometres in hand on average, less than the highest of nineteen; at ninety-nine members, 95.6 per cent and 5.4 kilometres. The count of members sets which member to plan to, and nineteen is simply the smallest count at which the highest member is the ninety-fifth-percentile one.

The margin goes where the streak can reach

The ensemble spends its margin where the streak can reach the gap, and none where it cannot. The highest member's floor less the ensemble centre's window floor, over 12,000 seeded crossings with a position error of 100 km and nineteen members. In 47.2 per cent of crossings it is within a kilometre of nothing — in 30.4 per cent because every member meets the same floor, the streak being unable to reach the gap during the crossing, and in the rest because no member's window is worse than the centre's. In 26.0 per cent it exceeds the fixed margin of 7.5 km (dashed), which is added to every crossing alike. Its largest is 47.7 km. On average the ensemble adds 5.5 km.
Fig. 3 The highest member’s floor less the ensemble centre’s window floor, over 12,000 seeded crossings with a position error of 100 km and nineteen members. In 47.2 per cent of crossings it is within a kilometre of nothing — in 30 per cent because every member meets the same floor, and in the rest because no member’s window is worse than the centre’s. In 26 per cent it exceeds the fixed margin (dashed), which is added to every crossing alike. Its largest is 47.7 km.

The two plans cover equally often and spend their margins in completely different ways. The fixed margin is 7.5 kilometres at a position error of a hundred — here, against 7.9 in the earlier trials, which drew the truth rather than the centre uniformly along the jet and used another seed — and every crossing carries it. The ensemble’s margin, the highest member less the centre’s own window, is nothing at all in nearly half of all crossings. In three crossings of ten every member meets the same floor, because the streak is too far from the gap to arrive during the crossing whatever the error, and in another sixth of them no member’s window is worse than the centre’s. In a quarter of crossings the ensemble’s margin is larger than the fixed margin, reaching forty-seven kilometres in the worst, and those are the crossings in which the streak is on the steep part of the curve and the members disagree about how far up it the truth will be.

That is exactly the distribution a margin ought to have. A streak along the jet moves the floor while the crossing is flown found the floor flat far from the gap and steep over a few hundred kilometres of the streak’s path, and an error in position is worth nothing where the floor is flat and a great deal where it is steep. A fixed margin cannot know which of those a crossing is. An ensemble knows, because its members sit on the curve.

A fifth to two fifths less in hand

At the same coverage, the highest of nineteen members holds a fifth to two fifths less in hand. What each plan holds in hand on average beyond the floor met, 12,000 seeded crossings at each position error. Dashed: the ensemble centre's window plus the fixed margin that covers 95 per cent, 3.9, 7.5, 10.9, 14.2, 20.7, 26.6 km at errors of 50, 100, 150, 200, 300, 400. Solid: the highest of nineteen members, covering 95.1 to 95.4 per cent, 3.0, 5.6, 7.7, 9.5, 12.8, 15.5 km. At 100 km of error the ensemble holds 5.6 against 7.5; at 400, 15.5 against 26.6. Dotted: ignoring the streak and carrying the margin that covers 95 per cent without it, 17.8 km held in hand at every error.
Fig. 4 What each plan holds in hand on average beyond the floor met, 12,000 seeded crossings at each position error. Dashed: the ensemble centre’s window plus the fixed margin that covers 95 per cent, 3.9, 7.5, 10.9, 14.2, 20.7 and 26.6 km at errors of 50, 100, 150, 200, 300 and 400. Solid: the highest of nineteen members, covering 95.1 to 95.4 per cent, 3.0, 5.6, 7.7, 9.5, 12.8 and 15.5 km. Dotted: ignoring the streak with the margin that covers 95 per cent without it, about 18 km at every error.

The saving is the allowance held in hand: the planned floor less the floor actually met, averaged over every crossing. It is what a margin costs in the crossings that did not need it, and at a position error of a hundred kilometres the fixed margin holds 7.5 kilometres and the highest of nineteen members 5.6, a quarter less, at the same coverage. The saving grows with the error: 23 per cent at an error of fifty kilometres, 33 at two hundred, 42 at four hundred, where the fixed margin holds 26.6 kilometres and the ensemble 15.5.

The earlier essay found that reading the forecast at all paid for itself, against simply ignoring the streak and carrying enough allowance for the worst it could do, only while the forecast’s error stayed under about two hundred and fifty kilometres. In these trials the plan that ignores the streak holds between 17.8 and 18.5 kilometres in hand, the dotted line in the figure. Read through an ensemble, the plan holds 12.8 kilometres at an error of three hundred and 15.5 at four hundred — still below it. The break-even moves out past the largest error measured, not because the forecast is better but because its error is spent where it matters.

And the ensemble did this with no error model. The fixed margin in the comparison was fitted to the very crossings it is scored on, which gives it the best case it can have; an operator using the earlier essay’s rate of eight kilometres per hundred would first have to know the error, and would be wrong about it by however much the forecast’s real skill differed from the one assumed. The ensemble’s plan uses nothing but its own members.

A wrong drift speed is carried by the members

An error in the drift speed widens the ensemble as it widens the fixed margin, and the ensemble keeps its lead. With the position error at 100 km and the drift speed also wrong by a normal error of the stated size, what each plan holds in hand beyond the floor met, 12,000 seeded crossings each. Dashed: the centre's window plus the fixed margin, recomputed at each drift error to cover 95 per cent — 7.5, 8.8, 11.2, 13.4, 15.1 km at 0, 25, 50, 75, 100 km/h. Solid: the highest of nineteen members, whose own drift speeds carry the error — 5.6, 6.2, 7.1, 7.9, 8.6 km, covering 95.2 to 95.6 per cent.
Fig. 5 With the position error at 100 km and the drift speed also wrong by a normal error of the stated size, what each plan holds in hand, 12,000 seeded crossings each. Dashed: the centre’s window plus a fixed margin recomputed at each drift error to cover 95 per cent — 7.5, 8.8, 11.2, 13.4 and 15.1 km at 0, 25, 50, 75 and 100 km/h. Solid: the highest of nineteen members, whose own drift speeds carry the error — 5.6, 6.2, 7.1, 7.9 and 8.6 km, covering 95.2 to 95.6 per cent.

The earlier essay found that an error in the drift speed is the same kind of error as one in position, because both move where the streak is at each moment of the crossing, and an error of 50 km/h moves the streak’s far end two hundred kilometres over four hours. For a fixed margin that means a second number to know and a larger margin to carry: with the drift speed wrong by 50 km/h on top of a hundred kilometres of position error, the margin rises from 7.5 kilometres to 10.7 and the allowance held from 7.5 to 11.2.

For the ensemble it means nothing new. Each member carries its own drift speed, so the members fan out further along the jet by the end of the crossing, and the highest of them rises only where that fan reaches a steeper part of the floor. At a drift error of 50 km/h the ensemble holds 7.1 kilometres, and at 100 km/h 8.6 against the fixed margin’s 15.1 — little more than half — while covering 95.6 per cent. The larger the part of the error that grows during the crossing, the more the ensemble saves, because a fixed margin has to be sized for the error at the end of the window and the ensemble spends it only where the end of the window matters.

The one thing the ensemble cannot report about itself

An ensemble that spreads too little covers too few crossings, and nothing inside it says so. The share of crossings the highest of nineteen members covers when the members are drawn with a stated fraction of the real error's spread — in starting position (100 km) and drift speed (50 km/h) alike — 12,000 seeded crossings each: 0.5 of the spread, 88.3 per cent holding 4.6 km; 0.67 of the spread, 91.7 per cent holding 5.6 km; 0.8 of the spread, 93.7 per cent holding 6.3 km; 1 of the spread, 95.5 per cent holding 7.1 km; 1.25 of the spread, 96.4 per cent holding 8.1 km; 1.5 of the spread, 97.0 per cent holding 8.8 km. The dotted line is the 95 per cent a calibrated ensemble of nineteen gives. Half the spread covers 88.3 per cent.
Fig. 6 The share of crossings the highest of nineteen members covers when the members are drawn with a stated fraction of the real error’s spread — 100 km in position and 50 km/h in drift speed, both scaled alike — 12,000 seeded crossings each: at half the spread 88.3 per cent holding 4.6 km; at two thirds 91.7 and 5.6; at 0.8, 93.7 and 6.3; calibrated 95.5 and 7.1; at 1.25, 96.4 and 8.1; at 1.5, 97.0 and 8.8. Dotted: the 95 per cent a calibrated ensemble of nineteen gives. The fixed margin holds 11.2 km at these errors.

Everything above rests on the truth being one more member, and that is a claim about the ensemble’s spread: that its members scatter as widely as the forecast’s real errors do. An ensemble whose members scatter too little puts the truth above all of them more often than one time in twenty, and nothing inside a single crossing can say so — the members agree with each other, which is exactly what makes them look trustworthy.

The trials put a number on it. With the members drawn at four fifths of the real spread, in position and drift speed alike, the highest of nineteen covers 93.7 per cent of crossings instead of 95. At two thirds it covers 91.7, and at half 88.3 — one crossing in nine short of the floor met instead of one in twenty. An ensemble that spreads too much errs the safe way: at one and a half times the spread it covers 97.0 per cent, and even then holds 8.8 kilometres in hand against the fixed margin’s 11.2.

So the ensemble’s plan trades one unknown for another. The fixed margin needed the size of the forecast’s error; the ensemble needs the ratio of its spread to that error, and a well-made ensemble is designed to make that ratio one. Whether it is can be checked, but only across many crossings: if the ensemble is calibrated, the truth’s rank among the members is uniform, so over a season the truth should be the highest in about one crossing in twenty, the lowest in one in twenty, and so on. Weather centres keep exactly that count for their ensembles — the rank histogram, introduced for this purpose in the mid-1990s — and its shape is the verification an operator would need before trusting the highest member. A histogram piled at its two ends says the spread is too small, and the trials above say how much coverage that costs.

The allowance was always a quantile

It is worth being clear about what has and has not changed from the earlier plan. The rule that keeps a route near land is pinned at both ends set out why the floor is a wall rather than a target: a plan that meets it on average is on the wrong side of it in half the crossings, so any forecast has to become an allowance for the bad cases. Both plans here are that. The fixed margin reaches the ninety-fifth percentile of the error by adding a constant; the ensemble reaches it by taking an order statistic of the members. The coverage is the same by construction. What differs is how evenly the allowance is spread across crossings, and an allowance spread according to where the streak actually is costs less for the same protection. Neither is rescued by planning again in the air: a crossing is a chain of decisions found the diversion allowance fixed before departure, so whatever the forecast gets wrong is paid for in the plan, and the ensemble changes only how the payment is divided.

A week of twilights learns what a sight is worth met the same trade in a navigator’s sights: a scatter that is known is worth using, and a scatter nobody knows is often cheaper to allow for than to estimate. The ensemble is a way of knowing the scatter crossing by crossing, and tonight’s residuals measure tonight’s air badly and weight tonight’s fix well is the warning that goes with it: an estimate of spread made from a handful of draws on one night is itself noisy, and the ensemble’s is no exception. Nineteen members place the ninety-fifth percentile to within the sampling that the coverage figure already includes; they do not place it for a crossing whose members were drawn from the wrong spread.

What an operator can take from this

Plan to the highest of nineteen members, or the ninety-fifth percentile of more. With a calibrated ensemble it covers the floor met in nineteen crossings of twenty whatever the forecast’s error, and needs no error stated.

It holds a fifth to two fifths less in hand than a fixed margin at the same coverage, more as the error grows and more again when the error is partly in the drift speed — because it spends nothing where the streak cannot reach the gap.

Its one requirement is a spread that matches the forecast’s real error, and that can only be checked across many crossings, by how often the truth falls above all the members. An ensemble spreading four fifths as widely as it should covers 94 crossings in a hundred; one spreading half as widely, 88.

The fixed margin remains the fallback. Without an ensemble, or with one whose rank counts say it spreads too little, the flat rate of about eight kilometres per hundred of error is what can be defended — and a jet moves the floor that a uniform wind cannot is the reminder that any such rate belongs to one streak on one gap.

What each number was held to

The rank argument must hold where it is claimed to. With a calibrated ensemble the highest of M members must cover at least M/(M + 1) of crossings, less two sampling standard errors, at every count from four to ninety-nine. It does at each, and exceeds the rule slightly for the reason the ties give.

No error, no margin. With no forecast error at all every member is the truth, and the highest member must hold nothing in hand. It holds nothing.

The finding must be there to fail, both ways. At nineteen members and a position error of a hundred kilometres the highest member must cover at least as often as the fitted fixed margin while holding at least a fifth less (95.3 against 95.0 per cent, 5.6 against 7.5 kilometres). And an ensemble with half the true spread must cover under ninety per cent (88.3), or its calibration would not be worth stating.

Every plan is scored on the same crossings. One seeded stream of centres, truths and members at each setting, shared by both plans.

Where the ensemble stops

Members independent and alike. Real members are not independent draws: they share a model and its biases, and a bias common to every member — a streak that the model always moves too slowly — moves the truth outside the ensemble in the same direction every time, which no count of members repairs and the rank histogram shows as a slope rather than a pile at both ends.

Errors normal, in two numbers. The streak is placed by a start and a drift speed, and every error is in those — the same flow field the quickest route is not the shortest steered a single aircraft through, reduced to where its fastest air is. A real forecast can also get the streak’s strength or length wrong, and the floor’s movement scales with its strength, so an ensemble whose members differ in strength would spread its floors differently.

One streak, one gap, a stated coverage. Nineteen crossings in twenty is a choice, and the ensemble’s plan moves with it by the rank argument exactly: ninety-nine in a hundred needs ninety-nine members, or a correction for the tail that the members themselves cannot supply.

The members are drawn, not forecast. Nothing here runs a weather model; the ensemble is a stated spread about a stated centre, which is the ensemble as its designers intend it. How far a real one departs from that is the quantity the rank histogram measures, and nothing here assumes any particular centre’s answer.

Still open: whether a season of crossings can correct the spread

The one failure that matters is an ensemble spreading too little, and it has a standard repair: stretch the members about their centre by a factor learned from past crossings, chosen so that the truth’s rank comes out uniform. The trials already say what the factor is worth — at four fifths of the true spread, stretching by a quarter restores the coverage from 93.7 per cent to 95.

What they do not say is how many crossings it takes to learn the factor well enough, when each crossing contributes one rank and most crossings are ties because the streak never reached the gap. Whether a season of transatlantic crossings, most of them uninformative, pins the factor to within the tenth that separates 93.7 per cent from 95, and whether the factor learned in winter serves the summer’s weaker jets, are questions a calibrated ensemble cannot ask.

Named alongside this one

Essays reaching for the same objects. Nobody chose these; they are what the concept index makes visible.

The objects this essay names

Each one links to every other essay that touches it.

BottleneckConventionCoverageEstimatorFlow fieldMarginRankingToleranceVerification