All posts
Optimisation·6 min read

2 km UK: what it would buy, and what it costs

4 km resolves the mesoscale. 2 km starts resolving the terrain that actually triggers thermals. It also costs roughly eight times the compute. Here is the trade, honestly, before any of it has been benchmarked on the live box.

4 km is the workhorse resolution for Convek. It is the point where a regional model starts producing the mesoscale features a global model cannot - sea breezes, convergence lines, valley flows - at a compute cost that fits four cycles a day on a single worker. Everything live today, six countries' worth, runs at 4 km. The obvious next rung is 2 km, and it is worth being clear about what that actually buys and what it costs before treating it as a done deal.

Start with the cost, because it is the part that decides everything else. Halving the grid spacing makes a WRF run roughly eight times slower: four times the columns, and a time step that has to halve to stay numerically stable. The live trimmed UK 4 km cycle lands in around 34 minutes end-to-end on the `raspuk` worker. A 2 km version of the same domain is estimated at four to five hours per cycle on a bigger box - and that estimate is a planning number from the compute scaling, not something measured on the live hardware yet. Four to five hours does not fit a four-cycle-a-day schedule on the current worker, full stop.

So 2 km is not a namelist change, it is an infrastructure decision. It needs a second, heavier worker - the plan is a VPS6-class box with more vCPU and memory - dedicated to the high-resolution domain, running alongside the existing 4 km batch rather than competing with it for the cycle window. And it needs `gb/4km` to be genuinely rock-solid first, because there is no sense building a more expensive product on top of a base layer that still occasionally needs hand-holding. The order is: stabilise 4 km, stand up the second worker, benchmark 2 km for real, then decide.

What it would buy pilots is terrain. At 4 km, a grid cell is about 16 km² - good enough for the sea breeze and the broad convergence line, but it still smears the things that happen at the scale of a single ridge. Ridge convergence zones, the trigger points on a south-facing face that fire an hour before the valley, the sharpness of a sea-breeze front rather than just its presence, valley wind detail in the Welsh and Lake District terrain. 2 km starts to resolve those. Cloudbase also tightens up, because the moisture profile is sampled on a finer grid where the terrain forcing is sharper.

On horizon, the 4 km product now targets a 72 hour window. The 2 km question is whether a new box can hold that same roughly three-day window without crowding the 4 km production batch. That is a runtime question, not a physics one, and it is downstream of the benchmark that has not been run.

The honest status, then: 2 km UK is planned, not in development, and not benchmarked on the live infrastructure. The numbers in this post are scaling estimates. When the second worker exists and the first real 2 km cycle runs, this post gets rewritten with measured wall-clock and a side-by-side against the 4 km output - because the only thing that justifies eight times the compute is evidence that the finer grid is actually a better forecast, not just a prettier one. When it ships it will be a Pro-tier resolution alongside the free 4 km base.

The discipline here is the same one that runs through the rest of this blog: ship the cheap thing solid, validate it, and only buy the expensive upgrade once there is a number to chase. 2 km is the upgrade. The validation pipeline and a stable 4 km base are the prerequisites. The optimisation series covers the rest of the stack, and the model page has the current shipping configuration.