Why wavelets
A three-level dual-tree complex wavelet pyramid separates scale and orientation while staying exactly invertible, so a forecast made in coefficient space maps back to a physical field with no loss.
Research series
A full research programme on forecasting tropical-Pacific sea-surface temperature as a field, in an exactly invertible complex-wavelet space — and on knowing, at every step, what the evidence is actually worth.
Eleven documents: a standalone publication, the charter that defines every metric and evidence tier, and nine reports tracing the work from a validated pipeline through two controlled studies that narrowed the design space, one pristine sealed test, and a probabilistic stack whose advantage only partly transferred. Each entry below opens its full report.
The theme
Tropical-Pacific sea-surface temperature is forecast as a field rather than an index, inside an exactly invertible complex-wavelet representation. The programme's real subject is what it takes to know whether such a system works: every report changes one factor, declares its gates in advance, and states which tier of evidence its claims are allowed to carry.
A three-level dual-tree complex wavelet pyramid separates scale and orientation while staying exactly invertible, so a forecast made in coefficient space maps back to a physical field with no loss.
Engineering checks, simulation results, repeatedly consulted validation, and one-shot sealed reads are never allowed into the same column. The tier labels are what stop development evidence being read as confirmation.
Start with the publication for the finished products, or the charter for the measurement rules. The nine numbered reports are the complete working record behind them, in the order the work was done.
Outside the arc
The publication reports the final products to an external audience; the charter governs how everything else is measured. Neither belongs to the chronology below.
The externally readable product of the whole programme: a CMIP6-pretrained deterministic forecaster and a 32-member residual-diffusion ensemble in the relative-SST frame, characterized on 2021–2025 development validation. Self-contained, and deliberately not ranked against the earlier generations.
The measurement contract: evidence tiers, the persistent metric identifiers, the house bootstrap, and the comparability rules. Read it first if you want to know exactly what any number in the series is allowed to mean.
The arc
Nine reports in reading order, each stating its incoming baseline, the single factor it changes, its predeclared gates, and the artifact it hands forward. Each entry opens the complete report, with its prose, figures, tables, and references preserved.
Builds the thing and proves it is built correctly: lossless coefficient packing, exact inverse reconstruction, auditable checkpoint surgery on the official backbone, and fail-closed access to the sealed period. The scope is the pipeline; skill claims come later.
Nine controlled families ask whether direct adaptation on 335 observed windows can forecast at all. It cannot: low error is bought with near-climatological damping, amplitude costs correlation, and residual learning spends itself cancelling persistence.
Masked coefficient pretraining over four SST products improves reconstruction monotonically and forecast transfer not at all. The durable products are a coefficient-energy explanation of why, and the statistical governance the rest of the series runs under.
Pretraining on the actual forecast task across eleven CMIP6 models transfers where reconstruction did not, and the resulting model spends the project's one pristine hash-gated read of the sealed period. This is the strongest confirmatory result in the series.
A controlled test of whether ocean heat content adds anything. Inside simulation it helps; through the observation pathway it resolvably hurts. The zeroed-input twin design is what makes that a causal statement rather than a guess.
Gaussian, latent-noise, and diffusion heads around the frozen center. Diffusion fixes the smudgy individual member — the symptom that started the program — but the regional calibration wall survives every change tried, including sixty-four times more data.
Six rungs of explicit calibration establish what post-processing can and cannot buy: average reliability is easy and can be lifted into sharp fields, but flow-dependent width does not transfer from simulation and cannot be manufactured after the fact.
Re-weighting the objective across wavelet bands is what unlocks temporal coherence — the joint sequence architecture alone does not reach it. The adopted stack is the best validation product of the series, and its one authorized test read shows the advantage only partly transferred.
The control the earlier reference finding was missing: five pipelines differing only in the anomaly reduction, each retrained from scratch. Retraining, not climatology updating, turns out to explain most of the degradation. Ends by freezing the deployment contract.