Edge filaments follow the magnetic field. The same n≠0 structure seen on the outboard-midplane XZ plane appears on the divertor-target XZ planes at the same instant, displaced toroidally by the field-line map zShift(x). Three tabs: the planes and the C(Δt,Δz) evidence; the 3-D torus with the filaments, the planes and the field lines; and the prediction atlas with every target, model and lag, plus the held-out movie of truth beside prediction.

















Question (Ben Zhu, 2026-08-28). Edge filaments are field-aligned tubes; a single divertor pixel loses them, a whole plane should not. Extract complete radial–toroidal planes at the outboard midplane (OMP) and at the two divertor targets, find whether and how they are connected, and then predict the target plane from the midplane plane.
Findings. (1) The planes are connected only through the equilibrium field-line map: a rigid toroidal shift finds nothing (C ≈ 0.03), rolling each radial row by its own displacement zShift(x, ytarget) − zShift(x, yOMP) gives 0.4–0.7 for Te and 0.1–0.3 for Ne on the inner target (n≠0, SOL rows), with the wrong-sign map at the noise floor and adjacent rows at 0.98. (2) The match is at zero lag in both cadences and at every one of the 32 planes along the leg, while the toroidally averaged density and momentum arrive 150–310 µs later on the inner leg: the pattern is coordinated along the rope by the electrons, the material drains along it at the ion sound speed. (3) The connection is carried by structures wider than ~5 cm (n ≲ 20); blob-scale structures are not coherent between the planes at any lag. (4) The inner target inherits more than the outer, and the outer strike-point rows, which carry most of the fluctuating heat flux, are the least connected. (5) A toroidally equivariant per-mode linear map predicts the target planes to 0.5–0.9 and beats every neural model tried, including a geometry-aware operator transformer in a pre-registered pilot (0 of 8 wins); with the full midplane plane as input, that linear transfer is at the ceiling set by upstream information. (6) Three times more data leaves the correlations unchanged and the predictors nearly so.
Data. Four Hermes-3/BOUT++ runs of TCV shots 85604 and 85606 (lower single null): the older public runs (624 frames, 3.13 µs per frame) and Yichen Fu's fit_profile+5 runs (765 / 800 frames, 2.61 µs); the fit_profile+4 segments join the +5 ones with a one-frame gap on the same refined grid, giving 1936 / 2726-frame corpora. Interior grid 64 × 32 × 81, separatrix at x = 16, inner target y = 0, outer target y = 31, OMP = row of largest R at the separatrix, 1/5 toroidal wedge. Time from Ωci = eB/mp = 9.58 × 10⁷ s⁻¹. All alignments use the run's own zShift; all splits are chronological with a 30-frame gap; shot 85606 is never used for model selection.
A pressure perturbation at the edge is polarised by the curvature and ∇B drifts, and the resulting E×B flow convects it radially outward; because charged particles stream freely along B, the perturbation is a tube tens of metres long, not a blob (Krasheninnikov, D'Ippolito & Myra 2008). Following one field line from the midplane, the tube winds several toroidal turns and passes near the X-point, where the pitch changes fastest with radius and the tube is sheared into a ribbon. The plate therefore sees a sheared, rotated image of the midplane, and the per-row field-line map is what undoes the shear.
The tube carries two kinds of information at two speeds. Temperature and potential are set by the electrons, whose parallel conduction is faster than one frame, so the temperature pattern is established along the whole tube at once: the n≠0 correlation is highest for Te and peaks at zero lag at every plane. Density and parallel momentum belong to the ions and drain toward the plates at the sound speed, so their toroidally averaged part arrives 150–310 µs later. The distance–lag surface shows both: a flat ridge for the pattern and for mean Te and φ, a sloping ridge along the sound-speed line for mean Ne, Vi and NVi. One sharpening of the standard story: the potential's mean is coordinated along the tube but its fluctuating pattern is not shared with the plate at any lag (0.05), so the plate's potential fluctuations are set locally by the sheath and the divertor's own dynamics. "Thermally coordinated" is the accurate phrase.
The same picture appears in experiment: NSTX midplane–divertor correlations of 0.7–0.8 at delays inside one 11 µs frame against ion transit times of 50–100 µs, explained by fast potential propagation and electron conduction (Maqueda & Stotler 2010); far-SOL correlation up to 0.7 decreasing toward the separatrix (Scotti et al. 2020); two filament populations on TCV, elongated far-SOL filaments that trace to the midplane and small circular ones born in the divertor, with not all upstream filaments surviving the X-point shear (Wüthrich et al. 2022); and divertor-volume gradients that drive local turbulence, interchange on the inner leg and drift-wave on the outer (Walkden et al. 2022).



Question. Given the whole outboard-midplane plane at one instant, does the divertor plane depend on it in a nonlinear, state-dependent way that a fixed per-mode linear transfer cannot capture?
Setup. Input: Ne and Te on the OMP SOL plane (48 × 81) at time t, nothing else. Targets: Ne and Te on the inner and outer plate planes at the same t, each rolled row by row by the run's own field-line map (zShift). Standardised with training statistics only. Frozen chronological split: frames 0–405 train, a 30-frame gap, frames 435–623 held out and scored as two chronological halves. Identical random toroidal rolls on input and both outputs during training; one seed; 400 epochs, 95 min on a Rusty A100; 85606 untouched.
Model A, ridge: one complex linear map per toroidal mode number from the midplane radial profiles to each target radial profile, penalised least squares, penalty chosen per mode on an inner validation slice. Model B, GAOT-lite (2.25 M parameters): a multiscale neighbourhood-attention encoder from the source cells to a 12 × 27 latent grid using only relative positions (radial offset, wrapped toroidal offset as sin/cos), a 6-layer transformer on the latent grid with a learned wrapped relative-position bias and no absolute embedding, and a neighbourhood-attention decoder to the target cells whose queries carry the row's downstream geometry (field-line displacement as a wedge-periodic phase, connection length, flux-expansion ratio, inner/outer id). It differs from the ViT of the ablation in exactly those respects.
Scoring. Pooled correlation per target and segment; paired block bootstrap (20 time blocks, 300 draws) of the difference GAOT − ridge; correlation and explained variance per mode; correlation per radial row; a shuffled-input null (input frames circularly shifted by 150 frames); toroidal-equivariance error (roll the input, compare with the rolled output). Pre-registered rule: GAOT must beat ridge outside the bootstrap band on at least two chronological segments, otherwise spatial architecture iteration stops.
| run | segment | target | ridge | GAOT | GAOT − ridge [16 %, 84 %] | n≠0 ridge / GAOT |
|---|---|---|---|---|---|---|
| old 85604 | early half | inner Ne | 0.599 | 0.501 | -0.097 [-0.121, -0.071] | 0.753 / 0.577 |
| old 85604 | early half | inner Te | 0.919 | 0.834 | -0.085 [-0.094, -0.076] | 0.886 / 0.808 |
| old 85604 | early half | outer Ne | 0.613 | 0.362 | -0.251 [-0.279, -0.222] | 0.596 / 0.357 |
| old 85604 | early half | outer Te | 0.921 | 0.821 | -0.099 [-0.109, -0.089] | 0.868 / 0.748 |
| old 85604 | late half | inner Ne | 0.656 | 0.503 | -0.154 [-0.177, -0.129] | 0.689 / 0.519 |
| old 85604 | late half | inner Te | 0.939 | 0.885 | -0.055 [-0.058, -0.051] | 0.859 / 0.776 |
| old 85604 | late half | outer Ne | 0.499 | 0.189 | -0.305 [-0.347, -0.268] | 0.372 / 0.190 |
| old 85604 | late half | outer Te | 0.920 | 0.859 | -0.063 [-0.069, -0.056] | 0.723 / 0.590 |
| full 85604 | early half | inner Ne | 0.631 | 0.548 | -0.085 [-0.097, -0.072] | 0.561 / 0.420 |
| full 85604 | early half | inner Te | 0.888 | 0.848 | -0.040 [-0.046, -0.036] | 0.732 / 0.659 |
| full 85604 | early half | outer Ne | 0.534 | 0.224 | -0.312 [-0.329, -0.297] | 0.488 / 0.107 |
| full 85604 | early half | outer Te | 0.821 | 0.738 | -0.084 [-0.094, -0.076] | 0.516 / 0.261 |
| full 85604 | late half | inner Ne | 0.770 | 0.652 | -0.119 [-0.129, -0.107] | 0.586 / 0.429 |
| full 85604 | late half | inner Te | 0.923 | 0.890 | -0.034 [-0.040, -0.029] | 0.724 / 0.636 |
| full 85604 | late half | outer Ne | 0.536 | 0.234 | -0.303 [-0.322, -0.287] | 0.477 / 0.078 |
| full 85604 | late half | outer Te | 0.864 | 0.808 | -0.058 [-0.071, -0.047] | 0.482 / 0.266 |
Controls. Shuffled-input null (ridge / GAOT): inner Ne 0.03 / -0.03 · inner Te 0.09 / 0.08 · outer Ne 0.14 / 0.05 · outer Te 0.15 / 0.12. Equivariance error: ridge 0.000, GAOT 0.230.
Decision. GAOT beats ridge outside the bootstrap band in 0 of 16 cells; ridge beats GAOT in all of them. The same comparison on the full 85604 corpus (1258 training frames, 2.6 × more than the old run) gives the same answer in every cell; GAOT reached epoch 192 of 400 before the 4-hour wall budget on the A100, at a training loss still falling slowly, so its numbers there are a lower bound on a fully trained model, and its equivariance error was 0.136. The hypothesis is not supported: with the full midplane plane as input, the fixed per-mode linear transfer is at the ceiling set by upstream information, and the missing variance at the outer strike point is not recoverable by a nonlinear operator of this class at this data size. Per mode, GAOT's explained variance goes negative above n ≈ 25, where ridge correctly predicts nothing. Spatial architecture iteration stops; the residual is characterised next, and the diagnostic-view experiment uses ridge as its baseline. The full-corpus repeat runs as confirmation.




Inputs: the three outboard-midplane fields on the SOL plane, uf(x, z, t) for f ∈ {Ne, Te, φ}, x = 16…63 (48 radial rows), z = 0…80 (81 toroidal cells of the 1/5 wedge), each standardised on the training block. Targets: yg(x, z, t+τ) on a divertor plane for g ∈ {Ne, Te, q} × {inner, outer}, standardised the same way. All fits use the first 65 % of frames; scores are on the last block after a 30-frame gap.
Scores: corr all = Pearson correlation of predicted and true fluctuation planes pooled over the test block; corr n≠0 = the same after removing the instantaneous toroidal mean from both; corr n=0 = the same on the toroidal means only; MSE skill = 1 − MSE / variance of the truth (0 = climatology); total-load corr = correlation of the plane-integrated time series; shift-aligned frame corr = mean per-frame correlation after the best rigid toroidal shift of the prediction.
The question. Given the outboard-midplane plane at time t, how much of each divertor-target plane at time t + τ is determined by it, which model class extracts it, and where on the plate the rest comes from. This is Ben's experiment 1 (whole plane in, whole plane out) and, at the end, the first pass of experiment 2 (a camera-sized window in).
How it is scored. Every model is fitted on the first 65 % of a run, a 30-frame gap is discarded, and the remaining frames are scored without ever being seen. Two numbers are reported for everything: the correlation between predicted and true planes (is the pattern right, regardless of amplitude) and the RMSE skill 1 − MSE/Var (is the error smaller than predicting the mean; 0 = predicting the mean, negative = worse). The n = 0 / n≠0 split separates the toroidally averaged profile from the fluctuating pattern.
The models, in one line each. Per-mode linear (ridge): one linear map per toroidal mode number, symmetry built in. PCA ridge: generic linear map on 256 principal components. Field-line identity: the aligned midplane plane times a gain per row, no learning. Downstream persistence / self-ridge: the plate's own past, the bar for what upstream adds. CNNs, U-Net, ViT, GAOT: the neural challengers. Full definitions at the bottom.
The answer. The per-mode linear map is the best model on every run, target and lag; temperature planes are predicted to 0.8–0.95, density to 0.5–0.8, the outer heat-flux plane worst; upstream beats the plate's own memory after about two frames; no neural model, including a geometry-aware operator transformer in a pre-registered pilot, beats the linear map; and a single fixed-angle window recovers most of what the whole plane can predict.