TCV emulator state of the union · definitions, evidence and open decisions

Research state of the union · 10 August 2026

TCV plasma emulator: what the present evidence supports

A self-contained account of the two-shot dataset, LOLA training procedures, representation choices, rollout calibration, ETKF sensor conditioning and the physics tests needed before diagnostic rankings can be interpreted as design guidance.

Shot 85604: training and validation windowsShot 85606: held-out rollout evaluationAll plotted quantities are defined belowNo cluster state was changed to build this report

Executive summary

  • The data resolve a one-fifth toroidal domain. The physical mode convention is n=5k, where k is the Fourier harmonic in the simulated wedge. Physical n=35 remains measurable and temporally coherent; the low-coherence tail begins near n≥80.
  • f8 is an effective aggregate-error baseline with a known spectral limit. Its 11-cell toroidal latent can represent at most k=5, or n=25. z44 reconstructs the n=35 edge component. z22 has enough nominal Nyquist capacity for n=35 but has neither demonstrated that reconstructed peak nor completed a matched held-out rollout gate.
  • Rollout-CRPS fine-tuning produced a modest checkpoint-level error improvement, not calibrated ensembles. Epoch 4 improves mean held-out free error by 6.8% and anchored error by 2.7% relative to the parent. Held-out free SSR increases from about 0.328 to 0.376, still far below the target value 1.
  • ETKF conditioning can reduce rollout error. φ-bearing idealized proxy suites rank strongly for the baseline f8 prior, but φ is gauge-dependent and the current channels are not complete diagnostic forward models. The rollout-CRPS checkpoint has not undergone the 11-suite ranking.
  • The emulator is not yet established as transport-faithful. Aggregate RMSE, spectral content, calibration, cross-field phase, radial flux and blob propagation are distinct requirements. The transport-sensitive battery remains mixed, so diagnostic rankings should be treated as model-conditional hypotheses.

01Dataset, cadence and toroidal modes

The study uses two simulations in one operating regime. Shot 85604 supplies training and same-shot validation windows; shot 85606 supplies a held-out temporal test. The stored frame interval is 3.1319 μs. This supports within-regime short-horizon tests, but it does not establish robustness across equilibria, forcing conditions or discharge regimes.

What the data can test

Short-horizon stochastic rollout skill; temporal transfer to shot 85606; codec fidelity; mode-by-mode predictability; and the response of an emulator ensemble to idealized sensor information.

What remains unidentified

Cross-regime generalization; real-instrument performance; uncertainty under distribution shift; and whether a sensor ranking persists across equilibria, noise levels and forward models.

Toroidal coordinates

The simulated interval is 2π/5 of the full torus and is repeated five times. Therefore only physical toroidal modes n=0,5,10,… occur:

n=Nperiodk=5k

For a field's wedge Fourier coefficient Fk(t,x,y), lag-1 complex coherence measures persistence of complex phase and amplitude direction between adjacent stored frames. The temporal mean is removed before the toroidal FFT:

γk=|t,x,yFk(t+1)Fk*(t)|t,x,y|Fk(t+1)|2t,x,y|Fk(t)|2

The non-axisymmetric power share pk is the energy in one wedge harmonic divided by the energy in all positive harmonics:

pk=t,x,y|Fk|2j1,t,x,y|Fj|2

Coherence and power answer different questions: a small mode can remain predictable, while an energetic mode can lose exact phase.

Toroidal domain and truth-only predictability. Panel A defines wedge harmonic k and physical full-torus mode n=5k. In panel B the blue curve is Ne lag-1 complex coherence and the orange curve is the percentage of Ne non-axisymmetric power in each mode; the dashed line marks n=35 and the shaded band starts at n=80. Panel C compares power in n=20–35 with the n≥80 tail for each field. Panel D shows each field's n=35 coherence and power share; point area also scales with power.
Figure 1. Toroidal domain and truth-only predictability. Panel A defines wedge harmonic k and physical full-torus mode n=5k. In panel B the blue curve is Ne lag-1 complex coherence and the orange curve is the percentage of Ne non-axisymmetric power in each mode; the dashed line marks n=35 and the shaded band starts at n=80. Panel C compares power in n=20–35 with the n≥80 tail for each field. Panel D shows each field's n=35 coherence and power share; point area also scales with power. Evidence: shot 85604 training HDF5; BOUT metadata ZMAX=0.2 and zperiod=5
fieldn=35 coherencen=35 powern=20–35 coherencen=20–35 powern≥80 coherencen≥80 power
Ne0.3252.42% 0.57664.59% 0.0080.35%
Te0.3161.53% 0.53954.33% 0.0110.18%
phi0.3111.65% 0.55744.47% 0.0080.28%
Vort0.1105.12% 0.28137.97% 0.00712.29%

The practical target suggested by these truth statistics is phase-sensitive prediction through approximately k=7–8 (n≈35–40), together with statistically faithful treatment of the low-coherence tail beginning around k=16 (n≈80). That tail cannot simply be discarded: vorticity retains 12.3% of its non-axisymmetric power there.

02Emulator pipeline and representation choices

Model pipeline and experimental branches. The top row is the LOLA path from a five-field state through an autoencoder, masked latent-diffusion sampler and decoder, followed by rollout composition or ETKF assimilation. Branch A changes the representation, branch B is a separate deterministic probabilistic-retrofit experiment, and branch C fine-tunes the f8 LOLA denoiser on sampled multi-block rollouts. CRPS is a scoring rule and does not replace diffusion sampling.
Figure 2. Model pipeline and experimental branches. The top row is the LOLA path from a five-field state through an autoencoder, masked latent-diffusion sampler and decoder, followed by rollout composition or ETKF assimilation. Branch A changes the representation, branch B is a separate deterministic probabilistic-retrofit experiment, and branch C fine-tunes the f8 LOLA denoiser on sampled multi-block rollouts. CRPS is a scoring rule and does not replace diffusion sampling. Evidence: external/lola and lola_ext experiment code

LOLA is a latent diffusion emulator. An autoencoder (AE) compresses a five-field 3D state into a smaller latent grid. A masked diffusion transformer samples future latent frames conditional on one or more known context frames. The decoder returns the sampled latent trajectory to field space. Independent Gaussian diffusion noise produces different ensemble members; in that precise sense, the model is still “rolling dice.”

The fields are electron density Ne, electron temperature Te, ion temperature Ti, electrostatic potential φ and parallel ion velocity Vi. f8, z22 and z44 are experiment labels for codecs with 11, 22 and 44 toroidal latent cells. Their direct real-FFT Nyquist ceilings are k=5, 11 and 22, or physical n=25, 55 and 110.

variantlatent z cellslatent Nyquist ceilingdominant reconstructed Ne edge modeAE reconstruction VRMSEheld-out rollout status
f811 k=5 / n=25 n=10 0.0366 held-out scored
f8-short11 k=5 / n=25 n=10 0.0366 held-out scored
z2222 k=11 / n=55 n=25 0.0322 not rollout-scored
z4444 k=22 / n=110 n=35 0.0392 held-out scored

AE reconstruction VRMSE is computed in physical field units for each field and frame, then averaged:

VRMSE=meanspace(x^x2)Varspace(x)+106

The “dominant reconstructed Ne edge mode” is the largest positive non-axisymmetric toroidal component in reconstructed edge density. It is an empirical content check, distinct from the Nyquist ceiling, which is only a mathematical capacity limit.

Representation and rollout comparison. In panel A each horizontal bar is the latent Nyquist ceiling and each dot is the dominant positive non-axisymmetric mode reconstructed in edge Ne; these are capacity and measured-content quantities, respectively. Panel B compares time-mean standardized RMSE for free rollouts (blue) and ETKF-anchored rollouts with inflation 2 (teal) on shot 85606. z22 is absent from panel B because no matching held-out rollout gate has been run.
Figure 3. Representation and rollout comparison. In panel A each horizontal bar is the latent Nyquist ceiling and each dot is the dominant positive non-axisymmetric mode reconstructed in edge Ne; these are capacity and measured-content quantities, respectively. Panel B compares time-mean standardized RMSE for free rollouts (blue) and ETKF-anchored rollouts with inflation 2 (teal) on shot 85606. z22 is absent from panel B because no matching held-out rollout gate has been run. Evidence: docs/frontier_points.json; jobs 6791660 and 6791711; gates 6791867, 6791868 and 6792464

03Two LOLA training procedures

Procedure A: masked-denoising continuation

Jobs 6791660 (z22) and 6791711 (f8 equal-budget control) continue training from an existing best checkpoint. “Continuation” means the denoiser weights begin from a pretrained run rather than random initialization; the script creates a new optimizer and learning-rate schedule. Each example contains five consecutive frames. A contiguous prefix of one to four frames is marked as context, and the model learns to denoise the remaining latent cells at a random diffusion time.

With clean latent trajectory x, Gaussian noise z, diffusion time t uniformly sampled on [0,1], and schedule coefficients αt and σt, the corrupted input and denoising objective are

xt=αtx+σtz,z𝒩(0,I)den=𝔼[meaniμθ(xt,t,m)x2stopgrad(vt,i)]

Here m is the context mask, μθ is the denoiser mean prediction, vt,i is the fixed schedule variance used as a weight, and stopgrad means that weight receives no gradient. Context cells are pinned to √(αt²+σt²)x. These runs use no CRPS or spectral term. Both use 40 continuation epochs, 1,024 training examples per epoch, batch 8, AdamW at 5×10−5 with cosine decay, and gradient clipping at 1.

Masked-denoising continuation histories. Blue is the arithmetic mean of 128 stochastic training-batch DenoiserLoss values in each continuation epoch. Teal is the mean over 32 shuffled validation batches, with newly sampled masks, diffusion times and Gaussian noise. The red point is the minimum validation loss. These curves measure the one-window denoising objective within a run; they are not rollout-error or physics-validity curves, and their absolute levels are not directly comparable across different latent grids.
Figure 4. Masked-denoising continuation histories. Blue is the arithmetic mean of 128 stochastic training-batch DenoiserLoss values in each continuation epoch. Teal is the mean over 32 shuffled validation batches, with newly sampled masks, diffusion times and Gaussian noise. The red point is the minimum validation loss. These curves measure the one-window denoising objective within a run; they are not rollout-error or physics-validity curves, and their absolute levels are not directly comparable across different latent grids. Evidence: Rusty configurations and histories for jobs 6791660 and 6791711
Figure 4 curveformal meaningappropriate interpretation
training mean DenoiserLossArithmetic mean of 128 stochastic batch losses in one epoch.Whether optimization is reducing the masked-denoising objective on sampled training windows.
validation mean DenoiserLossArithmetic mean of 32 shuffled validation-batch losses with fresh masks, t and z.A same-shot one-window generalization diagnostic; not a deterministic curve and not a rollout score.
red minimumEpoch with the smallest logged validation DenoiserLoss.The checkpoint saved as state_best under this selector.

Procedure B: sampled rollout-CRPS fine-tuning

Job 6791994 initializes the denoiser from the pretrained f8 state_best, freezes the autoencoder and updates the entire denoiser at a lower learning rate of 10−5. This is a fine-tune: it asks whether changing the training objective can reshape an existing sampler while holding the representation fixed. A from-scratch retrain would also change initialization and representation learning, making that narrow causal comparison harder.

Each training example samples M=4 diffusion members. Every member uses two composed five-frame windows with one-frame overlap: four new frames per block and eight predicted frames total. Each block uses 16 reverse-diffusion steps. Gradients pass through every reverse step and through the carried state between blocks. The 40 epochs contain 128 examples each, with batch 2 and four-step gradient accumulation for an effective batch of 8; toroidal-roll augmentation is enabled.

For predictions x(m), truth y, ensemble size M and α=0.95, define ε=(1−α)/M=0.0125. The index i below runs over scored latent cells:

A=1Mm=1Mmeani|xi(m)yi|D=2M(M1)m<jmeani|xi(m)xi(j)|afCRPSε=A12(1ε)DRCRPS=afCRPSε+0.3bandCRPSε

A rewards accuracy; D rewards ensemble separation; fair CRPS uses ε=0. bandCRPS applies a toroidal FFT, filters a band, divides by that truth band's detached RMS, evaluates afCRPS and averages bands. The requested edges [1,6,16] meet an f8 real-FFT grid containing only k=0,…,5, so the effective non-axisymmetric band is k∈[1,6). This objective cannot directly score k=7 (n=35) because that component is absent from the f8 latent.

Rollout-CRPS fine-tuning history. Panel A is the epoch-mean optimized objective afCRPS + 0.3×bandCRPS. Panels B–D use the same unaugmented validation slice and identical sampling seeds at every epoch: fair CRPS, latent ensemble-mean rollout RMSE, and aggregate validation spread-to-error ratio (SSR). Epoch −1 is the parent checkpoint. Vertical lines identify three checkpoint selectors: epoch 4 minimizes validation anchored RMSE, epoch 7 minimizes validation fair CRPS, and epoch 18 maximizes validation anchoring gain.
Figure 5. Rollout-CRPS fine-tuning history. Panel A is the epoch-mean optimized objective afCRPS + 0.3×bandCRPS. Panels B–D use the same unaugmented validation slice and identical sampling seeds at every epoch: fair CRPS, latent ensemble-mean rollout RMSE, and aggregate validation spread-to-error ratio (SSR). Epoch −1 is the parent checkpoint. Vertical lines identify three checkpoint selectors: epoch 4 minimizes validation anchored RMSE, epoch 7 minimizes validation fair CRPS, and epoch 18 maximizes validation anchoring gain. Evidence: job 6791994 history and checkpoint manifest
Figure 5 quantitydefinition and evaluation set
training objectiveEpoch mean of the optimized M=4, eight-frame ℒRCRPS on augmented training examples.
validation fair CRPSafCRPS with ε=0 on a fixed unaugmented shot-85604 validation slice, M=8 and identical sampling seeds each epoch.
validation rollout RMSESquare root of the mean squared difference between the M=8 latent ensemble mean and latent truth on the same fixed slice.
validation SSRMean validation-batch spread divided by mean validation-batch RMSE. This aggregate estimator is specific to Figure 5.

04Held-out rollout accuracy and calibration

All rollout errors below are measured in standardized model space: Ne is log-transformed and then standardized; Te, Ti, φ and Vi are standardized directly. Each standardized field and grid cell receives equal weight. These values are dimensionless and should not be read as physical-unit errors.

For an M-member ensemble at forecast frame t, the evaluator uses ensemble-mean error Et, ensemble spread St, spread-to-error ratio SSR and anchoring gain ρ:

Et=meancellsx¯y2St=meancellsVarm(x(m))SSR=meant=1:47StEt,ρ=mean(Etfree)mean(Etanchored)

SSR=1 means spread and ensemble-mean error have the same average magnitude under this estimator; SSR<1 is underdispersion. ρ>1 means anchoring lowers mean error. Figure 5 uses a different aggregate validation SSR estimator, so its ≈0.7 values are not a before/after pair with the ≈0.3–0.4 held-out framewise estimator.

Like-for-like held-out checkpoint gate on shot 85606. Panels A and B show time-mean standardized RMSE for free and ETKF-anchored rollouts. Each dot is one start (192, 288 or 384), and each translucent bar is their arithmetic mean. Panel C is the mean framewise free-rollout SSR; the dashed line at 1 denotes equality of ensemble spread and ensemble-mean error. All checkpoints use M=64 members, H=48 frames, inflation 2 and the same legacy 69-channel iter layout.
Figure 6. Like-for-like held-out checkpoint gate on shot 85606. Panels A and B show time-mean standardized RMSE for free and ETKF-anchored rollouts. Each dot is one start (192, 288 or 384), and each translucent bar is their arithmetic mean. Panel C is the mean framewise free-rollout SSR; the dashed line at 1 denotes equality of ensemble spread and ensemble-mean error. All checkpoints use M=64 members, H=48 frames, inflation 2 and the same legacy 69-channel iter layout. Evidence: jobs 6791995, 6791996 and 6792287
checkpointvalidation selection ruleheld-out free RMSEheld-out anchored RMSEfree change vs parentanchored change vs parent
parent f8pretrained state_best0.30240.2693referencereference
epoch 7minimum validation fair CRPS0.28800.2707+4.8%-0.5%
epoch 18maximum validation gain ρ0.30030.2744+0.7%-1.9%
epoch 4minimum validation anchored RMSE0.28180.2620+6.8%+2.7%

The epoch-4 result is modest but real under this gate: all three starts contribute to the means, and both free and anchored error improve relative to the parent. It is not the checkpoint chosen by minimum validation fair CRPS, and all evaluated checkpoints remain underdispersed. The training objective decreases from 0.2416 to 0.2246, while fixed-slice validation fair CRPS and validation SSR do not improve over the parent. The evidence therefore supports a useful early fine-tune checkpoint, but not a general calibration solution.

Field-level epoch-4 example on held-out shot 85606, start 192. Panel A compares free and ETKF-anchored whole-domain time-mean standardized RMSE; percentages are relative error reductions, so a negative value means anchoring increased error. Panel B compares anchored error and anchored ensemble spread on the identical edge support x≥16; S/E is their ratio. Values average forecast frames 1–47. This single start illustrates field behavior, while Figure 6 provides the three-start aggregate gate.
Figure 7. Field-level epoch-4 example on held-out shot 85606, start 192. Panel A compares free and ETKF-anchored whole-domain time-mean standardized RMSE; percentages are relative error reductions, so a negative value means anchoring increased error. Panel B compares anchored error and anchored ensemble spread on the identical edge support x≥16; S/E is their ratio. Values average forecast frames 1–47. This single start illustrates field behavior, while Figure 6 provides the three-start aggregate gate. Evidence: job 6792287, DALONG_anch_sel_s192_infl2.0/da_summary.json

05ETKF conditioning and sensor experiments

The ensemble transform Kalman filter (ETKF) is a data-assimilation method. At each analysis time it compares synthetic observations from each forecast member with an observed synthetic value. Ensemble anomalies estimate the covariance between measured channels and every latent-state direction. The observation innovation is then mapped into a state update. Multiplicative inflation expands forecast anomalies before the analysis; M is ensemble size and H is rollout horizon in frames.

In compact form, the ensemble-mean update has the Kalman structure

x¯a=x¯f+K(yh(x¯f))

Superscripts f and a denote forecast and analysis, h is the synthetic observation operator, y is the observation and K is estimated from the ensemble covariance and assumed observation error. Narrow or misoriented ensemble covariance can therefore produce an early update that raises global error even when later updates help. This behavior diagnoses covariance, localization, observation-error or forward-model mismatch; it does not by itself imply that early sensor samples should be removed.

Anchored rollout timing and ensemble-mean error. The dark-blue curve is the free rollout and the teal curve is the ETKF-anchored rollout; the horizontal axis is rollout frame and the vertical axis is ensemble-mean standardized error. Vertical guides mark ETKF analyses every four generated frames. One frame is 3.1319 μs, so analyses are ≈12.53 μs apart over the H=48 (≈150.3 μs) horizon. The configuration uses shot 85606, M=64, inflation 2 and the legacy 69-channel iter layout; ρ is mean free error divided by mean anchored error.
Figure 8. Anchored rollout timing and ensemble-mean error. The dark-blue curve is the free rollout and the teal curve is the ETKF-anchored rollout; the horizontal axis is rollout frame and the vertical axis is ensemble-mean standardized error. Vertical guides mark ETKF analyses every four generated frames. One frame is 3.1319 μs, so analyses are ≈12.53 μs apart over the H=48 (≈150.3 μs) horizon. The configuration uses shot 85606, M=64, inflation 2 and the legacy 69-channel iter layout; ρ is mean free error divided by mean anchored error. Evidence: docs/anchored_rollout_results/honest_6615586
Baseline f8 sensor-suite ranking. Each bar is the geometric mean anchoring gain ρ across three rollout starts; the overlaid ticks are individual starts. A bar above 1 means lower error after anchoring. gpi denotes idealized Ne samples on the GPI support, reflec an idealized Ne reflectometry chord, target the divertor-target support, and suffixes _phi and _te add direct φ or Te state proxies. This figure evaluates the baseline f8 prior with inflation 1; the rollout-CRPS checkpoint has not been evaluated across this full 11-suite set.
Figure 9. Baseline f8 sensor-suite ranking. Each bar is the geometric mean anchoring gain ρ across three rollout starts; the overlaid ticks are individual starts. A bar above 1 means lower error after anchoring. gpi denotes idealized Ne samples on the GPI support, reflec an idealized Ne reflectometry chord, target the divertor-target support, and suffixes _phi and _te add direct φ or Te state proxies. This figure evaluates the baseline f8 prior with inflation 1; the rollout-CRPS checkpoint has not been evaluated across this full 11-suite set. Evidence: jobs 6771840, 6771851 and 6771852; independent band-f8 replication 6790070–6790072

Figure 9 is a baseline model-conditional information ranking. The potential φ is gauge-dependent: replacing φ(x) with φ(x)+C leaves the electric field and E×B velocity unchanged because their dynamics depend on spatial gradients. A physical follow-up should use a declared gauge, assimilate φ anomalies or gradients, or implement a floating-potential forward model with realistic transfer and noise. The localization settings kmax=2 and 5 correspond to physical n≤10 and n≤25; a test that includes n=35 requires kmax≥7 and a representation capable of k=7.

06Physics validity beyond aggregate error

Low standardized RMSE is useful but insufficient for the research goal. A trustworthy edge/SOL emulator should also preserve the relationships that drive transport: spectral power in the coherent band, density–potential phase, radial E×B particle flux, intermittency and the direction and speed of localized structures.

Held-out edge-physics acceptance scorecard. Γn is the radial particle-flux proxy; its skewness sign tests the asymmetry of transport events. |cos αnv| is the transport projection from density–radial-velocity cross-phase. Qe,cond is conductive electron heat flux. Realizability requires Ti>0. Tail concentration is the fraction of |Γn| carried by the largest 1% of samples. λNe and λTe are midplane SOL e-folding widths. The final rows test the radial Ne-fluctuation skew transition, particle/energy step changes, end-to-start particle-count drift and whether blob radial velocity is identifiable. Green, amber, red and gray mean pass, partial, fail and open under the declared battery rules; they are not probabilities. The battery uses one sanctioned mid-toroidal rollout dump and does not provide a full-ensemble realizability test.
Figure 10. Held-out edge-physics acceptance scorecard. Γn is the radial particle-flux proxy; its skewness sign tests the asymmetry of transport events. |cos αnv| is the transport projection from density–radial-velocity cross-phase. Qe,cond is conductive electron heat flux. Realizability requires Ti>0. Tail concentration is the fraction of |Γn| carried by the largest 1% of samples. λNe and λTe are midplane SOL e-folding widths. The final rows test the radial Ne-fluctuation skew transition, particle/energy step changes, end-to-start particle-count drift and whether blob radial velocity is identifiable. Green, amber, red and gray mean pass, partial, fail and open under the declared battery rules; they are not probabilities. The battery uses one sanctioned mid-toroidal rollout dump and does not provide a full-ensemble realizability test. Evidence: docs/edge_physics_battery/physbat_6777298 and fluxtime2_f8cont_start216
Density–potential coupling in the edge physics battery. Curves labeled truth, free_member and anch_member are the simulation reference, one unassimilated ensemble member and one assimilated member. The spatial panels use poloidal wavenumber k_y in cycles per cell; the temporal panels use frequency f in cycles per stored frame. Cross-phase α is the argument of the Ne–φ cross-spectrum in degrees, and γ² is magnitude-squared coherence from 0 to 1. Radial E×B velocity uses v_r=−∂yφ/B.
Figure 11. Density–potential coupling in the edge physics battery. Curves labeled truth, free_member and anch_member are the simulation reference, one unassimilated ensemble member and one assimilated member. The spatial panels use poloidal wavenumber k_y in cycles per cell; the temporal panels use frequency f in cycles per stored frame. Cross-phase α is the argument of the Ne–φ cross-spectrum in degrees, and γ² is magnitude-squared coherence from 0 to 1. Radial E×B velocity uses v_r=−∂yφ/B. Evidence: docs/edge_physics_battery/physbat_6777298
Radial fluctuation and transport moments. Curves labeled truth, free_member and anch_member denote the simulation, one free member and one anchored member. The horizontal coordinate x is radial mesh index. Ne-tilde is the density fluctuation about its local mean; Γn is the radial particle-flux proxy formed from density fluctuation and radial E×B velocity. Skewness measures asymmetry and kurtosis measures tail weight. These profiles test transport structure that aggregate RMSE does not identify.
Figure 12. Radial fluctuation and transport moments. Curves labeled truth, free_member and anch_member denote the simulation, one free member and one anchored member. The horizontal coordinate x is radial mesh index. Ne-tilde is the density fluctuation about its local mean; Γn is the radial particle-flux proxy formed from density fluctuation and radial E×B velocity. Skewness measures asymmetry and kurtosis measures tail weight. These profiles test transport structure that aggregate RMSE does not identify. Evidence: docs/edge_physics_battery/physbat_6777298

Toroidal mode versus blob

n=35 is one Fourier component, whereas a localized blob is broadband. Blob fidelity is tested through coherent radial propagation and the coupled Ne–φ structure that sets E×B motion, not by the presence of one toroidal mode alone.

Unpredictable phase versus valid statistics

For the n≥80 tail, a stochastic emulator may be allowed to lose exact future phase if it reproduces conditional power, intermittency, dependence and transport contribution. Vorticity makes those tail statistics consequential.

07Evidence table and proposed sequence

topicsupported statementevidence limit / next decision
Toroidal learnabilityn=35 is present and coherent; the low-coherence tail begins near n≥80.Repeat the truth-only calculation on shot 85606 and any additional operating points.
f8 representationEfficient baseline with competitive aggregate RMSE and direct ceiling n=25.Cannot directly represent n=35; use as baseline and assimilation sandbox.
z44 representationReconstructs the n=35 edge peak.Its free/anchored aggregate gate is weaker than f8 and it allocates capacity into the low-coherence tail.
z22 representationNominal ceiling n=55; trained continuation completed.No held-out rollout gate, and reconstructed edge Ne peaks at n=25. It remains an untested middle candidate.
rollout CRPSEpoch 4 modestly improves held-out free and anchored RMSE.Underdispersion persists; fair-CRPS validation does not select that checkpoint; the full sensor ranking was not repeated.
φ-bearing suitesStrongest baseline-f8 idealized interventions in the three-start factorial.Needs a gauge-safe, noisy physical forward model and replication with a calibrated prior.
transport fidelitySome coarse gates pass, but cross-phase/flux and blob-motion tests are mixed or failing.These are acceptance gates before diagnostic-design conclusions.
  1. Complete the missing comparison. Evaluate z22 job 6791660 on shot 85606 with the same three starts, M=64, H=48 and stated ETKF protocol.
  2. Adopt an explicit codec gate. Require edge/SOL reconstruction and phase fidelity through k=7 (n=35), alongside physical-unit reconstruction error.
  3. Build a middle, band-aware representation. Test 16–22 toroidal latent cells or an explicit spectral branch, with reconstruction weight through k=7–8; use z44 as an upper-capacity reference.
  4. Separate predictable and stochastic targets. Use phase-sensitive rollout losses through n≈35–40, and distributional scores for the n≥80 tail, including vorticity statistics.
  5. Align selection with deployment. Evaluate checkpoints on absolute held-out free and anchored errors, calibration and physics gates; do not select on gain alone.
  6. Upgrade sensor realism. Declare observation operators, gauge handling, transfer functions, spatial averaging, correlations and noise for GPI, reflectometry and target probes.
  7. Expand the data. Add shots or operating points before interpreting architecture or diagnostic rankings as robust rather than within-regime.

08Definitions and reading guide

termdefinition in this report
AE / codecAutoencoder: encoder plus decoder used to move between five physical fields and a compressed latent grid.
LOLALatent diffusion emulator that samples future latent trajectories conditional on context frames.
epochOne configured training unit: 1,024 examples for Figure 4 runs; 128 rollout examples for job 6791994.
fine-tuneContinue from pretrained weights at a lower learning rate while holding the f8 autoencoder fixed; not a from-scratch retrain.
free / anchored rolloutFree uses only the emulator after the initial context. Anchored applies ETKF analyses using synthetic observations every four generated frames.
ensemble mean / spreadMean prediction across M sampled members / RMS of their unbiased cellwise variance.
RMSERoot mean squared error between ensemble mean and truth. Unless explicitly labeled AE VRMSE, rollout RMSE is dimensionless standardized-model-space error.
SSRSpread-to-error ratio. The precise aggregation differs between Figure 5 validation and Figures 6–7 held-out evaluation and is stated with each figure.
ρ, anchoring gainMean free error divided by mean anchored error; ρ>1 indicates improvement.
CRPS / afCRPS / bandCRPSContinuous ranked probability score; its almost-fair finite-ensemble form; and the same score applied after toroidal band filtering and truth-RMS normalization.
ETKF / inflationEnsemble transform Kalman filter / multiplicative expansion of forecast anomalies before assimilation.
M / H / startEnsemble-member count / rollout horizon in frames / first truth-frame index of an evaluated window.
Ne, Te, Ti, φ, Vi, VortElectron density, electron temperature, ion temperature, electrostatic potential, parallel ion velocity and vorticity.
edge / SOL / separatrixThe outer radial evaluation region / scrape-off layer on open field lines / boundary between closed and open magnetic flux surfaces. Edge-only Figure 7 uses cropped-grid x≥16.
R–Z / top-downPoloidal major-radius–vertical cross-section / view down the machine's vertical axis.
GPI / reflectometry / target proxyGas-puff-imaging-inspired support / radial microwave-diagnostic-inspired Ne chord / divertor-target point support. Current evaluator channels are idealized state samples, not complete instruments.
gauge dependenceφ and φ+C describe the same electric field for spatially constant C; gradients are gauge-invariant.
cross-phase α / coherence γ²Argument of the cross-spectrum between two fields / normalized magnitude-squared cross-spectral association.
Ne-tilde / ΓnDensity fluctuation about a local mean / radial particle-flux proxy based on density fluctuation and radial E×B velocity.
Qe,cond / λNe / λTeConductive electron heat-flux proxy / midplane SOL e-folding width of electron density / corresponding e-folding width of electron temperature.
realizability / tail concentrationRequirement that decoded physical temperatures remain positive / share of absolute flux carried by the largest 1% of events.
particle-count driftRatio of the final to initial domain-integrated particle count; a value near the truth ratio indicates comparable net drift over the rollout.
skewness / kurtosisNormalized third moment measuring asymmetry / normalized fourth moment measuring tail weight.

09Provenance, limitations and reproducibility

artifactlocation / jobsupports
BOUT metadata/mnt/home/sdelaurentiis/ceph/tcv-fresh-proj/85604/BOUT.dmp.0.ncZMAX=0.2, zperiod=5 and the one-fifth domain.
Truth HDF5/mnt/home/sdelaurentiis/ceph/tcv_well/TCV_85604/data/train/TCV_85604_train.hdf5Figure 1 coherence and power.
Representation manifestdocs/frontier_points.jsonFigure 3 codec metrics and f8/z44 gates.
z22 continuationjob 6791660Figure 4; training complete, held-out rollout gate absent.
f8 equal-budget controljob 6791711, gate 6792464Figures 3–4.
rollout-CRPS fine-tunejob 6791994Figure 5 and validation checkpoint selectors.
held-out checkpoint gatesjobs 6791995, 6791996, 6792287Figures 6–7.
baseline sensor factorialjobs 6771840, 6771851, 6771852Figure 9.
band-f8 sensor replicationjobs 67900706790072Independent φ/non-φ split at a smaller gain level.
physics batterydocs/edge_physics_battery/physbat_6777298/Figures 10–12.
machine-readable report inputsdocs/emulator_state_of_union_2026-08-10_data.jsonDerived Figures 1–7.
report builderdocs/build/build_emulator_state_of_union_2026_08_10.pyRecreates the derived figures and this self-contained HTML.

Primary method references

  1. BOUT++ input options: ZPERIOD defines the number of identical toroidal periods.
  2. LOLA: Lost in Latent Space (NeurIPS 2025): latent diffusion for probabilistic physical-system emulation.
  3. Probabilistic Retrofitting of Learned Simulators (2026): proper-scoring-rule retrofits for deterministic simulators; this motivates the separate comparison branch in Figure 2.
Scope boundary

The report summarizes what is supported by the current two-shot simulation dataset and logged experiments. It does not validate a real TCV diagnostic, reactor diagnostic selection, cross-regime generalization or transport-faithful digital-twin performance.