BeNTo Workshop @ NeurIPS 2026 Beyond Next-Token Prediction — Diffusion & Flow Models for Next-Generation Decoding

Which Geometry Predicts Few-Step Degradation in Latent Flow Matching?

Mohamed Aziz Khadraoui1  ·  Abdelkader Baggag2
1Higher School of Communication of Tunis (SUP'COM), Tunisia 2Qatar Computing Research Institute, Hamad Bin Khalifa University
a person walks forward and then turns around
a person jumps up and lands on both feet
a person waves both hands

Text-to-motion samples from the latent flow model (LFM) studied in the paper, with 50 Euler steps and guidance 2.5.

TL;DR. In a latent flow model the solver integrates in one space, but the sample is only evaluated after nonlinear maps. We follow the same trajectory through three spaces (𝒵 → 𝒳 → 𝒬) and find that only the integration-space geometry predicts few-step degradation. Downstream maps bend the trajectory more, and make its shape far more sample-dependent, but that bending says nothing about how hard the trajectory is to integrate. What the downstream map does contribute is local sensitivity, and that is strongly predictive.
Setup

One trajectory, three spaces

Flow matching with a linear interpolant prescribes straight conditional paths. The trajectory the learned field actually generates need not be straight, and in a latent model it is transformed twice before anyone looks at it. 𝒟φ is a frozen convolutional autoencoder's decoder. 𝒦 is the deterministic kinematic map (root integration and root-relative reconstruction).

CDFM is the control. It integrates directly in 𝒳, so any 𝒳 → 𝒬 effect it shows comes from the kinematic map alone, with no autoencoder involved.

one clip, noise → sample4 Euler steps

Trajectories are schematic; the bars are measured over 512 clips. Spread is the standard deviation of arc-to-chord across clips (log scale). Predicts is its Spearman correlation with 16-step degradation. Downstream maps bend trajectories more and more unevenly, while the predictive signal fades.

Theory

Where each quantity lives

A standard Euler bound (Proposition 1) shows which quantities can affect the mapped error, and in which space each one is defined:

$$\bigl\|\mathcal F^{(m)}(\widehat{\boldsymbol S}_1^{(m,n_{\mathrm{steps}})}) - \mathcal F^{(m)}(\boldsymbol S_1^{(m)})\bigr\|_{\mathcal Q} \;\le\; \underbrace{L_{\mathcal F}^{(m,n_{\mathrm{steps}})}}_{\text{sensitivity}}\,\frac{\Delta t}{2}\,\underbrace{\frac{e^{L_v^{(m)}}-1}{L_v^{(m)}}}_{\text{stability}}\,\underbrace{\sup_{t\in[0,1]}\bigl\|\ddot{\boldsymbol S}_t^{(m)}\bigr\|_{\mathcal S}}_{\text{2nd-order variation}}$$

Here \(m\in\{\mathrm{LFM},\mathrm{CDFM}\}\) indexes the model, \(\boldsymbol S_t^{(m)}\) is the trajectory in the integration space \(\mathcal S\) (\(\mathcal S^{(\mathrm{LFM})}=\mathcal Z\), \(\mathcal S^{(\mathrm{CDFM})}=\mathcal X\)), and \(\widehat{\boldsymbol S}_1^{(m,n_{\mathrm{steps}})}\) is the \(n_{\mathrm{steps}}\)-step Euler endpoint. The output map is \(\mathcal F^{(\mathrm{LFM})}=\mathcal K\circ\mathcal D_\varphi\) and \(\mathcal F^{(\mathrm{CDFM})}=\mathcal K\). \(L_v^{(m)}\) is the Lipschitz constant of the learned field \(v_\theta^{(m)}\), and \(L_{\mathcal F}^{(m,n_{\mathrm{steps}})}\) is the supremum of the local output-map operator norm over the segment joining the exact and Euler endpoints.

Output-map sensitivityThe only way the downstream map enters.
Stability & 2nd-order variationBoth are properties of the integration space \(\mathcal S\).
Bending of the mapped pathDoes not appear anywhere in the bound.

The measurements below decide which of these quantities actually matter.

Results

What the measurements show

HumanML3D text-to-motion. 512 fixed clips, 5 sampling replications, and 95% t-intervals across replications.

1

Distortion grows across the maps, but dispersion grows faster

We measure the arc-to-chord ratio of the same 50-step reference trajectory in each space. Its mean rises modestly, but its spread across clips grows by up to two orders of magnitude. CDFM's growth from 0.0034 to 0.3211 happens with no autoencoder in the pipeline, so the decoder is not the only source of distortion.

Arc-to-chord SD across clips (log scale)
ModelSpaceMeanSD across clipsMinMax
LFM𝒵1.01530.00521.0061.033
𝒳1.15720.09281.0361.577
𝒬1.26580.55261.0014.869
CDFM𝒳1.00870.00341.0031.026
𝒬1.10030.32111.0006.090
2

The predictive signal lives in the integration space

For LFM, the association is positive in 𝒵, collapses at the decoder (before the kinematic map is applied), and cannot be distinguished from zero in 𝒬. A representation map can make a trajectory look more bent while destroying the statistic's relation to integration difficulty. Geometric distortion and loss of predictive signal are different phenomena.

(a) Arc-to-chord, by space
(b) Predictors, LFM (in 𝒵)
(c) Predictors, CDFM (in 𝒳)

Spearman association between per-clip statistics of the 50-step reference trajectory and matched few-step degradation δ𝒬, plotted against the number of steps. Error bars are 95% t-intervals over five replications. Hover a legend entry to isolate its series.

3

Local beats global, and gain beats both

For LFM at one step, maximum turning angle clearly beats arc-to-chord (0.538 ± 0.032 vs 0.274 ± 0.021). This fits the bound depending on a local supremum rather than a path average. Empirical output-map gain is the strongest single predictor for both models at every budget.

0.72–0.78
gain, LFM
0.60–0.65
gain, CDFM
0.54 vs 0.27
turn vs arc, LFM @1
4

The gain is an amplifier, and more than one

A sensitive output map would correlate with degradation through amplification alone. We therefore run two controls, and the association survives both. The trajectory-mean gain is stronger than the terminal value for both models at every budget. So sensitivity met along the path carries information beyond the final state.

ControlLFMCDFM
vs. integration-space degradation δint
amplification impossible
0.465–0.4950.421–0.522
partial association with δ𝒬 given δint0.648–0.6900.471–0.495
One effect does not survive the controls, and we flag it. CDFM's segment-midpoint gain beats its terminal gain by +0.107 at one step against δ𝒬, but by only +0.002 against δint. The effect comes from the amplification stage, so we report it only as a 𝒬-space observation.
5

Arc-length schedules don't beat cosine

Placing timesteps by equalising cumulative arc length does not beat a cosine schedule, which gives the lowest FID at every budget for both models. For LFM at four steps, cosine reaches 0.453, against 1.851 for the integration-space arc-matched grid and 2.077 for the 𝒬-profile grid. Arc length is a first-order quantity, while Euler truncation error is governed by second-order variation. The representation space and the scheduling statistic are separate design choices.

Intervention

Straighten the predictive geometry, and few-step quality follows

A reflow student is fine-tuned on 20,096 self-generated pairs from the guidance-2.5 teacher. Its integration-space trajectories become much straighter. At matched guidance, one-step FID falls from 17.17 to 0.75, and a single function evaluation nearly matches the guidance-2.5 teacher at eight steps (0.76). The advantage is confined to the few-step regime: the guided teacher overtakes the student between 8 and 16 steps.

4.2×
lower max turning angle
17.2 → 0.75
one-step FID, guidance 1.0
0.43 → 0.55
one-step R@3 vs best teacher
FID vs sampling steps (log–log)

LFM teacher at classifier-free-guidance scales 1.0, 1.5 and 2.5, and the reflowed student at guidance 1.0. Three sampling replications, with 95% confidence intervals. The student's curve is nearly flat.

ModelGuid.Arc/chordMax turnFID @1@8@16R@3 @1
Teacher1.01.00880.16417.171.880.950.283
Teacher1.51.00960.22110.871.130.580.371
Teacher2.51.01530.3577.360.760.360.430
Reflow1.01.00200.0840.750.600.540.548
Guidance bends the trajectory while improving it, and that is not a counterexample. Raising guidance from 1.0 to 2.5 raises the turning angle from 0.164 to 0.357 and improves one-step FID from 17.17 to 7.36. Our associations are measured across clips at a fixed velocity field, and changing guidance changes the field. Trajectory geometry predicts difficulty within a field; it is not a criterion for choosing between models.
Demo

Body-part composition

The CDFM checkpoint analysed in the paper can also compose motions at sampling time. Two samples are generated from different prompts, and the upper body of one is blended with the lower body of the other using a feathered mask across the spine.

This is a qualitative illustration built on the same checkpoint. It is not part of the paper's evaluation.

Upper bodya person waves both hands
Lower bodya person walks forward
ComposedWalks forward while waving
Code

Reproduce every table and figure

All code is in azizkhadraoui/few-step-geometry-code. Everything except reflow training runs inference-only on frozen checkpoints, on a single 16 GB V100.

git clone https://github.com/azizkhadraoui/few-step-geometry-code
cd few-step-geometry-code
pip install -r requirements.txt

export HML3D_ROOT=/path/to/HumanML3D/humanml
export RVQ_CKPT=/path/to/rvq_vae_best.pt
export EVAL_ROOT=/path/to/evaluator
export WORK_DIR=./runs

# train the two base models once
VARIANT=latent python lfm_clfm_cdfm_experiment.py
VARIANT=direct python lfm_clfm_cdfm_experiment.py

# then run a measurement
CV_REPS=5 EVAL_N=512 python curvature_nfe_v2.py
MeasuresScript
Geometry in 𝒵, 𝒳, 𝒬three_space_geometry.py
Predictors & degradationcurvature_nfe_v2.py
Gain controlsgain_disentangle.py
Step schedulescurvature_schedule.py
Reflow studentreflow_experiment.py
Guidance vs. reflowguidance_control.py

Each script writes a JSON to $WORK_DIR. Slurm launchers are in slurm/: set the paths once in slurm/env.sh, then sbatch slurm/run_*.sh. The repository README maps every reported number to the function that produces it.

Sample the reflow student at guidance 1.0. Its training pairs were generated with classifier-free guidance already applied, so sampling it at 2.5 applies guidance twice and makes a working model look broken.

Conventions that matter

Fixed clip set
We use np.random.default_rng(0) and take the first 512 indices of a permutation of the test indices, sorted. The set is identical across every experiment and held fixed across replications.
Paired seeding
For a batch starting at index s in replication r, the seed is s + 100000*r. Each clip therefore gets the same noise at every step budget, which makes degradation a paired, within-clip comparison.
Two levels of averaging
Statistics are computed per clip and reduced to one value per replication (a mean, or a Spearman coefficient across the 512 clips). Intervals are then taken across replications, so they measure sampling variance. Dispersion across clips is reported separately.
Camera-ready

Additional Reproducibility and Clarification Notes

Precise statistical and reproducibility details that complement the final paper. They do not change any of its claims.

1

Dataset preprocessing

Our HumanML3D preprocessing contains 13,840 training motions and 2,644 test motions. The mirrored M-prefixed entries are not materialized. Absolute FID values should therefore not be compared directly with published HumanML3D leaderboard values. All comparisons in the paper use the same fixed preprocessing and evaluation protocol.

2

𝒬-space association for LFM

For LFM, the 𝒬-space arc-to-chord associations with few-step degradation are small in magnitude and negative, ranging from −0.091 to −0.024. Only the sixteen-step interval spans zero. This is the precise interpretation of the result.

Steps124816
Spearman ρ−0.055−0.074−0.091−0.061−0.024
95% interval±0.021±0.016±0.032±0.043±0.041
Spans zerononononoyes
3

Gain evaluation choice

The gain can be evaluated in four ways: at the terminal state, averaged along the trajectory, as the trajectory maximum, or at the endpoint-segment midpoint. Against integration-space degradation δint, the trajectory-mean gain has the highest point estimate of the four at every budget, for both models. Its interval overlaps that of the trajectory-maximum variant at every LFM budget and at 2–16 steps for CDFM, so the ordering between those two variants is descriptive.

ModelGain evaluation124816
LFMTrajectory mean0.591±0.0180.606±0.0130.595±0.0100.564±0.0170.547±0.021
Trajectory maximum0.584±0.0270.594±0.0220.576±0.0140.544±0.0160.528±0.022
CDFMTrajectory mean0.547±0.0110.561±0.0130.533±0.0180.483±0.0200.468±0.018
Trajectory maximum0.523±0.0090.538±0.0110.510±0.0160.459±0.0200.444±0.020

The terminal-state and endpoint-segment-midpoint variants are lower at every budget. The full table is in the paper's appendix.

4

Implementation details

Text encoder
Frozen sentence-transformers/sentence-t5-base (T5 encoder, 768-d). Captions are padded or truncated to 64 tokens. Both flow models are conditioned on the token-level hidden states and on their masked mean-pooled, L2-normalized embedding. Classifier-free guidance uses a learned null embedding, with conditioning dropped 10% of the time during training.
Autoencoder
Frozen 1-D convolutional autoencoder with GroupNorm and SiLU. The encoder maps the 263-d motion features through 256 → 512 → 256 channels with 4× temporal downsampling, so a 196-frame clip becomes a 49 × 256 continuous latent. The decoder mirrors it with nearest-neighbour upsampling. The residual quantizer stored in the checkpoint is bypassed: the flow acts on continuous codes, standardized per channel with training-set statistics.

The autoencoder is used frozen throughout; its architecture is defined in lfm_clfm_cdfm_experiment.py.

Limitations

Known issues and caveats

What the intervals do and don't cover
  • The intervals cover sampler variance only: a fixed clip set, one trained model per type, and 3–5 replications. Training and clip-set variance are not included, so the intervals understate total uncertainty.
  • Differences between two Spearman coefficients are descriptive. We test them only where the paper states an explicit dependent-correlation comparison.
Measurement proxies
  • Arc-to-chord and turning angle are computed on the common n = 50 grid and are comparable only at that resolution. Neither estimates the bound's second-order term.
  • The gain is a 4-probe finite difference, closer to a dimension-normalised Frobenius norm than to the operator norm. Integration-space dimensions differ (d = 12,544 for LFM vs 51,548 for CDFM), so compare gains only within a model.
Confounders and scope
  • We do not control for sequence length. It is a conditioning input and a plausible common cause of both gain and integration difficulty.
  • The stratified comparison uses 128-clip strata and retains a +0.87 residual at fifty steps. Read the few-step gap against that floor.
  • Reflow covers a single student and configuration. It shows that straightening can coincide with much better few-step behaviour, not that the size of the effect or the crossover point generalises.
Citation

BibTeX

@inproceedings{khadraoui2026geometry,
  title     = {Which Geometry Predicts Few-Step Degradation in Latent Flow Matching?},
  author    = {Mohamed Aziz Khadraoui and Abdelkader Baggag},
  booktitle = {NeurIPS 2026 Workshop on Beyond Next-Token Prediction:
               Diffusion \& Flow Models for Next-Generation Decoding (BeNTo)},
  year      = {2026}
}