Which Geometry Predicts Few-Step Degradation in Latent Flow Matching?
One trajectory, three spaces
Flow matching with a linear interpolant prescribes straight conditional paths. The trajectory the learned field actually generates need not be straight, and in a latent model it is transformed twice before anyone looks at it. 𝒟φ is a frozen convolutional autoencoder's decoder. 𝒦 is the deterministic kinematic map (root integration and root-relative reconstruction).
CDFM is the control. It integrates directly in 𝒳, so any 𝒳 → 𝒬 effect it shows comes from the kinematic map alone, with no autoencoder involved.
Trajectories are schematic; the bars are measured over 512 clips. Spread is the standard deviation of arc-to-chord across clips (log scale). Predicts is its Spearman correlation with 16-step degradation. Downstream maps bend trajectories more and more unevenly, while the predictive signal fades.
Where each quantity lives
A standard Euler bound (Proposition 1) shows which quantities can affect the mapped error, and in which space each one is defined:
Here \(m\in\{\mathrm{LFM},\mathrm{CDFM}\}\) indexes the model, \(\boldsymbol S_t^{(m)}\) is the trajectory in the integration space \(\mathcal S\) (\(\mathcal S^{(\mathrm{LFM})}=\mathcal Z\), \(\mathcal S^{(\mathrm{CDFM})}=\mathcal X\)), and \(\widehat{\boldsymbol S}_1^{(m,n_{\mathrm{steps}})}\) is the \(n_{\mathrm{steps}}\)-step Euler endpoint. The output map is \(\mathcal F^{(\mathrm{LFM})}=\mathcal K\circ\mathcal D_\varphi\) and \(\mathcal F^{(\mathrm{CDFM})}=\mathcal K\). \(L_v^{(m)}\) is the Lipschitz constant of the learned field \(v_\theta^{(m)}\), and \(L_{\mathcal F}^{(m,n_{\mathrm{steps}})}\) is the supremum of the local output-map operator norm over the segment joining the exact and Euler endpoints.
The measurements below decide which of these quantities actually matter.
What the measurements show
HumanML3D text-to-motion. 512 fixed clips, 5 sampling replications, and 95% t-intervals across replications.
Distortion grows across the maps, but dispersion grows faster
We measure the arc-to-chord ratio of the same 50-step reference trajectory in each space. Its mean rises modestly, but its spread across clips grows by up to two orders of magnitude. CDFM's growth from 0.0034 to 0.3211 happens with no autoencoder in the pipeline, so the decoder is not the only source of distortion.
| Model | Space | Mean | SD across clips | Min | Max |
|---|---|---|---|---|---|
| LFM | 𝒵 | 1.0153 | 0.0052 | 1.006 | 1.033 |
| 𝒳 | 1.1572 | 0.0928 | 1.036 | 1.577 | |
| 𝒬 | 1.2658 | 0.5526 | 1.001 | 4.869 | |
| CDFM | 𝒳 | 1.0087 | 0.0034 | 1.003 | 1.026 |
| 𝒬 | 1.1003 | 0.3211 | 1.000 | 6.090 |
The predictive signal lives in the integration space
For LFM, the association is positive in 𝒵, collapses at the decoder (before the kinematic map is applied), and cannot be distinguished from zero in 𝒬. A representation map can make a trajectory look more bent while destroying the statistic's relation to integration difficulty. Geometric distortion and loss of predictive signal are different phenomena.
Spearman association between per-clip statistics of the 50-step reference trajectory and matched few-step degradation δ𝒬, plotted against the number of steps. Error bars are 95% t-intervals over five replications. Hover a legend entry to isolate its series.
Local beats global, and gain beats both
For LFM at one step, maximum turning angle clearly beats arc-to-chord (0.538 ± 0.032 vs 0.274 ± 0.021). This fits the bound depending on a local supremum rather than a path average. Empirical output-map gain is the strongest single predictor for both models at every budget.
The gain is an amplifier, and more than one
A sensitive output map would correlate with degradation through amplification alone. We therefore run two controls, and the association survives both. The trajectory-mean gain is stronger than the terminal value for both models at every budget. So sensitivity met along the path carries information beyond the final state.
| Control | LFM | CDFM |
|---|---|---|
| vs. integration-space degradation δint amplification impossible | 0.465–0.495 | 0.421–0.522 |
| partial association with δ𝒬 given δint | 0.648–0.690 | 0.471–0.495 |
Arc-length schedules don't beat cosine
Placing timesteps by equalising cumulative arc length does not beat a cosine schedule, which gives the lowest FID at every budget for both models. For LFM at four steps, cosine reaches 0.453, against 1.851 for the integration-space arc-matched grid and 2.077 for the 𝒬-profile grid. Arc length is a first-order quantity, while Euler truncation error is governed by second-order variation. The representation space and the scheduling statistic are separate design choices.
Straighten the predictive geometry, and few-step quality follows
A reflow student is fine-tuned on 20,096 self-generated pairs from the guidance-2.5 teacher. Its integration-space trajectories become much straighter. At matched guidance, one-step FID falls from 17.17 to 0.75, and a single function evaluation nearly matches the guidance-2.5 teacher at eight steps (0.76). The advantage is confined to the few-step regime: the guided teacher overtakes the student between 8 and 16 steps.
LFM teacher at classifier-free-guidance scales 1.0, 1.5 and 2.5, and the reflowed student at guidance 1.0. Three sampling replications, with 95% confidence intervals. The student's curve is nearly flat.
| Model | Guid. | Arc/chord | Max turn | FID @1 | @8 | @16 | R@3 @1 |
|---|---|---|---|---|---|---|---|
| Teacher | 1.0 | 1.0088 | 0.164 | 17.17 | 1.88 | 0.95 | 0.283 |
| Teacher | 1.5 | 1.0096 | 0.221 | 10.87 | 1.13 | 0.58 | 0.371 |
| Teacher | 2.5 | 1.0153 | 0.357 | 7.36 | 0.76 | 0.36 | 0.430 |
| Reflow | 1.0 | 1.0020 | 0.084 | 0.75 | 0.60 | 0.54 | 0.548 |
Body-part composition
The CDFM checkpoint analysed in the paper can also compose motions at sampling time. Two samples are generated from different prompts, and the upper body of one is blended with the lower body of the other using a feathered mask across the spine.
This is a qualitative illustration built on the same checkpoint. It is not part of the paper's evaluation.
a person waves both hands
a person walks forward
Reproduce every table and figure
All code is in azizkhadraoui/few-step-geometry-code. Everything except reflow training runs inference-only on frozen checkpoints, on a single 16 GB V100.
git clone https://github.com/azizkhadraoui/few-step-geometry-code
cd few-step-geometry-code
pip install -r requirements.txt
export HML3D_ROOT=/path/to/HumanML3D/humanml
export RVQ_CKPT=/path/to/rvq_vae_best.pt
export EVAL_ROOT=/path/to/evaluator
export WORK_DIR=./runs
# train the two base models once
VARIANT=latent python lfm_clfm_cdfm_experiment.py
VARIANT=direct python lfm_clfm_cdfm_experiment.py
# then run a measurement
CV_REPS=5 EVAL_N=512 python curvature_nfe_v2.py
| Measures | Script |
|---|---|
| Geometry in 𝒵, 𝒳, 𝒬 | three_space_geometry.py |
| Predictors & degradation | curvature_nfe_v2.py |
| Gain controls | gain_disentangle.py |
| Step schedules | curvature_schedule.py |
| Reflow student | reflow_experiment.py |
| Guidance vs. reflow | guidance_control.py |
Each script writes a JSON to $WORK_DIR. Slurm launchers are in slurm/:
set the paths once in slurm/env.sh, then sbatch slurm/run_*.sh. The repository README
maps every reported number to the function that produces it.
Conventions that matter
Fixed clip set
np.random.default_rng(0) and take the first 512
indices of a permutation of the test indices, sorted. The set is identical across every experiment and held fixed
across replications.Paired seeding
s in replication
r, the seed is s + 100000*r. Each clip therefore gets the same noise at every step budget,
which makes degradation a paired, within-clip comparison.Two levels of averaging
Additional Reproducibility and Clarification Notes
Precise statistical and reproducibility details that complement the final paper. They do not change any of its claims.
Dataset preprocessing
Our HumanML3D preprocessing contains 13,840 training motions and 2,644 test motions. The mirrored
M-prefixed entries are not materialized. Absolute FID values should therefore not be compared
directly with published HumanML3D leaderboard values. All comparisons in the paper use the same fixed
preprocessing and evaluation protocol.
𝒬-space association for LFM
For LFM, the 𝒬-space arc-to-chord associations with few-step degradation are small in magnitude and negative, ranging from −0.091 to −0.024. Only the sixteen-step interval spans zero. This is the precise interpretation of the result.
| Steps | 1 | 2 | 4 | 8 | 16 |
|---|---|---|---|---|---|
| Spearman ρ | −0.055 | −0.074 | −0.091 | −0.061 | −0.024 |
| 95% interval | ±0.021 | ±0.016 | ±0.032 | ±0.043 | ±0.041 |
| Spans zero | no | no | no | no | yes |
Gain evaluation choice
The gain can be evaluated in four ways: at the terminal state, averaged along the trajectory, as the trajectory maximum, or at the endpoint-segment midpoint. Against integration-space degradation δint, the trajectory-mean gain has the highest point estimate of the four at every budget, for both models. Its interval overlaps that of the trajectory-maximum variant at every LFM budget and at 2–16 steps for CDFM, so the ordering between those two variants is descriptive.
| Model | Gain evaluation | 1 | 2 | 4 | 8 | 16 |
|---|---|---|---|---|---|---|
| LFM | Trajectory mean | 0.591±0.018 | 0.606±0.013 | 0.595±0.010 | 0.564±0.017 | 0.547±0.021 |
| Trajectory maximum | 0.584±0.027 | 0.594±0.022 | 0.576±0.014 | 0.544±0.016 | 0.528±0.022 | |
| CDFM | Trajectory mean | 0.547±0.011 | 0.561±0.013 | 0.533±0.018 | 0.483±0.020 | 0.468±0.018 |
| Trajectory maximum | 0.523±0.009 | 0.538±0.011 | 0.510±0.016 | 0.459±0.020 | 0.444±0.020 |
The terminal-state and endpoint-segment-midpoint variants are lower at every budget. The full table is in the paper's appendix.
Implementation details
- Text encoder
- Frozen
sentence-transformers/sentence-t5-base(T5 encoder, 768-d). Captions are padded or truncated to 64 tokens. Both flow models are conditioned on the token-level hidden states and on their masked mean-pooled, L2-normalized embedding. Classifier-free guidance uses a learned null embedding, with conditioning dropped 10% of the time during training. - Autoencoder
- Frozen 1-D convolutional autoencoder with GroupNorm and SiLU. The encoder maps the 263-d motion features through 256 → 512 → 256 channels with 4× temporal downsampling, so a 196-frame clip becomes a 49 × 256 continuous latent. The decoder mirrors it with nearest-neighbour upsampling. The residual quantizer stored in the checkpoint is bypassed: the flow acts on continuous codes, standardized per channel with training-set statistics.
The autoencoder is used frozen throughout; its architecture is defined in lfm_clfm_cdfm_experiment.py.
Known issues and caveats
What the intervals do and don't cover
- The intervals cover sampler variance only: a fixed clip set, one trained model per type, and 3–5 replications. Training and clip-set variance are not included, so the intervals understate total uncertainty.
- Differences between two Spearman coefficients are descriptive. We test them only where the paper states an explicit dependent-correlation comparison.
Measurement proxies
- Arc-to-chord and turning angle are computed on the common n = 50 grid and are comparable only at that resolution. Neither estimates the bound's second-order term.
- The gain is a 4-probe finite difference, closer to a dimension-normalised Frobenius norm than to the operator norm. Integration-space dimensions differ (d = 12,544 for LFM vs 51,548 for CDFM), so compare gains only within a model.
Confounders and scope
- We do not control for sequence length. It is a conditioning input and a plausible common cause of both gain and integration difficulty.
- The stratified comparison uses 128-clip strata and retains a +0.87 residual at fifty steps. Read the few-step gap against that floor.
- Reflow covers a single student and configuration. It shows that straightening can coincide with much better few-step behaviour, not that the size of the effect or the crossover point generalises.
BibTeX
@inproceedings{khadraoui2026geometry,
title = {Which Geometry Predicts Few-Step Degradation in Latent Flow Matching?},
author = {Mohamed Aziz Khadraoui and Abdelkader Baggag},
booktitle = {NeurIPS 2026 Workshop on Beyond Next-Token Prediction:
Diffusion \& Flow Models for Next-Generation Decoding (BeNTo)},
year = {2026}
}