WiT: Waypoint Diffusion Transformers for Alleviating Trajectory Conflict in Pixel-Space Image Generation
Give pixel-space generation an explicit sense of structure. WiT dynamically predicts
compact semantic waypoints from noisy pixels and uses them to guide image prediction,
keeping the evolving generative state entirely in pixel space.
Hainuo Wang1,*Mingjia Li2,*Xiaojie Guo2,†
1 School of Computer Science and Technology 2 School of Software
Tianjin University, Tianjin, China
The repository is public, but the training and inference code will be open-sourced
after the paper is formally published.
Abstract
Direct pixel-space Flow Matching avoids the reconstruction bottleneck of latent
autoencoders, but noisy RGB observations provide little explicit semantic
organization. Nearby states can correspond to different target structures, even
under the same class condition. A shared, finite-capacity vector field must resolve
these locally mixed transport requirements, whose learning signals can interfere.
We call this optimization difficulty trajectory conflict.
WiT introduces explicit semantic routing through compact waypoints projected from
pretrained vision representations. At each integration timestep, a lightweight
Waypoints Predictor infers the waypoint from the current noisy pixel state.
Just-Pixel AdaLN exposes this structure to the Pixel Space Generator through
spatially varying modulation, while the generator directly predicts clean RGB pixels.
Experiments show consistent FID improvements over corresponding JiT baselines
across model scales, generator training durations, and ImageNet resolutions of
256×256 and 512×512. WiT-H/16 reaches FID 1.79.
Toy experiments, local semantic-neighborhood measurements, and state-matched
gradient analysis provide complementary evidence for the proposed mechanism.
The problem
Similar noisy pixels. Different target structures.
Noise hides pose, layout, scale, and spatial arrangement. Even within one class,
nearby noisy observations can require different transport directions. Semantic
waypoints make target-relevant structure explicit for the pixel generator.
Paper, Fig. 1. Standard pixel prediction and semantic routing (a–d), training loss and samples (e–f), and an analytic two-spiral illustration (g–j). The toy distribution shows its strongest intrinsic semantic conflict at intermediate timesteps. WiT improves recovery of the target manifold.
The role of a waypoint. It is a spatial routing condition recomputed
from the evolving noisy state. Semantic prediction and pixel prediction have distinct
roles within the architecture; the actual generative trajectory stays in pixel space.
Method
Predict structure. Modulate pixels. Repeat.
WiT separates semantic representation prediction from high-dimensional image
prediction, with a dynamically updated connection between them.
01
Build compact semantic targets
Frozen DINOv3-S/16 features supply dense spatial structure. PCA, fitted on 50,000
ImageNet training images, projects each patch feature to 64 dimensions. Semantic
organization comes from the pretrained encoder; PCA keeps the prediction target compact.
02
Infer the current waypoint
A 21M-parameter ViT-S/16 Waypoints Predictor takes the current noisy pixels,
timestep, and class condition and predicts a clean semantic waypoint. Its weights
are frozen for image generation, but its prediction changes as the pixel state evolves.
03
Guide clean-pixel prediction
Just-Pixel AdaLN converts each semantic token into spatially varying affine
modulation of the generator's attention and MLP branches. The native pixel-token
sequence is preserved, and the generator directly predicts the clean RGB image.
Paper, Fig. 2. The Waypoints Predictor infers a compact semantic representation from the current noisy pixel state. The Pixel Space Generator uses it as a spatial condition and directly predicts the clean image.
Paper, Fig. 3. Just-Pixel AdaLN supplies spatially varying modulation (a). As noisy pixel states evolve, the predicted semantic waypoints are recomputed throughout inference (b).
Training: learn semantics, then pixels
Train the Waypoints Predictor for 600 epochs, freeze its EMA weights, then train
the Pixel Space Generator with predicted waypoints as conditions. The generator
follows JiT's clean-image prediction with a velocity-space loss. The same frozen
predictor is reused across the reported experiments.
Inference: one evolving pixel state
At each vector-field evaluation, predict the waypoint from the current noisy state,
condition the pixel generator, and obtain a velocity update from its clean-image
prediction. Sampling uses a 50-step Heun solver. No separate semantic flow or
latent image decoder is sampled.
Algorithm 1. Training Procedure of WiT. Two-stage training: learn the semantic predictor, then freeze it to condition pixel-space generation.
Algorithm 2. Inference Procedure of WiT. Dynamic waypoint prediction and Just-Pixel AdaLN guide each pixel-space ODE evaluation.
ImageNet results
Gains across scale, training duration, and resolution
Semantic routing consistently improves FID over the corresponding JiT baselines,
from Base to Huge generators and from 256×256 to 512×512 images.
WiT-H/16 · 256×256
1.79 FID
Scaling to Huge
At 600 generator epochs, with IS 302.3. JiT-H/16 reaches 1.86 FID;
the larger, 2B-parameter JiT-G/16 reaches 1.82.
WiT-L/16 · 256×256
2.02 FID
Stronger Large-scale results
At 600 generator epochs, versus JiT-L/16's 2.36.
At 200 epochs, WiT-L/16 reaches 2.28 versus JiT-L/16's 2.79.
WiT-B/32 · 512×512
3.38 FID
Extending to higher resolution
At 600 generator epochs, versus JiT-B/32's 4.02.
The improvement also holds at 200 epochs: 3.89 versus 4.64.
Reading the comparisons. FID-50K is lower-is-better; IS is
higher-is-better. WiT and direct JiT baselines use the 50-step Heun solver.
Epochs count Pixel Space Generator training only: WiT additionally
uses a 21M predictor pretrained for 600 epochs, then frozen and reused. These are
comparisons at matched generator training durations, not matched total training cost.
ImageNet 256×256: direct baselines and WiT
ImageNet 256 by 256 direct comparisons, paper Table 2
Method
Params
Epochs
IS ↑
FID-50K ↓
JiT-B/16
131M
200
-
4.37
WiT-B/16
131M + 21M
200
270.7
3.34
JiT-B/16
131M
600
275.1
3.66
WiT-B/16
131M + 21M
600
280.2
3.03
JiT-L/16
459M
200
-
2.79
LF-DiT-L/16
465M
200
-
2.48
WiT-L/16
459M + 21M
200
276.6
2.28
JiT-L/16
459M
600
298.5
2.36
WiT-L/16
459M + 21M
600
293.3
2.02
WiT-XL/16
676M + 21M
200
288.9
2.16
WiT-XL/16
676M + 21M
600
301.0
1.89
JiT-H/16
953M
600
303.4
1.86
JiT-G/16
2B
600
292.6
1.82
WiT-H/16
953M + 21M
600
302.3
1.79
Paper, Table 2. WiT parameter counts show the pixel generator
plus the 21M Waypoints Predictor. A dash denotes a value not reported in the table.
The remaining entries of Table 2, on ImageNet 256×256.
These results come from the respective publications and use different architectures
and training setups; they are not controlled compute-matched comparisons.
Broader ImageNet 256 by 256 comparisons, paper Table 2
Method
Params
Epochs
IS ↑
FID-50K ↓
Latent-space diffusion models
DiT-XL/2
675M + 49M
-
278.2
2.27
SiT-XL/2
675M + 49M
-
277.5
2.06
REPA (SiT-XL/2)
675M + 49M
-
305.7
1.42
LightningDiT-XL/2
675M + 49M
-
295.3
1.35
DDT-XL/2
675M + 49M
-
310.6
1.26
RAE (DiTDH-XL/2)
839M + 415M
-
262.6
1.13
Pixel-space models (non-diffusion)
JetFormer
2.8B
-
-
6.64
FractalMAR-H
848M
-
348.9
6.15
Pixel-space diffusion models
ADM-G
554M
-
186.7
4.59
RIN
410M
-
182.0
3.42
SiD (UViT/2)
2B
-
256.3
2.44
PixelFlow (XL/4)
677M
-
282.1
1.98
PixNerd (XL/16)
700M
-
297.0
2.15
Understanding trajectory conflict
From local semantic mixing to gradient interference
Three complementary analyses examine whether target-relevant semantic routing
improves manifold recovery, organizes local neighborhoods, and reduces conflicting
learning signals in the shared generator.
01 · Controlled toy experiment
Target-relevant routes improve manifold recovery
On two-spiral data embedded in 4, 8, 16, 32, 64, and 128 dimensions, WiT lowers
normalized nearest-manifold RMSE at every evaluated dimension. Relative reductions
versus the jointly optimized JiT baseline are 7.0%, 11.5%, 9.3%, 21.7%,
15.9%, and 14.1%, respectively. A parameter-matched Random Route baseline
with irrelevant routing labels does not yield comparable, consistently positive gains.
Paper, Fig. 6. Samples are shown in a 2D projection. Colors and nRMSE measure distances in the original observation space, normalized by dimension. Lower nRMSE means better manifold recovery; red percentages are WiT reductions relative to JiT.
02 · Representation-level evidence
More coherent semantic neighborhoods
Normalized Local Conditional Variance (NLCV) measures the dispersion of ground-truth
semantic endpoints within local neighborhoods. Holding the ImageNet class fixed,
neighborhoods formed from predicted waypoints have 23.8% to 35.6% lower
semantic dispersion than those formed from noisy pixel observations.
Within-condition NLCV, lower is better, paper Table 7
Timestep t
Raw state
Predicted semantic
Reduction
0.16
0.9713
0.7401
23.8%
0.31
0.9668
0.6578
32.0%
0.51
0.9636
0.6208
35.6%
Paper, Table 7. Each group contains an anchor and four
same-class nearest neighbors in a 64D PCA-whitened query space. Raw queries use
patch-mean noisy pixels; semantic queries use mean-pooled predicted waypoint tokens.
NLCV is normalized by within-class semantic variance and macro-averaged across classes.
03 · Optimization-level evidence
Less disagreement-dependent gradient interference
Sample pairs are matched by class condition, timestep, and noisy-state distance,
then separated into high and low semantic-target disagreement groups. Greater
disagreement is associated with more output-layer gradient opposition in both
models, but the effect is weaker in WiT.
State-matched output-layer gradient interference, paper Table 8
High-minus-low effect
JiT
WiT
Relative mitigation
Gradient cosine
−0.15961
−0.13677
14.3%
Negative gradient mass
0.07434
0.05833
21.5%
Paper, Table 8. Less-negative cosine and a smaller increase
in negative mass indicate weaker disagreement-dependent interference. Absolute
mitigation: 0.02284 [0.00200, 0.04391] and 0.01601 [0.00451, 0.02730], respectively.
Brackets are the paper's confidence intervals, estimated with 10,000
ImageNet-class cluster bootstrap samples.
What the evidence supports. Predicted waypoints reorganize the
representation presented to a finite-capacity pixel generator and reduce
disagreement-dependent learning interference. These measurements support semantic
routing as an aid to optimization; they do not establish that geometric path
intersections disappear or that a separate semantic trajectory is generated.
Design choices
Why compact waypoints and spatial modulation?
Ablations compare representation guidance, source encoders, waypoint dimension,
and injection strategy within the evaluated ImageNet settings.
Representation-guided alternatives
Representation guidance at 200 generator epochs, paper Table 4
JiT-B/16 variant
FID ↓
JiT baseline
4.37
+ REPA
5.14
+ RCG
4.96
+ PixelREPA
4.00
WiT-B/16
3.34
Paper, Table 4. Same JiT-B/16 backbone, 200 generator epochs.
Results describe these implementations under the evaluated setting.
Choice of pretrained representation
Waypoint source with B/16 at 600 generator epochs, paper Table 5
Waypoint source
FID ↓
JiT (no waypoint)
3.66
MoCoV3-S/16
3.11
MAE-B/16
3.32
DINOv2-S/14
3.36
DINOv3-S/16 (default)
3.03
Paper, Table 5. B/16 generator, 600 generator epochs.
Every evaluated encoder improves over JiT; DINOv3-S/16 gives the lowest FID.
Waypoint dimension and injection mechanism
WiT-B/16 waypoint and injection ablation at 200 epochs, paper Table 6
PCA dimension
Injection
IS ↑
FID ↓
32
Just-Pixel AdaLN
210.40
5.11
128
Just-Pixel AdaLN
211.33
4.12
64
Channel Concat
221.19
3.93
64
In-context Concat
238.92
3.63
64 (WiT default)
Just-Pixel AdaLN
270.73
3.34
Paper, Table 6. WiT-B/16, 200 generator epochs. Among tested
configurations, 64D waypoints balance compactness and retained structure, while
Just-Pixel AdaLN outperforms channel and token concatenation.
Guidance behavior (paper, Fig. 5). The minimum-FID CFG scale shifts
from 3.8 for WiT-B/16 at 200 epochs to 3.1 at 600 epochs and 2.9 for WiT-L/16 at
600 epochs. Excessively strong guidance degrades FID in these sweeps.
Qualitative samples
Global structure and fine pixel detail
ImageNet 256×256 samples from Large, Extra-Large,
and Huge models. Select any figure to inspect it at full size.
WiT-L/16 · Paper, Fig. 4. Representative samples illustrate coherent object layouts alongside fine local textures, with generation performed directly in pixel space.
WiT-XL/16 · Paper, Fig. 7. ImageNet 256×256.
WiT-H/16 · Paper, Fig. 8. ImageNet 256×256.
Repository Status
Training and inference code will be released here
This repository is the official landing point for WiT. The current public state is
the project page plus a placeholder code entry. Training and inference code will be
released in this same repository after the paper is formally published.
We thank Qiming Hu
for insightful discussions and feedback. This work was partially supported by
computational resources from TPU Research Cloud (TRC).
Citation
BibTeX
@article{wang2026wit,
title={WiT: Waypoint Diffusion Transformers for Alleviating Trajectory Conflict in Pixel-Space Image Generation},
author={Wang, Hainuo and Li, Mingjia and Guo, Xiaojie},
journal={arXiv preprint arXiv:2603.15132},
year={2026}
}