The problem

Similar noisy pixels. Different target structures.

Noise hides pose, layout, scale, and spatial arrangement. Even within one class, nearby noisy observations can require different transport directions. Semantic waypoints make target-relevant structure explicit for the pixel generator.

Paper, Fig. 1. Standard pixel prediction and semantic routing (a–d), training loss and samples (e–f), and an analytic two-spiral illustration (g–j). The toy distribution shows its strongest intrinsic semantic conflict at intermediate timesteps. WiT improves recovery of the target manifold.
The role of a waypoint. It is a spatial routing condition recomputed from the evolving noisy state. Semantic prediction and pixel prediction have distinct roles within the architecture; the actual generative trajectory stays in pixel space.

Method

Predict structure. Modulate pixels. Repeat.

WiT separates semantic representation prediction from high-dimensional image prediction, with a dynamically updated connection between them.

01

Build compact semantic targets

Frozen DINOv3-S/16 features supply dense spatial structure. PCA, fitted on 50,000 ImageNet training images, projects each patch feature to 64 dimensions. Semantic organization comes from the pretrained encoder; PCA keeps the prediction target compact.

02

Infer the current waypoint

A 21M-parameter ViT-S/16 Waypoints Predictor takes the current noisy pixels, timestep, and class condition and predicts a clean semantic waypoint. Its weights are frozen for image generation, but its prediction changes as the pixel state evolves.

03

Guide clean-pixel prediction

Just-Pixel AdaLN converts each semantic token into spatially varying affine modulation of the generator's attention and MLP branches. The native pixel-token sequence is preserved, and the generator directly predicts the clean RGB image.

Paper, Fig. 2. The Waypoints Predictor infers a compact semantic representation from the current noisy pixel state. The Pixel Space Generator uses it as a spatial condition and directly predicts the clean image.
Paper, Fig. 3. Just-Pixel AdaLN supplies spatially varying modulation (a). As noisy pixel states evolve, the predicted semantic waypoints are recomputed throughout inference (b).

Training: learn semantics, then pixels

Train the Waypoints Predictor for 600 epochs, freeze its EMA weights, then train the Pixel Space Generator with predicted waypoints as conditions. The generator follows JiT's clean-image prediction with a velocity-space loss. The same frozen predictor is reused across the reported experiments.

Inference: one evolving pixel state

At each vector-field evaluation, predict the waypoint from the current noisy state, condition the pixel generator, and obtain a velocity update from its clean-image prediction. Sampling uses a 50-step Heun solver. No separate semantic flow or latent image decoder is sampled.

Algorithm 1. Training Procedure of WiT. Two-stage training: learn the semantic predictor, then freeze it to condition pixel-space generation.
Algorithm 2. Inference Procedure of WiT. Dynamic waypoint prediction and Just-Pixel AdaLN guide each pixel-space ODE evaluation.

ImageNet results

Gains across scale, training duration, and resolution

Semantic routing consistently improves FID over the corresponding JiT baselines, from Base to Huge generators and from 256×256 to 512×512 images.

WiT-H/16 · 256×256

1.79 FID

Scaling to Huge

At 600 generator epochs, with IS 302.3. JiT-H/16 reaches 1.86 FID; the larger, 2B-parameter JiT-G/16 reaches 1.82.

WiT-L/16 · 256×256

2.02 FID

Stronger Large-scale results

At 600 generator epochs, versus JiT-L/16's 2.36. At 200 epochs, WiT-L/16 reaches 2.28 versus JiT-L/16's 2.79.

WiT-B/32 · 512×512

3.38 FID

Extending to higher resolution

At 600 generator epochs, versus JiT-B/32's 4.02. The improvement also holds at 200 epochs: 3.89 versus 4.64.

Reading the comparisons. FID-50K is lower-is-better; IS is higher-is-better. WiT and direct JiT baselines use the 50-step Heun solver. Epochs count Pixel Space Generator training only: WiT additionally uses a 21M predictor pretrained for 600 epochs, then frozen and reused. These are comparisons at matched generator training durations, not matched total training cost.

ImageNet 256×256: direct baselines and WiT

ImageNet 256 by 256 direct comparisons, paper Table 2
MethodParamsEpochsIS ↑FID-50K ↓
JiT-B/16131M200-4.37
WiT-B/16131M + 21M200270.73.34
JiT-B/16131M600275.13.66
WiT-B/16131M + 21M600280.23.03
JiT-L/16459M200-2.79
LF-DiT-L/16465M200-2.48
WiT-L/16459M + 21M200276.62.28
JiT-L/16459M600298.52.36
WiT-L/16459M + 21M600293.32.02
WiT-XL/16676M + 21M200288.92.16
WiT-XL/16676M + 21M600301.01.89
JiT-H/16953M600303.41.86
JiT-G/162B600292.61.82
WiT-H/16953M + 21M600302.31.79

Paper, Table 2. WiT parameter counts show the pixel generator plus the 21M Waypoints Predictor. A dash denotes a value not reported in the table.

ImageNet 512×512: the gains carry over

ImageNet 512 by 512 FID, paper Table 3
MethodFID ↓ · 200 epochsFID ↓ · 600 epochs
JiT-B/324.644.02
WiT-B/323.893.38

Paper, Table 3. B/32 models generate 512×512 images with 32×32 pixel patches.

Broader comparisons from the paper

The remaining entries of Table 2, on ImageNet 256×256. These results come from the respective publications and use different architectures and training setups; they are not controlled compute-matched comparisons.

Broader ImageNet 256 by 256 comparisons, paper Table 2
MethodParamsEpochsIS ↑FID-50K ↓
Latent-space diffusion models
DiT-XL/2675M + 49M-278.22.27
SiT-XL/2675M + 49M-277.52.06
REPA (SiT-XL/2)675M + 49M-305.71.42
LightningDiT-XL/2675M + 49M-295.31.35
DDT-XL/2675M + 49M-310.61.26
RAE (DiTDH-XL/2)839M + 415M-262.61.13
Pixel-space models (non-diffusion)
JetFormer2.8B--6.64
FractalMAR-H848M-348.96.15
Pixel-space diffusion models
ADM-G554M-186.74.59
RIN410M-182.03.42
SiD (UViT/2)2B-256.32.44
PixelFlow (XL/4)677M-282.11.98
PixNerd (XL/16)700M-297.02.15

Understanding trajectory conflict

From local semantic mixing to gradient interference

Three complementary analyses examine whether target-relevant semantic routing improves manifold recovery, organizes local neighborhoods, and reduces conflicting learning signals in the shared generator.

01 · Controlled toy experiment

Target-relevant routes improve manifold recovery

On two-spiral data embedded in 4, 8, 16, 32, 64, and 128 dimensions, WiT lowers normalized nearest-manifold RMSE at every evaluated dimension. Relative reductions versus the jointly optimized JiT baseline are 7.0%, 11.5%, 9.3%, 21.7%, 15.9%, and 14.1%, respectively. A parameter-matched Random Route baseline with irrelevant routing labels does not yield comparable, consistently positive gains.

Paper, Fig. 6. Samples are shown in a 2D projection. Colors and nRMSE measure distances in the original observation space, normalized by dimension. Lower nRMSE means better manifold recovery; red percentages are WiT reductions relative to JiT.

02 · Representation-level evidence

More coherent semantic neighborhoods

Normalized Local Conditional Variance (NLCV) measures the dispersion of ground-truth semantic endpoints within local neighborhoods. Holding the ImageNet class fixed, neighborhoods formed from predicted waypoints have 23.8% to 35.6% lower semantic dispersion than those formed from noisy pixel observations.

Within-condition NLCV, lower is better, paper Table 7
Timestep tRaw statePredicted semanticReduction
0.160.97130.740123.8%
0.310.96680.657832.0%
0.510.96360.620835.6%

Paper, Table 7. Each group contains an anchor and four same-class nearest neighbors in a 64D PCA-whitened query space. Raw queries use patch-mean noisy pixels; semantic queries use mean-pooled predicted waypoint tokens. NLCV is normalized by within-class semantic variance and macro-averaged across classes.

03 · Optimization-level evidence

Less disagreement-dependent gradient interference

Sample pairs are matched by class condition, timestep, and noisy-state distance, then separated into high and low semantic-target disagreement groups. Greater disagreement is associated with more output-layer gradient opposition in both models, but the effect is weaker in WiT.

State-matched output-layer gradient interference, paper Table 8
High-minus-low effectJiTWiTRelative mitigation
Gradient cosine−0.15961−0.1367714.3%
Negative gradient mass0.074340.0583321.5%

Paper, Table 8. Less-negative cosine and a smaller increase in negative mass indicate weaker disagreement-dependent interference. Absolute mitigation: 0.02284 [0.00200, 0.04391] and 0.01601 [0.00451, 0.02730], respectively. Brackets are the paper's confidence intervals, estimated with 10,000 ImageNet-class cluster bootstrap samples.

What the evidence supports. Predicted waypoints reorganize the representation presented to a finite-capacity pixel generator and reduce disagreement-dependent learning interference. These measurements support semantic routing as an aid to optimization; they do not establish that geometric path intersections disappear or that a separate semantic trajectory is generated.

Design choices

Why compact waypoints and spatial modulation?

Ablations compare representation guidance, source encoders, waypoint dimension, and injection strategy within the evaluated ImageNet settings.

Representation-guided alternatives

Representation guidance at 200 generator epochs, paper Table 4
JiT-B/16 variantFID ↓
JiT baseline4.37
+ REPA5.14
+ RCG4.96
+ PixelREPA4.00
WiT-B/163.34

Paper, Table 4. Same JiT-B/16 backbone, 200 generator epochs. Results describe these implementations under the evaluated setting.

Choice of pretrained representation

Waypoint source with B/16 at 600 generator epochs, paper Table 5
Waypoint sourceFID ↓
JiT (no waypoint)3.66
MoCoV3-S/163.11
MAE-B/163.32
DINOv2-S/143.36
DINOv3-S/16 (default)3.03

Paper, Table 5. B/16 generator, 600 generator epochs. Every evaluated encoder improves over JiT; DINOv3-S/16 gives the lowest FID.

Waypoint dimension and injection mechanism

WiT-B/16 waypoint and injection ablation at 200 epochs, paper Table 6
PCA dimensionInjectionIS ↑FID ↓
32Just-Pixel AdaLN210.405.11
128Just-Pixel AdaLN211.334.12
64Channel Concat221.193.93
64In-context Concat238.923.63
64 (WiT default)Just-Pixel AdaLN270.733.34

Paper, Table 6. WiT-B/16, 200 generator epochs. Among tested configurations, 64D waypoints balance compactness and retained structure, while Just-Pixel AdaLN outperforms channel and token concatenation.

Guidance behavior (paper, Fig. 5). The minimum-FID CFG scale shifts from 3.8 for WiT-B/16 at 200 epochs to 3.1 at 600 epochs and 2.9 for WiT-L/16 at 600 epochs. Excessively strong guidance degrades FID in these sweeps.

Repository Status

Training and inference code will be released here

This repository is the official landing point for WiT. The current public state is the project page plus a placeholder code entry. Training and inference code will be released in this same repository after the paper is formally published.

Acknowledgments

Support and thanks

We thank Qiming Hu for insightful discussions and feedback. This work was partially supported by computational resources from TPU Research Cloud (TRC).

Citation

BibTeX

@article{wang2026wit,
  title={WiT: Waypoint Diffusion Transformers for Alleviating Trajectory Conflict in Pixel-Space Image Generation},
  author={Wang, Hainuo and Li, Mingjia and Guo, Xiaojie},
  journal={arXiv preprint arXiv:2603.15132},
  year={2026}
}