Diffusion language models work with a discrete output space, unlike image models that repeatedly refine a continuous output, so they don't do the noise-prediction thing anyway.
Ok, you are correct. These models don't train noise predictors at all unlike first gen image diffusion
Ok, you are correct. These models don't train noise predictors at all unlike first gen image diffusion