I'm very interested in Diffusion text models. The concept of taking noise and adding words starting randomly all over the response, and filling in the noise from there on breaks my brain.
I'm sure I have a fundamental misunderstanding of the technology, though.
Denoising is probably the weakest part of this model. There is recent research from Kaiming He showing that predicting the noise is actually not the best strategy for image generation since noise space is so large.
Simply predicting the surface of the data you are trying to generate is far more representationally efficient, and perhaps in the next few months we'll see a version of that for LLMs
How does that break your brain? It's how basically every human writes and iterates on text..?
DiffusionGemma goes one step further even, and does this denoising over multiple "canvases" which lets it do reasoning and separate out a "final reply" canvas, looks something like this: https://gist.github.com/embedding-shapes/f4cb46bad704b6d0168...
Diffusion text models for me is the more interesting type of LLMs for local usage, as it really makes good use of single GPUs for single responses, rather than auto-regressive ones, and is a lot faster! Probably the fastest model I've been able to run so far, ending up doing ~670 tok/s (depending on the type of text) on a Pro 6000