I re-implemented this one for macOS over the last couple of months: https://github.com/mmastrac/diffgemma
I like the model a lot and it's fairly good at reasoning. You can also really bend it to your needs. It's designed for machines with more compute than memory bandwidth but IMO does really well on metal.
I've got it up to ~15tok/s on M3-class machines, but I wager there's a bunch of perf on M5 that I just don't have hardware access to unlock.
I tried to implement MTP using the other Gemma MTP heads but I failed to move that perf needle. There's some interesting research to be done about pre-seeding the diffusion canvas from draft models. DiffusionGemma with the right drafter can hit 20-30 tok/s on my machine, but I've been unable to combine the two together to make it faster than what it's been running at so far.
If these models get good at coding it's going to force a rethink of how languages, compilers and test suite runners work. "AI changes everything" is a cliché by this point but I think it's actually true.
If your model can reason and write code at 1500 toks/sec, then you should end up totally bottlenecked on CPU time the entire time a prompt is active. If you aren't, then you're losing wall time versus competitors. But our whole development stack is based around the idea that programmers spend most of their time thinking, talking and coding, not waiting for the CPU (melting CI clusters being a painful exception to that).
What I'm imagining here is some sort of hybrid mode in which compiling code and running it through an interpreter can be overlapped, so an LLM can propose a change and immediately begin running unit tests while type errors that might affect some other module are found in parallel. And the tests would always run sharded, potentially on a remote cluster, even in local dev.
Obviously this approach is to some extent what made the JVM popular. Java compiles very fast because javac does little more than type checking, and the type system is simple. Then the JVM does the heavy lifting of compilation in parallel with it running. So although Java has a reputation for poor startup times, turnaround times for the JVM can be really excellent compared to something like C++, Swift or Rust. And a lot of startup time pain is just poor frameworks like old Springs that want to reflectively scan the app's files and do other inefficient stuff. More modern frameworks push more to incremental build tasks and can get startup down to very little, <0.5secs for a web server with DB connections to start for instance.
But it feels like this approach should be pushed much further. The model should spend all its time waiting on unit tests to run.
I'm very interested in Diffusion text models. The concept of taking noise and adding words starting randomly all over the response, and filling in the noise from there on breaks my brain.
I'm sure I have a fundamental misunderstanding of the technology, though.
Appealing results... do we think there is scope to close the accuracy gap against AR models? or even leverage the "Bidirectional Reasoning and Self-Correction" into an overall advantage?
there's still JEPA to be integrated before AGI.
Would DiffusionGemma be suitable candidate for DFlash 2?
[flagged]
Just wanted to share this, I found it was a really nice resource to understand how diffusion Gemma worked: https://newsletter.maartengrootendorst.com/p/a-visual-guide-...
The really interesting thing to me was that they didn’t need to train this model from scratch they just used their existing MOE checkpoint:
“To convert a decoder-only model (Gemma 4 26B A4B) into a denoiser, we can make use of something it is not directly using when generating tokens, namely the logits of all tokens!”
What makes me hopeful about this release is that possibly this same conversion can be applied to other open models and we might see a bunch of diffusion versions of existing local models. It’s exciting stuff!