Is this an example of the bitter lesson? Could diffusion models not be this good with equal scaling/compute as LLMs?
Well the LLMs are much bigger than diffusion models so I think that's the bitter lesson. You could scale compute for diffusion models though.
Well the LLMs are much bigger than diffusion models so I think that's the bitter lesson. You could scale compute for diffusion models though.