logoalt Hacker News

Training Text-to-Image Models Without a VAE

43 points • by schopra909 • last Tuesday at 8:24 PM • 14 comments • view on HN

Comments

schopra909 • last Tuesday at 8:31 PM

Hi HN, author here!

For context, we're a 2-person lab training generative video models. Goal is a new set of controllable, animation tools (you can read more about that here https://www.linum.ai/about if you're curious).

The biggest bottleneck for our last text-to-video model in terms of training and inference cost is attention. Video models are incredibly token dense (e.g. 110K tokens for a several second clip). If we can condense that context window more aggressively, we can train bigger models for a lot less $$ and offer them to prosumers at reasonable price points (unlike the big models today like Seedance, which cost an arm and a leg to run).

Traditionally, image and video models have two disjoint components: VAE (Variational Autoencoder) and Diffusion Transformer (DiT). They're trained separately, and empirically VAEs seems to struggle to get past 16x16 token reduction.

Here, we're switching to pixel-space, throwing away the VAE, and achieving 32x32 token reduction (4x smaller context windows) while learning a better overall model in a fraction of the training samples.

The central thesis is "simpler is better". If we can put the compression problem into the more powerful Diffusion Transformer (DiT) would should be able to learn a "latent space" optimized for generation and get better compression without hurting generation quality.

I'll be checking this post off and on the next couple of hours, so feel free to drop questions below. And I'll try to answer them to the best of my ability.

P.S. The model checkpoints from this blog are Apache 2.0, so feel to try playing with it yourself on a GPU!

➕ show 1 reply
vunderba • today at 8:00 PM

Neat. What are your thoughts on the recently released Iris-3B, a pixel-space model which also bypasses the need for a traditional VAE?

https://arxiv.org/pdf/2610.09450

https://huggingface.co/speridlabs/iris-3b

➕ show 1 reply
joefourier • today at 7:29 PM

I wonder how much of that is due to the use of a suboptimal VAE? The same optimisations could be applied to it, and my intuition tells me that your total compute spend would be even more optimal if you retrained the VAE yourself with a better approach (esp. ensuring translation and rotation invariance, + ability to rescale/blur the latents).

You could also have the same advantages of a draft in low resolution latent space, with an easier to learn data distribution that's more robust to perturbation, but instead of 512x512, you could get a 4096x4096 output for the same compute (assuming a 8x VAE).

There's also the advantage of being able to use a high number of diffusion/flow matching steps for the main LDM while the VAE can be single step and much smaller, since it does not have to handle language or significant scene understanding, just perceptual compression. This sounds especially important for a video model where I would be extremely hesitant to train a generative model without relying on interframe compression.

➕ show 2 replies
bitpush • today at 5:44 PM

Is there a way to use this in ComfyUI today?

➕ show 1 reply