Im running this on my 4070ti super (16 gb vram), and it takes 10 minutes for a 10-seconds 480p video. but the results are spectacular.
I am particularly curious how multimodal models will work with types of knowledge that are inherently non-text. For example, SOTA LLMs really suck at electronics, especially analog electronics.
Is MiniMax H3 capable of logical / technical reasoning, or is it purely art oriented?
If you wouldn't mind sharing, what's your Comfy workflow for this? I have the same video card setup and would like to give it a shot.
1 minute for a second of footage, that is awesome! Thanks for sharing.
Huh, interesting. I tried to generate a 10 second 1080p clip on a bigger machine and the results were quite poor. Unusable for anything, in fact.
How much RAM does your machine have?
FWIW on a 5080 16GB it takes 3 minutes for 10 seconds 480p video (the mouse video workflow with length changed from 5 seconds to 10 seconds)