logoalt Hacker News

ericdtoday at 4:26 PM1 replyview on HN

Was recently thinking about doing something similar, do you basically just have the qwens describe what they see for the flashes?

Was considering adding a LoRa/vision head to Flash, but seems like it could take a while to get it right.

If DSv4 Flash was multimodal, I’d probably be done model shopping for a while


Replies

arjietoday at 4:52 PM

Same, with a multimodal DSv4 Flash I would just stop paying attention to things. Very smart, and at 260 tok/s it's too fast to care about anything else. If you ever graft something like that I would love to hear about it.

Yes, I have a very dumb flow. The harness has a describe_image tool that takes an image and a prompt and so DSv4 Flash uses it to get an idea of what it's looking at.

show 1 reply