I'd be interested to see if using DiffusionGemma-as-Jev helps as you can feed the image directly into the model and it'll make decisions based on the image embeddings.
not sure how this is innovative they show the System-1 model can play Doom right in the announcement [1] :
>Doom >We love how this doomo doomonstrates real-time intelligence and what can be doone with code + AI. The engineer behind it was worried about making 10 queries a second (which ends up costing ~$7/hour), but the rest of us agreed that was lower than expected! This is so fun we intend to not only release an in-depth walkthrough, but also host some events to hack on this.
[1] https://typesafe.ai/blog/introducing-system-one-models-and-j...
How does it do on OSWorld-verified? Recently read that even Fable 5 is just at 85% .
I did not understand what this is all about. Anyone with more brain than me can explain please?
Super cool. Hope waitlist will move soon. I have a use case for it too.
are you the author? If so - what are your notes on using Jev in this scenario?
It seems like the more honest comparison would be to OCR the screen and send that as input to the LLM?
This is only tangentially related but is Jev trained on the same kind of data as the rest of the LLM world? (i.e. unethical)
At a glance it addresses two out of three of my "load bearing points" against AI, which is the cost to run the things, and that they can be used to generate slop.
If it was trained ethically that would "close the gap".
[dead]
I trailed off a few lines into the README. No human ever edited any of this. « LLM detected, project rejected ».