I think it’s about sample efficiency. You could finetune your own jev using Lora with very little data