The approach is cool; but the results feel ~sus~.
GenEval and GenEval2 prompts are very terse (~10 tokens) while these models are primarily trained on long-dense captions. So, it's pretty common to upsample prompts before running them through benchmarks. That way you're testing the model as it was trained, rather than testing it on prompts that are out of distribution.
If you look at the Qwen-Image Report before they enhanced it for their 12-25 release, upsampled prompts score 0.87 on GenEval* In this paper they take it from 0.74 to 0.81.
To me that reads to me that they're basically getting the model to adapt to the benchmarks' terse prompts. And it's unclear to me if simply finetuning the model on shorter prompts would work just as well as the RL solution.