Wow! TIL! I've been running a for loop around the two ordering variations to catch the winner of each turn and the difference is quite noticeable. In the options-after-body case in 47 of 100 attempts it classifies as phishing, whereas in the options-before-body case it classifies clearly as rickroll (94 out of 100 attempts)
Payroll sends you an email with a link to a Youtube video that plays a song.
Options after body:
Average probabilities:
Rickroll 0.5158 ( 51 wins)
Phishing 0.4561 ( 47 wins)
Spam 0.0281 ( 2 wins)
Joke 0.0000 ( 0 wins)
Legitimate 0.0000 ( 0 wins)
Options before body: Average probabilities:
Rickroll 0.9293 ( 94 wins)
Joke 0.0549 ( 5 wins)
Phishing 0.0140 ( 1 wins)
Spam 0.0018 ( 0 wins)
Legitimate 0.0000 ( 0 wins)
This was Gemma4-26B-A4B-NVFP4 by the way.EDIT
Gemma4-12B-it-NVFP4 seems way less sensitive to option/body ordering:
Options after body:
Average probabilities:
Rickroll 0.9867 ( 99 wins)
Phishing 0.0133 ( 1 wins)
Joke 0.0000 ( 0 wins)
Spam 0.0000 ( 0 wins)
Legitimate 0.0000 ( 0 wins)
Options before body: Average probabilities:
Rickroll 0.9401 ( 93 wins)
Phishing 0.0336 ( 3 wins)
Spam 0.0250 ( 4 wins)
Joke 0.0010 ( 0 wins)
Legitimate 0.0002 ( 0 wins)
Anyway, this for-looping stuff doing 100 calls to even a local VLLM API takes around 5 seconds in total, so this isn't anywhere close to sub-second Jev territory.
Yah, that's what I use: https://rcarmo.github.io/projects/go-system-one uses Gemma, and that's partly why. Seems less prone to getting distracted with ordering.