Because of masked attention in LLMs, if you put the options before the body (the email to analyze), the transformer already knows what it needs to look for, and can use more tokens to create state to address that specific task (BERT has no mask in the attention, so tokens attend also to next tokens). You could also do a few examples in the system prompt to improve calibration.
Another trick that works is to repeat the question two times: "I'm repeating the task and labels for clarity: ..."
What a time to be alive, repeating questions to a model twice to increase accuracy.
This works very very well :).
Prompt Repetition Improves Non-Reasoning LLMs: https://arxiv.org/abs/2512.14982
Wow! TIL! I've been running a for loop around the two ordering variations to catch the winner of each turn and the difference is quite noticeable. In the options-after-body case in 47 of 100 attempts it classifies as phishing, whereas in the options-before-body case it classifies clearly as rickroll (94 out of 100 attempts)
Payroll sends you an email with a link to a Youtube video that plays a song.
Options after body:
Options before body: This was Gemma4-26B-A4B-NVFP4 by the way.EDIT
Gemma4-12B-it-NVFP4 seems way less sensitive to option/body ordering:
Options after body:
Options before body: Anyway, this for-looping stuff doing 100 calls to even a local VLLM API takes around 5 seconds in total, so this isn't anywhere close to sub-second Jev territory.