This is all quantifiable. I regularly have my model run benchmarks against all the config permutations and then choose the best based on my criteria, which typically boil down to trading prefill and decode times
What model has to trade between those? I have both. You just have different, independently-optimized forward passes for each.
What model has to trade between those? I have both. You just have different, independently-optimized forward passes for each.