Yes, 3.6x cheaper when it comes to training!
When it comes to inference it should also be cheaper (since we’ve cut attention sequence length by 4x), but I don’t have a hard number for you how much cheaper for inference