You can do this on x86 as well at a cost a merely tens to hundreds (possibly lots of hundreds) of thousands of cycles. This is part of why x86 is so popular in the embedded space.
(I’m being sarcastic, obviously. x86 interrupts and interrupt returns are hilariously slow. FRED may improve this by quite a bit.)
What's the technical reason for them being slow? Book keeping with caches or something?
But you’re also comparing dozens of cycles at 10mhz to 100ks at 5ghz. That’s probably comparable in terms of wall clock, no?