There are two questions about that stability I have.
One, things like catastrophic forgetting and falling into incoherence.
Two, less likely but far more worrying, falling into unwanted attractor states. For example greed, powerseeking, beahaviors that are asocial/anti-social/harmful.
Aren't such attractors also problems during training? Presumably alignment constraints would need to apply to continuous learning as well.