What you describe sounds a lot like the HTML/CSS split, with having a separate semantic/display data. I think that's a rather large jump from the 80x25 model, retrofitting that much data sounds like quite the challange.
Not sure how you're planning on doing that, can you incrementally extend the current model, or will it require a completely new protocol?