This why those RL env startups are able to charge frontier labs so much for their work. LLMs still generalize poorly outside of self-verifiable tasks like coding and math.
Labs have to compensate with post-training in RL env that embeds these expertise well, which is non-trivial both in terms of domain knowledge and technical expertise.