That isn't what they did here. They took the output of the LLM as one feature, then added 17 other features, then piped it into a crappy model and got a 3% improvement.