I think the main limit of their model is that it was prompted and with LLMs you get what you prompt for. It's as truthful as any complex models. Could be good, could be bad, could be meh.