The model was trained on CC0, "royalty-free" music, so the outputs are going to sound a bit bland and generic. Model is not currently possible to finetune without the semantic tokenizer.