What evals did you use to compare that to using an existing language with a large corpus of training data?