Not today, but Interesting idea. The main motivation was a drop-in for an existing Jev setup, so the only teacher right now is Jev and the audit measures agreement with Jev. A correction would have to become a second label source that overrides Jev's for that input.
The head is a multinomial logistic regression: one linear layer plus softmax on top of a frozen sentence-embedding model (bge-small by default, swappable). That head is the entire local model, the encoder is off the shelf and never changes.
Yeah, I kinda figured that head was the whole model, the way you phrased it. scikit-learn?