logoalt Hacker News

rahimnathwani • yesterday at 11:10 PM • 1 reply • view on HN

"Using embeddings + logistic classifier, the architecture matches or beats Jev and Laya in all basic classification"

Have I understood correctly that you trained only the logistic classifier, but didn't need to train the embedding model?

If so, I'm curious whether you compared that approach (A) with:

B) Jev only, with a single output.

C) Jev with multiple outputs fed into a logistic classifier.

Obviously C has cons (can't be self-hosted, needs some up-front work on deciding the shape of the output) but it might be somewhat more interpretable. (And I suppose it might have better performance?)


Replies

nico • yesterday at 11:46 PM

You are correct, I didn’t train the embeddings model

Here's a gist with code you can use to test the Banking77 dataset: https://gist.github.com/nicobrenner/056a5aaff5d0119c0032ecda...

The gist uses BAAI/bge-large-en-v1.5, which is 1.2GB approx. You can replace it for all-MiniLM-L6-v2 (91 MB @ fp32 or 45 MB quantized fp16) small enough for mobile/edge. With all-MiniLM-L6-v2 it still gets 93.0% on Banking77, only 1.3 points behind bge-large at 15x smaller

I haven’t compared different ways of sending requests to Jev

The data to train the classifiers comes from the datasets used to test them (not from Jev)