I tried something similar by putting together a "global ledger" of speaker identities. This was more happening during the call than having a predefined one but I just couldn't get it to work properly. The issue being that as soon as some speech gets assigned to a new label or an incorrect one, everything ends up getting misaligned and gets messy quickly.
I might take another look into doing it in a different way that gradually builds up from successful calls, I just need to think of how to do this in a simple(ish) way for non-technical users and a way that still works well enough on low to mid tier laptops.
I know this is hn and not a product development meeting, but:
If you’re targeting data for people to use on the same call then requires more intense work, while if you’re targeting data for people to use after or on an ongoing basis (eg an established business meeting with staff/vendors/etc) then having a predefined one seems good. Same interface as a CRM: picture, some audio clips for users to choose from, and confidence scores on each section of the recording for cases where people sound similar or are talking over each other. Over time the diarization gets better as users accept/reject/tag samples and that work helps them feel more aligned with the tool.