Knowing who said what
paperite works out how many people spoke in a recording and which parts belong to each of them, then labels the transcript accordingly. This runs on your Mac using Apple's Neural Engine. The audio is not sent anywhere to do it.
How does paperite tell voices apart?
It analyses the voice characteristics across the whole recording and clusters them into speakers, then assigns each part of the transcript to a cluster. The result is a transcript where every line carries a speaker and a timestamp rather than a single undifferentiated block of text.
What if it gets the number of speakers wrong?
You correct it. The recording panel shows how many speakers were detected with a control to raise or lower that number, and re-processing reassigns the transcript to the corrected count. This matters in practice: a quiet participant or a noisy room can merge two people into one.
Can I give the speakers names?
Yes. Each detected speaker can be renamed in place, and the name propagates through the transcript. paperite also shows how much of the conversation each person accounted for, which is a faster way to identify who is who than reading from the top.
Does any of this need an internet connection?
No. Speaker separation runs entirely on the machine, like the recording and the transcription before it.