Speaker Diarization and Transcription
A reusable audio pipeline that answers who spoke when and what they said.
Combined speaker embeddings, clustering, diarization, and speech recognition into a reusable Python pipeline. The work moved the process away from a single notebook and toward a batch job that could continue after interruption.
Long audio jobs fail for ordinary reasons: a session ends, a GPU run stops, or one file needs to be checked again. The checkpointed harness was built around that reality.
The pipeline
Audio was represented by speaker embeddings, grouped into speaker clusters, aligned into speaker segments, and then passed to speech recognition. The result was a timeline that connected a speaker with the words spoken in that part of the recording.
The useful engineering part
The reusable class and checkpointed batch runner made it possible to process many files, resume a partial run, and work with Darija audio instead of treating each experiment as a fresh manual session.
selected tools
- Python
- Resemblyzer
- pyannote
- spectralcluster
- DeepSpeech
- GPU