Defects Missed in Transcription — AI Speaks After 0.5-Second Silence
DEV Community
Defects Missed in Transcription — AI Speaks After 0.5-Second Silence
When inspecting the quality of a speech synthesis model using only STT, we missed all added sounds lasting just 0.1 seconds after silence. I’ll explain how we built a detector using waveform envelopes and how all 12 models produced false positives due to commas.
0 comments
No comments yet.