Meyti Eka Apriyani, Moch. Nadi Rafli Maulana, Elok Nur Hamdana
WebRTC-based real-time communication is highly vulnerable to background noise, which negatively affects both perceptual speech quality and intelligibility. Traditional suppression techniques, such as spectral subtraction and Wiener filtering, are computationally efficient but perform poorly in non-stationary environments. To address this gap, this paper proposes a client-side implementation of the Dual-Signal Transformation LSTM Network (DTLN) within a WebRTC framework. Unlike previous server-side or offline approaches, the system is deployed entirely in the browser using ONNX Runtime-Web and Web Workers, ensuring low-latency and privacy-preserving processing. The novelty of this work lies in its real-time, browser-based, and multilingual evaluation, covering English and Indonesian datasets. The system was tested across 1,360 audio mixtures combining clean speech with 17 noise types from six categories and four SNR levels. Results demonstrated significant improvements in perceptual quality (average PESQ = 1.656), particularly in stationary noise such as transportation (PESQ = 1.877), while intelligibility gains were more modest (average STOI = 0.537). Statistical analysis confirmed the significance of improvements, and a small-scale Mean Opinion Score (MOS) test with bilingual listeners validated perceptual quality results. These findings underscore the potential of browser-based noise suppression while highlighting current limitations and avenues for improvement through hybrid architectures and multilingual optimization. © 2025 IEEE.
State Polytechnic of Malang, Dept. of Information Technology, Malang, Indonesia