Score: 0

Resnet-conformer network with shared weights and attention mechanism for sound event localization, detection, and distance estimation

Published: July 23, 2025 | arXiv ID: 2507.17941v1

By: Quoc Thinh Vo, David Han

Potential Business Impact:

Helps computers pinpoint sounds in noisy places.

Business Areas:

Audio Media and Entertainment, Music and Audio

This technical report outlines our approach to Task 3A of the Detection and Classification of Acoustic Scenes and Events (DCASE) 2024, focusing on Sound Event Localization and Detection (SELD). SELD provides valuable insights by estimating sound event localization and detection, aiding in various machine cognition tasks such as environmental inference, navigation, and other sound localization-related applications. This year's challenge evaluates models using either audio-only (Track A) or audiovisual (Track B) inputs on annotated recordings of real sound scenes. A notable change this year is the introduction of distance estimation, with evaluation metrics adjusted accordingly for a comprehensive assessment. Our submission is for Task A of the Challenge, which focuses on the audio-only track. Our approach utilizes log-mel spectrograms, intensity vectors, and employs multiple data augmentations. We proposed an EINV2-based [1] network architecture, achieving improved results: an F-score of 40.2%, Angular Error (DOA) of 17.7 degrees, and Relative Distance Error (RDE) of 0.32 on the test set of the Development Dataset [2 ,3].

A Robust framework for sound event localization and detection on real recordings

Sound

Finds sounds and where they come from.

16 Dec 2025 0

91%

Spatial and Semantic Embedding Integration for Stereo Sound Event Localization and Detection in Regular Videos

Audio and Speech Processing

Finds sounds and their direction in videos.

7 Jul 2025 0

91%

Integrating Spatial and Semantic Embeddings for Stereo Sound Event Localization in Videos

Audio and Speech Processing

Helps computers understand sounds and sights together.

8 Sep 2025 0

View PDF Login to Bookmark

Country of Origin

🇺🇸 United States

Page Count

4 pages

Resnet-conformer network with shared weights and attention mechanism for sound event localization, detection, and distance estimation

Helps computers pinpoint sounds in noisy places.

Technical Abstract

A Robust framework for sound event localization and detection on real recordings

Spatial and Semantic Embedding Integration for Stereo Sound Event Localization and Detection in Regular Videos

Integrating Spatial and Semantic Embeddings for Stereo Sound Event Localization in Videos