Mid Sweden University

miun.sePublications
Change search
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf
CNN Models for Microphone Array Covariance Matrix Upsampling and Acoustic Imaging
Mid Sweden University, Faculty of Science, Technology and Media, Department of Computer and Electrical Engineering (2023-).ORCID iD: 0000-0001-7410-0483
Tampere University. (Audio Research Group)ORCID iD: 0009-0009-3751-6469
Tampere University. (Audio Research Group)ORCID iD: 0000-0002-1041-0498
Mid Sweden University, Faculty of Science, Technology and Media, Department of Computer and Electrical Engineering (2023-).ORCID iD: 0000-0002-8253-7535
Show others and affiliations
2026 (English)In: 2026 IEEE International Symposium on Artificial Intelligence for Instrumentation and Measurement (AI4IM), IEEE, 2026Conference paper, Published paper (Refereed)
Abstract [en]

Acoustic imaging visualization is a core methodology in acoustics, enabling spatial analysis of sound sources and acoustic scenes. However, limited sensor availability in practical systems motivate approaches that enhance spatial resolution without increasing the hardware complexity. In this paper, we focus on upsampling virtually a tetrahedral 4-microphone array to a spherical 32-microphone array by estimating the covariance matrices of the channels employing deep learning techniques. Five neural network architectures are investigated for covariance upsampling for acoustic imaging using the real-world STARSS23 dataset. These models are developed to estimate a 32-microphone, time–frequency covariance matrix from a 4-microphone input covariance representation. The proposed architectures are based on 2D convolutional layers to capture the underlying spatial–spectral structure of covariance matrices, and are further enhanced with frequency dynamic convolution to model their frequency-dependent properties. The proposed architectures are evaluated in terms of root mean square error (RMSE) and using delay-and-sum beamforming acoustic imaging. Quantitative results show that all models outperform a random-guess baseline, which yields an RMSE of 0.548, with the best-performing architecture achieving an RMSE of 0.432. We analyze qualitatively the performance of the proposed models through beamforming heatmap visualizations derived from the 4-channel input covariance, the 32-channel ground truth, and the predicted 32-channel covariance matrices. These results demonstrate that covariance upsampling significantly enhances the effective performance of the 4-channel microphone array, producing sound maps that closely resemble those obtained with the 32-channel array.

Place, publisher, year, edition, pages
IEEE, 2026.
Keywords [en]
Acoustic imaging, microphone array upsampling, spherical microphone arrays, sparse microphone arrays, spatial covariance matrix, beamforming, deep learning, convolutional neural networks
National Category
Electrical Engineering, Electronic Engineering, Information Engineering
Identifiers
URN: urn:nbn:se:miun:diva-57943DOI: 10.1109/AI4IM69129.2026.11558232ISBN: 979-8-3315-5176-6 (print)ISBN: 979-8-3315-5175-9 (electronic)OAI: oai:DiVA.org:miun-57943DiVA, id: diva2:2080496
Conference
2026 IEEE International Symposium on Artificial Intelligence for Instrumentation and Measurement (AI4IM), Amalfi, Italy, 21-23 May, 2026
Available from: 2026-06-26 Created: 2026-06-26 Last updated: 2026-07-06Bibliographically approved

Open Access in DiVA

fulltext(4417 kB)66 downloads
File information
File name FULLTEXT02.pdfFile size 4417 kBChecksum SHA-512
40760dd19f2215b8d75a01b00274f5afd330f27362477c8333e59b32bcb9207747496f09f9a48a3cd2cc3638f73ca633f7878204c8083e94ba83fc78beb463f6
Type fulltextMimetype application/pdf

Other links

Publisher's full text

Authority records

Adamopoulou-Soulantika, MarianthiJiang, MengSeyed Jalaleddin, MousaviradLundgren, Jan

Search in DiVA

By author/editor
Adamopoulou-Soulantika, MarianthiSudarsanam, ParthasaarathyDiaz-Guerra, DavidJiang, MengPolitis, ArchontisSeyed Jalaleddin, MousaviradVirtanen, TuomasLundgren, Jan
By organisation
Department of Computer and Electrical Engineering (2023-)
Electrical Engineering, Electronic Engineering, Information Engineering

Search outside of DiVA

GoogleGoogle Scholar
Total: 66 downloads
The number of downloads is the sum of all downloads of full texts. It may include eg previous versions that are now no longer available

doi
isbn
urn-nbn

Altmetric score

doi
isbn
urn-nbn
Total: 90 hits
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf