Explora I+D+i UPV

Volver atrás Publicación

GENERACIÓN DE RIRS PARA MODELADO DE SALAS MEDIANTE VITGAN

Compartir
Autores UPV

Año

CONGRESO

GENERACIÓN DE RIRS PARA MODELADO DE SALAS MEDIANTE VITGAN

Abstract

Deep Learning (DL) models have been used for some years in sound applications such as sound event detection, echo cancellation, voice enhancement or music genre classification, among others. Recently, neural networks based on attention modules (Transformers) have demonstrated their great ability to capture relationships between distinct parts of a natural language text or an image. These DL models can identify relationships globally, rather than relying solely on nearby local information. In this paper we present an exploratory investigation of the ability of such networks to model the Room Impulse Response (RIR) between any two positions of the transmitter and receiver. For this purpose, RIRs measured in six different rooms with a circular array of sixty loudspeakers and two microphone arrays, a planar array and circular array of 64 and 30 microphones, respectively, have been used. The results show that the parameters of the time-frequency representation of the RIR and the type of loss function used can improve the quality of the RIRs inferred by the Transformer.