Cross-attention Inspired Selective State Space Models for Target Sound Extraction

The Transformer model, particularly its cross-attention module, is widely used for feature fusion in target sound extraction which extracts the signal of interest based on given clues. Despite its effectiveness, this approach suffers from low computational efficiency. Recent advancements in state sp...

Ausführliche Beschreibung

Gespeichert in:

Bibliographische Detailangaben
Veröffentlicht in:	arXiv.org 2024-12
Hauptverfasser:	Wu, Donghang, Wang, Yiwen, Wu, Xihong, Qu, Tianshu
Format:	Artikel
Sprache:	eng
Schlagworte:	Effectiveness Mixtures State space models Task complexity Transformers
Online-Zugang:	Volltext
Tags:	Tag hinzufügen Keine Tags, Fügen Sie den ersten Tag hinzu!

container_end_page
container_issue
container_start_page
container_title	arXiv.org
container_volume
creator	Wu, Donghang Wang, Yiwen Wu, Xihong Qu, Tianshu
description	The Transformer model, particularly its cross-attention module, is widely used for feature fusion in target sound extraction which extracts the signal of interest based on given clues. Despite its effectiveness, this approach suffers from low computational efficiency. Recent advancements in state space models, notably the latest work Mamba, have shown comparable performance to Transformer-based methods while significantly reducing computational complexity in various tasks. However, Mamba's applicability in target sound extraction is limited due to its inability to capture dependencies between different sequences as the cross-attention does. In this paper, we propose CrossMamba for target sound extraction, which leverages the hidden attention mechanism of Mamba to compute dependencies between the given clues and the audio mixture. The calculation of Mamba can be divided to the query, key and value. We utilize the clue to generate the query and the audio mixture to derive the key and value, adhering to the principle of the cross-attention mechanism in Transformers. Experimental results from two representative target sound extraction methods validate the efficacy of the proposed CrossMamba.
format	Article
fullrecord	<record><control><sourceid>proquest</sourceid><recordid>TN_cdi_proquest_journals_3103018841</recordid><sourceformat>XML</sourceformat><sourcesystem>PC</sourcesystem><sourcerecordid>3103018841</sourcerecordid><originalsourceid>FETCH-proquest_journals_31030188413</originalsourceid><addsrcrecordid>eNqNjM0KgkAURocgSMp3uNBamB8t92LUokXoPga9hiIzNvcaPX4GPUCb7yzO4VuJSBujkjzVeiNiokFKqQ9HnWUmErcieKLEMqPj3ju4OJr6gC1UOGLD_QuhYsvLTrZBuPoWR4LOB6hteCBD5WfXQvnmYJvvw06sOzsSxj9uxf5U1sU5mYJ_zkh8H_wc3KLuRkkjVZ6nyvxXfQCyGT9I</addsrcrecordid><sourcetype>Aggregation Database</sourcetype><iscdi>true</iscdi><recordtype>article</recordtype><pqid>3103018841</pqid></control><display><type>article</type><title>Cross-attention Inspired Selective State Space Models for Target Sound Extraction</title><source>Free E- Journals</source><creator>Wu, Donghang ; Wang, Yiwen ; Wu, Xihong ; Qu, Tianshu</creator><creatorcontrib>Wu, Donghang ; Wang, Yiwen ; Wu, Xihong ; Qu, Tianshu</creatorcontrib><description>The Transformer model, particularly its cross-attention module, is widely used for feature fusion in target sound extraction which extracts the signal of interest based on given clues. Despite its effectiveness, this approach suffers from low computational efficiency. Recent advancements in state space models, notably the latest work Mamba, have shown comparable performance to Transformer-based methods while significantly reducing computational complexity in various tasks. However, Mamba's applicability in target sound extraction is limited due to its inability to capture dependencies between different sequences as the cross-attention does. In this paper, we propose CrossMamba for target sound extraction, which leverages the hidden attention mechanism of Mamba to compute dependencies between the given clues and the audio mixture. The calculation of Mamba can be divided to the query, key and value. We utilize the clue to generate the query and the audio mixture to derive the key and value, adhering to the principle of the cross-attention mechanism in Transformers. Experimental results from two representative target sound extraction methods validate the efficacy of the proposed CrossMamba.</description><identifier>EISSN: 2331-8422</identifier><language>eng</language><publisher>Ithaca: Cornell University Library, arXiv.org</publisher><subject>Effectiveness ; Mixtures ; State space models ; Task complexity ; Transformers</subject><ispartof>arXiv.org, 2024-12</ispartof><rights>2024. This work is published under http://arxiv.org/licenses/nonexclusive-distrib/1.0/ (the “License”). Notwithstanding the ProQuest Terms and Conditions, you may use this content in accordance with the terms of the License.</rights><oa>free_for_read</oa><woscitedreferencessubscribed>false</woscitedreferencessubscribed></display><links><openurl>$$Topenurl_article</openurl><openurlfulltext>$$Topenurlfull_article</openurlfulltext><thumbnail>$$Tsyndetics_thumb_exl</thumbnail><link.rule.ids>780,784</link.rule.ids></links><search><creatorcontrib>Wu, Donghang</creatorcontrib><creatorcontrib>Wang, Yiwen</creatorcontrib><creatorcontrib>Wu, Xihong</creatorcontrib><creatorcontrib>Qu, Tianshu</creatorcontrib><title>Cross-attention Inspired Selective State Space Models for Target Sound Extraction</title><title>arXiv.org</title><description>The Transformer model, particularly its cross-attention module, is widely used for feature fusion in target sound extraction which extracts the signal of interest based on given clues. Despite its effectiveness, this approach suffers from low computational efficiency. Recent advancements in state space models, notably the latest work Mamba, have shown comparable performance to Transformer-based methods while significantly reducing computational complexity in various tasks. However, Mamba's applicability in target sound extraction is limited due to its inability to capture dependencies between different sequences as the cross-attention does. In this paper, we propose CrossMamba for target sound extraction, which leverages the hidden attention mechanism of Mamba to compute dependencies between the given clues and the audio mixture. The calculation of Mamba can be divided to the query, key and value. We utilize the clue to generate the query and the audio mixture to derive the key and value, adhering to the principle of the cross-attention mechanism in Transformers. Experimental results from two representative target sound extraction methods validate the efficacy of the proposed CrossMamba.</description><subject>Effectiveness</subject><subject>Mixtures</subject><subject>State space models</subject><subject>Task complexity</subject><subject>Transformers</subject><issn>2331-8422</issn><fulltext>true</fulltext><rsrctype>article</rsrctype><creationdate>2024</creationdate><recordtype>article</recordtype><sourceid>ABUWG</sourceid><sourceid>AFKRA</sourceid><sourceid>AZQEC</sourceid><sourceid>BENPR</sourceid><sourceid>CCPQU</sourceid><sourceid>DWQXO</sourceid><recordid>eNqNjM0KgkAURocgSMp3uNBamB8t92LUokXoPga9hiIzNvcaPX4GPUCb7yzO4VuJSBujkjzVeiNiokFKqQ9HnWUmErcieKLEMqPj3ju4OJr6gC1UOGLD_QuhYsvLTrZBuPoWR4LOB6hteCBD5WfXQvnmYJvvw06sOzsSxj9uxf5U1sU5mYJ_zkh8H_wc3KLuRkkjVZ6nyvxXfQCyGT9I</recordid><startdate>20241221</startdate><enddate>20241221</enddate><creator>Wu, Donghang</creator><creator>Wang, Yiwen</creator><creator>Wu, Xihong</creator><creator>Qu, Tianshu</creator><general>Cornell University Library, arXiv.org</general><scope>8FE</scope><scope>8FG</scope><scope>ABJCF</scope><scope>ABUWG</scope><scope>AFKRA</scope><scope>AZQEC</scope><scope>BENPR</scope><scope>BGLVJ</scope><scope>CCPQU</scope><scope>DWQXO</scope><scope>HCIFZ</scope><scope>L6V</scope><scope>M7S</scope><scope>PIMPY</scope><scope>PQEST</scope><scope>PQQKQ</scope><scope>PQUKI</scope><scope>PRINS</scope><scope>PTHSS</scope></search><sort><creationdate>20241221</creationdate><title>Cross-attention Inspired Selective State Space Models for Target Sound Extraction</title><author>Wu, Donghang ; Wang, Yiwen ; Wu, Xihong ; Qu, Tianshu</author></sort><facets><frbrtype>5</frbrtype><frbrgroupid>cdi_FETCH-proquest_journals_31030188413</frbrgroupid><rsrctype>articles</rsrctype><prefilter>articles</prefilter><language>eng</language><creationdate>2024</creationdate><topic>Effectiveness</topic><topic>Mixtures</topic><topic>State space models</topic><topic>Task complexity</topic><topic>Transformers</topic><toplevel>online_resources</toplevel><creatorcontrib>Wu, Donghang</creatorcontrib><creatorcontrib>Wang, Yiwen</creatorcontrib><creatorcontrib>Wu, Xihong</creatorcontrib><creatorcontrib>Qu, Tianshu</creatorcontrib><collection>ProQuest SciTech Collection</collection><collection>ProQuest Technology Collection</collection><collection>Materials Science & Engineering Collection</collection><collection>ProQuest Central (Alumni Edition)</collection><collection>ProQuest Central UK/Ireland</collection><collection>ProQuest Central Essentials</collection><collection>ProQuest Central</collection><collection>Technology Collection</collection><collection>ProQuest One Community College</collection><collection>ProQuest Central Korea</collection><collection>SciTech Premium Collection</collection><collection>ProQuest Engineering Collection</collection><collection>Engineering Database</collection><collection>Access via ProQuest (Open Access)</collection><collection>ProQuest One Academic Eastern Edition (DO NOT USE)</collection><collection>ProQuest One Academic</collection><collection>ProQuest One Academic UKI Edition</collection><collection>ProQuest Central China</collection><collection>Engineering Collection</collection></facets><delivery><delcategory>Remote Search Resource</delcategory><fulltext>fulltext</fulltext></delivery><addata><au>Wu, Donghang</au><au>Wang, Yiwen</au><au>Wu, Xihong</au><au>Qu, Tianshu</au><format>book</format><genre>document</genre><ristype>GEN</ristype><atitle>Cross-attention Inspired Selective State Space Models for Target Sound Extraction</atitle><jtitle>arXiv.org</jtitle><date>2024-12-21</date><risdate>2024</risdate><eissn>2331-8422</eissn><abstract>The Transformer model, particularly its cross-attention module, is widely used for feature fusion in target sound extraction which extracts the signal of interest based on given clues. Despite its effectiveness, this approach suffers from low computational efficiency. Recent advancements in state space models, notably the latest work Mamba, have shown comparable performance to Transformer-based methods while significantly reducing computational complexity in various tasks. However, Mamba's applicability in target sound extraction is limited due to its inability to capture dependencies between different sequences as the cross-attention does. In this paper, we propose CrossMamba for target sound extraction, which leverages the hidden attention mechanism of Mamba to compute dependencies between the given clues and the audio mixture. The calculation of Mamba can be divided to the query, key and value. We utilize the clue to generate the query and the audio mixture to derive the key and value, adhering to the principle of the cross-attention mechanism in Transformers. Experimental results from two representative target sound extraction methods validate the efficacy of the proposed CrossMamba.</abstract><cop>Ithaca</cop><pub>Cornell University Library, arXiv.org</pub><oa>free_for_read</oa></addata></record>
fulltext	fulltext
identifier	EISSN: 2331-8422
ispartof	arXiv.org, 2024-12
issn	2331-8422
language	eng
recordid	cdi_proquest_journals_3103018841
source	Free E- Journals
subjects	Effectiveness Mixtures State space models Task complexity Transformers
title	Cross-attention Inspired Selective State Space Models for Target Sound Extraction
url	https://sfx.bib-bvb.de/sfx_tum?ctx_ver=Z39.88-2004&ctx_enc=info:ofi/enc:UTF-8&ctx_tim=2025-01-03T09%3A06%3A35IST&url_ver=Z39.88-2004&url_ctx_fmt=infofi/fmt:kev:mtx:ctx&rfr_id=info:sid/primo.exlibrisgroup.com:primo3-Article-proquest&rft_val_fmt=info:ofi/fmt:kev:mtx:book&rft.genre=document&rft.atitle=Cross-attention%20Inspired%20Selective%20State%20Space%20Models%20for%20Target%20Sound%20Extraction&rft.jtitle=arXiv.org&rft.au=Wu,%20Donghang&rft.date=2024-12-21&rft.eissn=2331-8422&rft_id=info:doi/&rft_dat=%3Cproquest%3E3103018841%3C/proquest%3E%3Curl%3E%3C/url%3E&disable_directlink=true&sfx.directlink=off&sfx.report_link=0&rft_id=info:oai/&rft_pqid=3103018841&rft_id=info:pmid/&rfr_iscdi=true