Details der Publikation - MiM-ISTD: Mamba-in-Mamba for Efficient Infrared Small Target Detection

MiM-ISTD: Mamba-in-Mamba for Efficient Infrared Small Target Detection

Thanks to the development of basic models, infrared small target detection (ISTD) algorithms have made significant progress. Specifically, the structures combining convolutional networks with transformers can well extract both local and global features. At the same time, they also inherit defects from the basic model, e.g., the quadratic computational complexity of transformers, which impacts efficiency. Inspired by a recent basic model with linear complexity for long-distance modeling, called Mamba, we explore the potential of this state space model in ISTD in this paper. However, direct application is unsuitable since local features, which are critical to detecting small targets, cannot be fully exploited. Instead, we tailor a Mamba-in-Mamba (MiM-ISTD) structure for efficient ISTD. For example, we treat the local patches as "visual sentences" and further decompose them into sub-patches as "visual words" to further explore the locality. The interactions among each word in a given visual sentence will be calculated with negligible computational costs. By aggregating the word and sentence features, the representation ability of MiM-ISTD can be significantly bolstered. Experiments on NUAA-SIRST and IRSTD-1k prove the superior accuracy and efficiency of our method. Specifically, MiM-ISTD is $10 \times$ faster than the SOTA and reduces GPU memory usage by 73.4$\%$ per $2048 \times 2048$ image during inference, overcoming the computation$\&$memory constraints on performing Mamba-based understanding on high-resolution infrared images.Source code is available at https://github.com/txchen-USTC/MiM-ISTD..

Medienart:	Preprint

Erscheinungsjahr:	2024
Erschienen:	2024

Enthalten in:	arXiv.org - (2024) vom: 04. März Zur Gesamtaufnahme - year:2024

Sprache:	Englisch

Beteiligte Personen:	Chen, Tianxiang [VerfasserIn] Tan, Zhentao [VerfasserIn] Gong, Tao [VerfasserIn] Chu, Qi [VerfasserIn] Wu, Yue [VerfasserIn] Liu, Bin [VerfasserIn] Ye, Jieping [VerfasserIn] Yu, Nenghai [VerfasserIn]

Links:	Volltext [kostenfrei]

Themen:	000 Computer Science - Computer Vision and Pattern Recognition

Förderinstitution / Projekttitel:

PPN (Katalog-ID):	XCH042764599

Internformat


LEADER	01000naa a22002652 4500
001	XCH042764599
003	DE-627
005	20240306114448.0
007	cr uuu---uuuuu
008	240306s2024 xx \|\|\|\|\|o 00\| \|\|eng c
035			\|a (DE-627)XCH042764599
035			\|a (chemrXiv)2403.02148
040			\|a DE-627 \|b ger \|c DE-627 \|e rakwb
041			\|a eng
100	1		\|a Chen, Tianxiang \|e verfasserin \|4 aut
245	1	0	\|a MiM-ISTD: Mamba-in-Mamba for Efficient Infrared Small Target Detection
264		1	\|c 2024
336			\|a Text \|b txt \|2 rdacontent
337			\|a Computermedien \|b c \|2 rdamedia
338			\|a Online-Ressource \|b cr \|2 rdacarrier
520			\|a Thanks to the development of basic models, infrared small target detection (ISTD) algorithms have made significant progress. Specifically, the structures combining convolutional networks with transformers can well extract both local and global features. At the same time, they also inherit defects from the basic model, e.g., the quadratic computational complexity of transformers, which impacts efficiency. Inspired by a recent basic model with linear complexity for long-distance modeling, called Mamba, we explore the potential of this state space model in ISTD in this paper. However, direct application is unsuitable since local features, which are critical to detecting small targets, cannot be fully exploited. Instead, we tailor a Mamba-in-Mamba (MiM-ISTD) structure for efficient ISTD. For example, we treat the local patches as "visual sentences" and further decompose them into sub-patches as "visual words" to further explore the locality. The interactions among each word in a given visual sentence will be calculated with negligible computational costs. By aggregating the word and sentence features, the representation ability of MiM-ISTD can be significantly bolstered. Experiments on NUAA-SIRST and IRSTD-1k prove the superior accuracy and efficiency of our method. Specifically, MiM-ISTD is $10 \times$ faster than the SOTA and reduces GPU memory usage by 73.4$\%$ per $2048 \times 2048$ image during inference, overcoming the computation$\&$memory constraints on performing Mamba-based understanding on high-resolution infrared images.Source code is available at https://github.com/txchen-USTC/MiM-ISTD.
650		4	\|a Computer Science - Computer Vision and Pattern Recognition \|7 (dpeaa)DE-84
650		4	\|a 000 \|7 (dpeaa)DE-84
700	1		\|a Tan, Zhentao \|4 aut
700	1		\|a Gong, Tao \|4 aut
700	1		\|a Chu, Qi \|4 aut
700	1		\|a Wu, Yue \|4 aut
700	1		\|a Liu, Bin \|4 aut
700	1		\|a Ye, Jieping \|4 aut
700	1		\|a Yu, Nenghai \|4 aut
773	0	8	\|i Enthalten in \|t arXiv.org \|g (2024) vom: 04. März
773	1	8	\|g year:2024 \|g day:04 \|g month:03
856	4	0	\|u https://arxiv.org/abs/2403.02148 \|z kostenfrei \|3 Volltext
912			\|a GBV_XCH
951			\|a AR
952			\|j 2024 \|b 04 \|c 03

MiM-ISTD: Mamba-in-Mamba for Efficient Infrared Small Target Detection

Zugang & Verfügbarkeit

Zugehörige Publikationen/Bände