Details der Publikation

DiffStyler : Controllable Dual Diffusion for Text-Driven Image Stylization

Despite the impressive results of arbitrary image-guided style transfer methods, text-driven image stylization has recently been proposed for transferring a natural image into a stylized one according to textual descriptions of the target style provided by the user. Unlike the previous image-to-image transfer approaches, text-guided stylization progress provides users with a more precise and intuitive way to express the desired style. However, the huge discrepancy between cross-modal inputs/outputs makes it challenging to conduct text-driven image stylization in a typical feed-forward CNN pipeline. In this article, we present DiffStyler, a dual diffusion processing architecture to control the balance between the content and style of the diffused results. The cross-modal style information can be easily integrated as guidance during the diffusion process step-by-step. Furthermore, we propose a content image-based learnable noise on which the reverse denoising process is based, enabling the stylization results to better preserve the structure information of the content image. We validate the proposed DiffStyler beyond the baseline methods through extensive qualitative and quantitative experiments. The code is available at https://github.com/haha-lisa/Diffstyler.

Medienart:	E-Artikel

Erscheinungsjahr:	2024
Erschienen:	2024

Enthalten in:	Zur Gesamtaufnahme - volume:PP
Enthalten in:	IEEE transactions on neural networks and learning systems - PP(2024) vom: 10. Jan.

Sprache:	Englisch

Beteiligte Personen:	Huang, Nisha [VerfasserIn] Zhang, Yuxin [VerfasserIn] Tang, Fan [VerfasserIn] Ma, Chongyang [VerfasserIn] Huang, Haibin [VerfasserIn] Dong, Weiming [VerfasserIn] Xu, Changsheng [VerfasserIn]

Links:	Volltext

Themen:	Journal Article

Anmerkungen:	Date Revised 10.01.2024 published: Print-Electronic Citation Status Publisher

doi:	10.1109/TNNLS.2023.3342645

funding:
Förderinstitution / Projekttitel:

PPN (Katalog-ID):	NLM366888560

Internformat


LEADER	01000naa a22002652 4500
001	NLM366888560
003	DE-627
005	20240114233954.0
007	cr uuu---uuuuu
008	240114s2024 xx \|\|\|\|\|o 00\| \|\|eng c
024	7		\|a 10.1109/TNNLS.2023.3342645 \|2 doi
028	5	2	\|a pubmed24n1256.xml
035			\|a (DE-627)NLM366888560
035			\|a (NLM)38198263
040			\|a DE-627 \|b ger \|c DE-627 \|e rakwb
041			\|a eng
100	1		\|a Huang, Nisha \|e verfasserin \|4 aut
245	1	0	\|a DiffStyler \|b Controllable Dual Diffusion for Text-Driven Image Stylization
264		1	\|c 2024
336			\|a Text \|b txt \|2 rdacontent
337			\|a ƒaComputermedien \|b c \|2 rdamedia
338			\|a ƒa Online-Ressource \|b cr \|2 rdacarrier
500			\|a Date Revised 10.01.2024
500			\|a published: Print-Electronic
500			\|a Citation Status Publisher
520			\|a Despite the impressive results of arbitrary image-guided style transfer methods, text-driven image stylization has recently been proposed for transferring a natural image into a stylized one according to textual descriptions of the target style provided by the user. Unlike the previous image-to-image transfer approaches, text-guided stylization progress provides users with a more precise and intuitive way to express the desired style. However, the huge discrepancy between cross-modal inputs/outputs makes it challenging to conduct text-driven image stylization in a typical feed-forward CNN pipeline. In this article, we present DiffStyler, a dual diffusion processing architecture to control the balance between the content and style of the diffused results. The cross-modal style information can be easily integrated as guidance during the diffusion process step-by-step. Furthermore, we propose a content image-based learnable noise on which the reverse denoising process is based, enabling the stylization results to better preserve the structure information of the content image. We validate the proposed DiffStyler beyond the baseline methods through extensive qualitative and quantitative experiments. The code is available at https://github.com/haha-lisa/Diffstyler
650		4	\|a Journal Article
700	1		\|a Zhang, Yuxin \|e verfasserin \|4 aut
700	1		\|a Tang, Fan \|e verfasserin \|4 aut
700	1		\|a Ma, Chongyang \|e verfasserin \|4 aut
700	1		\|a Huang, Haibin \|e verfasserin \|4 aut
700	1		\|a Dong, Weiming \|e verfasserin \|4 aut
700	1		\|a Xu, Changsheng \|e verfasserin \|4 aut
773	0	8	\|i Enthalten in \|t IEEE transactions on neural networks and learning systems \|d 2012 \|g PP(2024) vom: 10. Jan. \|w (DE-627)NLM23236897X \|x 2162-2388 \|7 nnns
773	1	8	\|g volume:PP \|g year:2024 \|g day:10 \|g month:01
856	4	0	\|u http://dx.doi.org/10.1109/TNNLS.2023.3342645 \|3 Volltext
912			\|a GBV_USEFLAG_A
912			\|a GBV_NLM
951			\|a AR
952			\|d PP \|j 2024 \|b 10 \|c 01

DiffStyler : Controllable Dual Diffusion for Text-Driven Image Stylization

Zugang & Verfügbarkeit

Zugehörige Publikationen/Bände