Details der Publikation - SynArtifact: Classifying and Alleviating Artifacts in Synthetic Images via Vision-Language Model

SynArtifact: Classifying and Alleviating Artifacts in Synthetic Images via Vision-Language Model

In the rapidly evolving area of image synthesis, a serious challenge is the presence of complex artifacts that compromise perceptual realism of synthetic images. To alleviate artifacts and improve quality of synthetic images, we fine-tune Vision-Language Model (VLM) as artifact classifier to automatically identify and classify a wide range of artifacts and provide supervision for further optimizing generative models. Specifically, we develop a comprehensive artifact taxonomy and construct a dataset of synthetic images with artifact annotations for fine-tuning VLM, named SynArtifact-1K. The fine-tuned VLM exhibits superior ability of identifying artifacts and outperforms the baseline by 25.66%. To our knowledge, this is the first time such end-to-end artifact classification task and solution have been proposed. Finally, we leverage the output of VLM as feedback to refine the generative model for alleviating artifacts. Visualization results and user study demonstrate that the quality of images synthesized by the refined diffusion model has been obviously improved..

Medienart:	Preprint

Erscheinungsjahr:	2024
Erschienen:	2024

Enthalten in:	arXiv.org - (2024) vom: 28. Feb. Zur Gesamtaufnahme - year:2024

Sprache:	Englisch

Beteiligte Personen:	Cao, Bin [VerfasserIn] Yuan, Jianhao [VerfasserIn] Liu, Yexin [VerfasserIn] Li, Jian [VerfasserIn] Sun, Shuyang [VerfasserIn] Liu, Jing [VerfasserIn] Zhao, Bo [VerfasserIn]

Links:	Volltext [kostenfrei]

Themen:	000 Computer Science - Computer Vision and Pattern Recognition

Förderinstitution / Projekttitel:

PPN (Katalog-ID):	XCH042793289

Internformat


LEADER	01000naa a22002652 4500
001	XCH042793289
003	DE-627
005	20240306114517.0
007	cr uuu---uuuuu
008	240306s2024 xx \|\|\|\|\|o 00\| \|\|eng c
035			\|a (DE-627)XCH042793289
035			\|a (chemrXiv)2402.18068
040			\|a DE-627 \|b ger \|c DE-627 \|e rakwb
041			\|a eng
100	1		\|a Cao, Bin \|e verfasserin \|4 aut
245	1	0	\|a SynArtifact: Classifying and Alleviating Artifacts in Synthetic Images via Vision-Language Model
264		1	\|c 2024
336			\|a Text \|b txt \|2 rdacontent
337			\|a Computermedien \|b c \|2 rdamedia
338			\|a Online-Ressource \|b cr \|2 rdacarrier
520			\|a In the rapidly evolving area of image synthesis, a serious challenge is the presence of complex artifacts that compromise perceptual realism of synthetic images. To alleviate artifacts and improve quality of synthetic images, we fine-tune Vision-Language Model (VLM) as artifact classifier to automatically identify and classify a wide range of artifacts and provide supervision for further optimizing generative models. Specifically, we develop a comprehensive artifact taxonomy and construct a dataset of synthetic images with artifact annotations for fine-tuning VLM, named SynArtifact-1K. The fine-tuned VLM exhibits superior ability of identifying artifacts and outperforms the baseline by 25.66%. To our knowledge, this is the first time such end-to-end artifact classification task and solution have been proposed. Finally, we leverage the output of VLM as feedback to refine the generative model for alleviating artifacts. Visualization results and user study demonstrate that the quality of images synthesized by the refined diffusion model has been obviously improved.
650		4	\|a Computer Science - Computer Vision and Pattern Recognition \|7 (dpeaa)DE-84
650		4	\|a 000 \|7 (dpeaa)DE-84
700	1		\|a Yuan, Jianhao \|4 aut
700	1		\|a Liu, Yexin \|4 aut
700	1		\|a Li, Jian \|4 aut
700	1		\|a Sun, Shuyang \|4 aut
700	1		\|a Liu, Jing \|4 aut
700	1		\|a Zhao, Bo \|4 aut
773	0	8	\|i Enthalten in \|t arXiv.org \|g (2024) vom: 28. Feb.
773	1	8	\|g year:2024 \|g day:28 \|g month:02
856	4	0	\|u https://arxiv.org/abs/2402.18068 \|z kostenfrei \|3 Volltext
912			\|a GBV_XCH
951			\|a AR
952			\|j 2024 \|b 28 \|c 02

SynArtifact: Classifying and Alleviating Artifacts in Synthetic Images via Vision-Language Model

Zugang & Verfügbarkeit

Zugehörige Publikationen/Bände