TY - GEN
T1 - Co-Painter
T2 - 2025 IEEE/CVF International Conference on Computer Vision, ICCV 2025
AU - Fu, Bowen
AU - Wei, Wei
AU - Tang, Jiaqi
AU - Nie, Jiangtao
AU - Ye, Yanyu
AU - Xu, Xiaogang
AU - Chen, Ying Cong
AU - Zhang, Lei
N1 - Publisher Copyright:
© 2025 IEEE.
PY - 2025
Y1 - 2025
N2 - Controllable diffusion models have been widely applied in image stylization. However, existing methods often treat the style in the reference image as a single, indivisible entity, which makes it difficult to transfer specific stylistic attributes. To address this issue, we propose a fine-grained controllable image stylization framework, Co-PAINTER, to decouple multiple attributes embedded in the reference image and adaptively inject them into the diffusion model. We first build a multi-condition image stylization framework based on the text-to-image generation model. Then, to drive it, we develop a fine-grained decoupling mechanism to implicitly separate the attributes from the image. Finally, we design a gated feature injection mechanism to adaptively regulate the importance of multiple attributes. To support the above procedure, we also build a dataset with fine-grained styles. It comprises nearly 48,000 image-text pairs samples. Extensive experiments demonstrate that the proposed model achieves an optimal balance between text alignment and style similarity to reference images, both in standard and fine-grained settings. Our code: https://github.com/bowen310/Co-Painter
AB - Controllable diffusion models have been widely applied in image stylization. However, existing methods often treat the style in the reference image as a single, indivisible entity, which makes it difficult to transfer specific stylistic attributes. To address this issue, we propose a fine-grained controllable image stylization framework, Co-PAINTER, to decouple multiple attributes embedded in the reference image and adaptively inject them into the diffusion model. We first build a multi-condition image stylization framework based on the text-to-image generation model. Then, to drive it, we develop a fine-grained decoupling mechanism to implicitly separate the attributes from the image. Finally, we design a gated feature injection mechanism to adaptively regulate the importance of multiple attributes. To support the above procedure, we also build a dataset with fine-grained styles. It comprises nearly 48,000 image-text pairs samples. Extensive experiments demonstrate that the proposed model achieves an optimal balance between text alignment and style similarity to reference images, both in standard and fine-grained settings. Our code: https://github.com/bowen310/Co-Painter
KW - image stylization; diffusion model; image synthesis;
UR - https://www.scopus.com/pages/publications/105044207428
U2 - 10.1109/ICCV51701.2025.01563
DO - 10.1109/ICCV51701.2025.01563
M3 - 会议稿件
AN - SCOPUS:105044207428
T3 - Proceedings of the IEEE International Conference on Computer Vision
SP - 16830
EP - 16839
BT - Proceedings - 2025 IEEE/CVF International Conference on Computer Vision, ICCV 2025
PB - Institute of Electrical and Electronics Engineers Inc.
Y2 - 19 October 2025 through 23 October 2025
ER -