Image inpainting is a crucial research area in computer vision. Despite significant advancements with deep learning methods, challenges such as information loss and weak adaptability remain. This paper introduces a Transformer-based image inpainting method named Swin2FII, which integrates SwinV2 Transformer and fast Fourier convolution structure to address information loss and bottleneck issues, significantly enhancing inpainting accuracy and expanding its application scope. Swin2FII incorporates a super-resolution model, enhancing feature extraction and information transmission through efficient reconstruction, thereby improving detail recovery and stability. We employ the Charbonnier loss function to address gradient explosion, accurately estimating low-frequency signals and enhancing the precision of detail and texture reconstruction. Furthermore, combining mixed-precision training and data augmentation significantly boosts the model's adaptability and generalization ability. Experimental results show that our Swin2FII method outperforms the existing techniques on multiple public datasets. Notably, it exhibits excellent generalization and performance in a variety of scenarios and mask scales. In addition, Swin2FII also demonstrates strong capabilities in fluid image inpainting and mural image inpainting tasks.
[1] ZHAO L L, SHEN L, HONG R C. Survey on image inpainting research progress [J]. Computer Science, 2021, 48(3): 14-26 (in Chinese).
[2] PATHAK D, KRÄHENBÜHL P, DONAHUE J, et al. Context encoders: Feature learning by inpainting [C]//2016 IEEE Conference on Computer Vision and Pattern Recognition. Las Vegas: IEEE, 2016: 2536-2544.
[3] LUO H Y, ZHENG Y H. Survey of research on image inpainting methods [J]. Journal of Frontiers of Computer Science & Technology, 2022, 16(10): 2193-2218 (in Chinese).
[4] DOSOVITSKIY A, BEYER L, KOLESNIKOV A, et al. An image is worth 16×16 words: Transformers for image recognition at scale 16 words: Transformers for image recognition at scale [DB/OL]. (2020-10-20). https://arxiv.org/abs/2010.11929
[5] LIU Z, LIN Y T, CAO Y, et al. Swin transformer: Hierarchical vision transformer using shifted windows [C]//2021 IEEE/CVF International Conference on Computer Vision. Montreal: IEEE, 2021: 9992-10002.
[6] BERTALMIO M, SAPIRO G, CASELLES V, et al. Image inpainting [M]// Proceedings of the 27th annual conference on Computer graphics and interactive techniques. New York: ACM Press, 2000: 417-424.
[7] CHAN T F, SHEN J H. Nontexture inpainting by curvature-driven diffusions [J]. Journal of Visual Communication and Image Representation, 2001, 12(4): 436-449.
[8] CRIMINISI A, PÉREZ P, TOYAMA K. Region filling and object removal by exemplar-based image inpainting [J]. IEEE Transactions on Image Processing, 2004, 13(9): 1200-1212.
[9] BARNES C, SHECHTMAN E, FINKELSTEIN A, et al. PatchMatch: a randomized correspondence algorithm for structural image editing[J]. ACM Transactions on Graphics, 2009, 28(3): 24.
[10] QIN Z, ZENG Q L, ZONG Y X, et al. Image inpainting based on deep learning: A review [J]. Displays, 2021, 69: 102028.
[11] YU J H, LIN Z, YANG J M, et al. Free-form image inpainting with gated convolution [C]//2019 IEEE/CVF International Conference on Computer Vision. Seoul: IEEE, 2019: 4470-4479.
[12] ZHU M Y, HE D L, LI X, et al. Image inpainting by end-to-end cascaded refinement with mask awareness [J]. IEEE Transactions on Image Processing, 2021, 30: 4855-4866.
[13] VASWANI A, SHAZEER N, PARMAR N, et al. Attention is all you need [C]// 31st Conference on Neural Information Processing Systems. Long Beach: NIPS, 2017: 1-11.
[14] LIU H Y, JIANG B, XIAO Y, et al. Coherent semantic attention for image inpainting [C]//2019 IEEE/CVF International Conference on Computer Vision. Seoul: IEEE, 2019: 4169-4178.
[15] DENG Y, HUI S Q, MENG R Y, et al. Hourglass attention network for Image inpainting [M]// Computer Vision – ECCV 2022. Cham: Springer, 2022: 483-501.
[16] LI W B, LIN Z, ZHOU K, et al. MAT: Mask-aware transformer for large hole image inpainting [C]//2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition. New Orleans: IEEE, 2022: 10748-10758.
[17] DONG Q L, CAO C J, FU Y W. Incremental transformer structure enhanced image inpainting with masking positional encoding [C]//2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition. New Orleans: IEEE, 2022: 11348-11358.
[18] LIANG J Y, CAO J Z, SUN G L, et al. SwinIR: Image restoration using swin transformer [C]//2021 IEEE/CVF International Conference on Computer Vision Workshops. Montreal: IEEE, 2021: 1833-1844.
[19] CHI L, JIANG B, MU Y. Fast Fourier convolution [C]// 34th Conference on Neural Information Processing Systems. Vancouver: NIPS, 2020: 4479-4488.
[20] SHCHEKOTOV I, ANDREEV P, IVANOV O, et al. FFC-SE: Fast Fourier convolution for speech enhancement [DB/OL]. (2022-04-06). https://arxiv.org/abs/2204.03042
[21] BERENGUEL-BAETA B, BERMUDEZ-CAMEO J, GUERRERO J J. FreDSNet: Joint monocular depth and semantic segmentation with fast Fourier convolutions [DB/OL]. (2022-10-04). https://arxiv.org/abs/2210.01595
[22] ZHANG D F, HUANG F Y, LIU S Z, et al. SwinFIR: Revisiting the SwinIR with fast Fourier convolution and improved training for image super-resolution [DB/OL]. (2022-08-24). https://arxiv.org/abs/2208.11247
[23] SUVOROV R, LOGACHEVA E, MASHIKHIN A, et al. Resolution-robust large mask inpainting with Fourier convolutions [C]//2022 IEEE/CVF Winter Conference on Applications of Computer Vision. Waikoloa: IEEE, 2022: 3172-3182.
[24] CONDE M V, CHOI U J, BURCHI M, et al. Swin2SR: SwinV2 Transformer for Compressed Image Super-Resolution and Restoration [M]// Computer Vision – ECCV 2022 Workshops. Cham: Springer, 2023: 669-687.
[25] LIU Z, HU H, LIN Y T, et al. Swin transformer V2: Scaling up capacity and resolution [C]//2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition. New Orleans: IEEE, 2022: 11999-12009.
[26] LAI W S, HUANG J B, AHUJA N, et al. Fast and accurate image super-resolution with deep Laplacian pyramid networks [J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2019, 41(11): 2599-2613.
[27] DOERSCH C, SINGH S, GUPTA A, et al. What makes Paris look like Paris? [J]. Communications of the ACM, 2015, 58(12): 103-110.
[28] KARRAS T, AILA T M, LAINE S, et al. Progressive growing of GANs for improved quality, stability, and variation [DB/OL]. (2017-10-27). https://arxiv.org/abs/1710.10196
[29] ZHOU B L, LAPEDRIZA A, KHOSLA A, et al. Places: A 10 million image database for scene recognition [J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2018, 40(6): 1452-1464.
[30] LIU G L, REDA F A, SHIH K J, et al. Image inpainting for irregular holes using partial convolutions [M]// Computer Vision – ECCV 2018. Cham: Springer, 2018: 89-105.
[31] ZHENG C X, CHAM T J, CAI J F. Pluralistic image completion [C]//2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition. Long Beach: IEEE, 2019: 1438-1447.
[32] NAZERI K, NG E, JOSEPH T, et al. EdgeConnect: Generative image inpainting with adversarial edge learning [DB/OL]. (2019-01-01). https://arxiv.org/abs/1901.00212
[33] LI J Y, WANG N, ZHANG L F, et al. Recurrent feature reasoning for image inpainting [C]//2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition. Seattle: IEEE, 2020: 7760-7768.
[34] GUO X F, YANG H Y, HUANG D. Image inpainting via conditional texture and structure dual generation [C]//2021 IEEE/CVF International Conference on Computer Vision. Montreal: IEEE, 2021: 14114-14123.
[35] LI X G, GUO Q, LIN D, et al. MISF: Multi-level interactive Siamese filtering for high-fidelity image inpainting [C]//2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition. New Orleans: IEEE, 2022: 1859-1868.
[36] KINGMA D P, BA J, HAMMAD M M. Adam: A method for stochastic optimization [DB/OL]. (2014-12-22). https://arxiv.org/abs/1412.6980