Integrating image and depth information for RGB-D salient object detection has become a research hotspot in the field of saliency detection. Balancing the efficiency and performance of salient object detection models under resource constraints is a key challenge. To address this, this paper proposes an asymmetric lightweight network suitable for real-time RGB-D salient object detection tasks. The network reduces the number of network parameters by designing different lightweight feature extraction networks for different input modalities. Additionally, a multi-modal feature enhancement fusion module is designed to effectively fuse multi-modal features while compensating for the information loss caused by the lightweight backbone network. Moreover, this paper utilizes a global context module for dense decoding, aggregating local and global information of multi-scale features without significantly increasing computational complexity. The experimental results on five benchmarks show that the proposed lightweight RGB-D salient object detection network not only outperforms most mainstream models quantitatively and qualitatively, but also significantly outperforms other models in terms of efficiency, only with a parameter count of 5.1 million and a computational load of 0.77 gigaflops. This achievement validates the proposed method's ability to achieve lightweight salient object detection while maintaining high efficiency.
[1] LEE M, LEE S, LEE J, et al. Saliency as pseudo-pixel supervision for weakly and semi-supervised semantic segmentation [J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2023, 45(10): 12341-12357.
[2] WANG W G, SHEN J B, YU Y Z, et al. Stereoscopic thumbnail creation via efficient stereo saliency detection [J]. IEEE Transactions on Visualization and Computer Graphics, 2017, 23(8): 2014-2027.
[3] WU Y H, GAO S H, MEI J, et al. JCS: An explainable COVID-19 diagnosis system by joint classification and segmentation [J]. IEEE Transactions on Image Processing, 2021, 30: 3113-3126.
[4] ZHANG M, FEI S X, LIU J, et al. Asymmetric two-stream architecture for accurate RGB-D saliency detection [M]// Computer Vision – ECCV 2020. Cham: Springer, 2020: 374-390.
[5] HOWARD A G, ZHU M L, CHEN B, et al. MobileNets: Efficient convolutional neural networks for mobile vision applications [DB/OL]. (2017-04-17). https://arxiv.org/abs/1704.04861
[6] ZHANG X Y, ZHOU X Y, LIN M X, et al. ShuffleNet: An extremely efficient convolutional neural network for mobile devices [C]//2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition. Salt Lake City: IEEE, 2018: 6848-6856.
[7] WU Y H, LIU Y, XU J, et al. MobileSal: Extremely efficient RGB-D salient object detection [J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2022, 44(12): 10261-10269.
[8] CHEN S H, FU Y. Progressively guided alternate refinement network for RGB-D salient object detection [M]//Lecture notes in computer science. Cham: Springer International Publishing, 2020: 520-538.
[9] ZHAO X Q, ZHANG L H, PANG Y W, et al. A single stream network for robust and real-time RGB-D salient object detection [M]// Computer Vision – ECCV 2020. Cham: Springer, 2020: 646-662.
[10] PIAO Y R, RONG Z K, ZHANG M, et al. A2dele: Adaptive and attentive depth distiller for efficient RGB-D salient object detection [C]//2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition. Seattle: IEEE, 2020: 9057-9066.
[11] JI W, LI J, ZHANG M, et al. Accurate RGB-D salient object detection via collaborative learning [M]// Computer Vision – ECCV 2020. Cham: Springer, 2020: 52-69.
[12] JIN X, YI K, XU J. MoADNet: Mobile asymmetric dual-stream networks for real-time and lightweight RGB-D salient object detection [J]. IEEE Transactions on Circuits and Systems for Video Technology, 2022, 32(11): 7632-7645.
[13] ZHANG W B, JI G P, WANG Z, et al. Depth quality-inspired feature manipulation for efficient RGB-D salient object detection [C]// 29th ACM International Conference on Multimedia. Online: ACM, 2021: 731-740.
[14] SONG K C, WANG H, ZHAO Y, et al. Lightweight multi-level feature difference fusion network for RGB-D-T salient object detection [J]. Journal of King Saud University - Computer and Information Sciences, 2023, 35(8): 101702.
[15] PENG H W, LI B, XIONG W H, et al. RGBD salient object detection: A benchmark and algorithms [M]// Computer Vision – ECCV 2014. Cham: Springer, 2014: 92-109.
[16] QU L Q, HE S F, ZHANG J W, et al. RGBD salient object detection via deep fusion [J]. IEEE Transactions on Image Processing, 2017, 26(5): 2274-2285.
[17] LIU Z Y, SHI S, DUAN Q T, et al. Salient object detection for RGB-D image by single stream recurrent convolution neural network [J]. Neurocomputing, 2019, 363: 46-57.
[18] WANG N N, GONG X J. Adaptive fusion for RGB-D salient object detection [J]. IEEE Access, 2019, 7: 55277-55284.
[19] SHIGEMATSU R, FENG D, YOU S D, et al. Learning RGB-D salient object detection using background enclosure, depth contrast, and top-down features [C]//2017 IEEE International Conference on Computer Vision Workshops. Venice: IEEE, 2017: 2749-2757.
[20] HAN J W, CHEN H, LIU N, et al. CNNs-based RGB-D saliency detection via cross-view transfer and multiview fusion [J]. IEEE Transactions on Cybernetics, 2018, 48(11): 3171-3183.
[21] FU K R, FAN D P, JI G P, et al. JL-DCF: Joint learning and densely-cooperative fusion framework for RGB-D salient object detection [C]//2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition. Seattle: IEEE, 2020: 3049-3059.
[22] CHEN H, LI Y F. Progressively complementarity-aware fusion network for RGB-D salient object detection [C]//2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition. Salt Lake City: IEEE, 2018: 3051-3060.
[23] PIAO Y R, JI W, LI J J, et al. Depth-induced multi-scale recurrent attention network for saliency detection [C]//2019 IEEE/CVF International Conference on Computer Vision. Seoul: IEEE, 2019: 7253-7262.
[24] FAN D P, ZHAI Y J, BORJI A, et al. BBS-net: RGB-D salient object detection with a bifurcated backbone strategy network [M]// Computer Vision – ECCV 2020. Cham: Springer, 2020: 275-292.
[25] HOWARD A, SANDLER M, CHEN B, et al. Searching for MobileNetV3 [C]//2019 IEEE/CVF International Conference on Computer Vision. Seoul: IEEE, 2019: 1314-1324.
[26] SANDLER M, HOWARD A, ZHU M L, et al. MobileNetV2: Inverted residuals and linear bottlenecks [C]//2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition. Salt Lake City: IEEE, 2018: 4510-4520.
[27] ZHAO J X, CAO Y, FAN D P, et al. Contrast prior and fluid pyramid integration for RGBD salient object detection [C]//2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition. Long Beach: IEEE, 2019: 3922-3931.
[28] ZHANG J, FAN D P, DAI Y C, et al. UC-net: Uncertainty inspired RGB-D saliency detection via conditional variational autoencoders [C]//2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition. Seattle: IEEE, 2020: 8579-8588.
[29] ZHANG M, REN W S, PIAO Y R, et al. Select, supplement and focus for RGB-D saliency detection [C]//2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition. Seattle: IEEE, 2020: 3469-3478.
[30] CHEN L C, PAPANDREOU G, KOKKINOS I, et al. DeepLab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected CRFs [J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2018, 40(4): 834-848.