Designing a JPEG-compatible receiver-side residual refinement pipeline for unmanned aerial vehicle image transmission and object detection
Unmanned aerial vehicle (UAV) image transmission commonly relies on JPEG because of its interoperability and low deployment cost, yet compression can weaken structures used by downstream object detectors. This study presents JPEG-compatible receiver-side residual refinement (JCRR), a post-decoding pipeline that preserves the transmitter, JPEG syntax, and transmitted payload. JCRR comprises bounded residual reconstruction (BRR) and structure-prior residual refinement (SPRR), which adds a normalized Sobel edge prior. Both variants refine JPEG-decoded images before a fixed YOLO26s detector. An additional robustness check was conducted with a separately trained YOLO26n detector without retraining the receiver-side refinement models. Experiments used 548 VisDrone validation images at JPEG Q15, Q20, and Q25. BRR and SPRR were each trained with three seeds, and Artifact Reduction Convolutional Neural Network (ARCNN) was trained as an external post-decoding baseline under the same Q20-pair protocol. At the same transmitted JPEG payload, BRR increased mAP50 by 1.67–2.14 percentage points and SPRR by 1.67–2.22 percentage points. Under strict size bins defined after 640 × 640 letterboxing, SPRR improved small-target matched-target recall by 0.85–1.18 percentage points over JPEG; seed-specific paired bootstrap intervals were positive at every operating point. On the 300-image Q20 reconstruction holdout, SPRR achieved 37.75 ± 0.02 dB PSNR and 0.9779 ± 0.0002 SSIM, compared with 37.18 dB and 0.9763 for JPEG. Under the FP32 receiver-side forward-only timing protocol, BRR and SPRR processed images at approximately 62.0 and 61.6 frames per second, respectively, excluding JPEG decoding, detector inference, communication, and file I/O. WebP remained competitive at nearby, rather than bitrate-matched, operating points. The results support JCRR as a compatibility-oriented receiver-side enhancement for UAV systems that must retain JPEG transmission, while not implying bitrate reduction or end-to-end real-time flight deployment.
- Jia C, Ye F, Sun H, Ma S, Gao W. Learning to compress unmanned aerial vehicle (UAV) captured video: benchmark and analysis. In: Proceedings of the 2023 Data Compression Conference (DCC). Snowbird, UT, USA; 2023. doi: 10.1109/DCC55655.2023.00060
- Varga LA, Koch S, Zell A. Comprehensive analysis of the object detection pipeline on UAVs. Remote Sens. 2022;14(21):5508. doi: 10.3390/rs14215508
- Zhu P, Wen L, Du D, Bian X, Hu Q, Ling H. VisDrone-DET2018: the Vision Meets Drones Object Detection in Image Challenge. In: Leal-Taixe L, Roth S, eds. Computer Vision - ECCV 2018 Workshops. Lecture Notes in Computer Science. Vol 11131. Cham: Springer; 2019:205-220. doi: 10.1007/978-3-030-11021-5_27
- Ding J, Xue N, Xia G-S, Bai X, Yang W, Yang MY, et al. Object detection in aerial images: a large-scale benchmark and challenges. IEEE Trans Pattern Anal Mach Intell. 2022;44(11):7778-7796. doi: 10.1109/TPAMI.2021.3117983
- Varga LA, Kiefer B, Messmer M, Zell A. SeaDronesSee: a maritime benchmark for detecting humans in open water. In: Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision. Waikoloa, HI, USA; 2022:2260-2270. doi: 10.1109/WACV51458.2022.00374
- Wu J, Wu C, Lin Y, et al. Semantic segmentation-based semantic communication system for image transmission. Digit Commun Netw. 2024;10:519-527. doi: 10.1016/j.dcan.2023.02.006
- Wallace GK. The JPEG still picture compression standard. Commun ACM. 1991;34(4):30-44. doi: 10.1145/103085.103089
- Jocher G, Qiu J, Liu M, Lyu S, Akyon FC, Kalfaoglu ME. Ultralytics YOLO26: Unified Real-Time End-to-End Vision Models. arXiv. Preprint posted online 2026. doi: 10.48550/arXiv.2606.03748
- Buckler M, Jayasuriya S, Sampson A. Reconfiguring the imaging pipeline for computer vision. In: Proceedings of the IEEE International Conference on Computer Vision. Venice, Italy; 2017:975-984. doi: 10.1109/ICCV.2017.111
- Blasinski H, Farrell JE, Lian T, Liu Z, Wandell BA. Optimizing image acquisition systems for autonomous driving. Electron Imaging. 2018;2018(15):161-1-161-7. doi: 10.2352/ISSN.2470-1173.2018.05.PMII-161
- Liu Z, Lian T, Farrell JE, Wandell BA. Soft prototyping camera designs for car detection based on a convolutional neural network. In: Proceedings of the IEEE/CVF International Conference on Computer Vision. Seoul, Korea; 2019:2383-2392. doi: 10.1109/ICCVW.2019.00292
- Lu G, Ge X, Zhong T, Hu Q, Geng J. Preprocessing enhanced image compression for machine vision. IEEE Trans Circuits Syst Video Technol. 2024;34(12):13556-13568. doi: 10.1109/tcsvt.2024.3441049
- Guleryuz OG, Chou PA, Hoppe H, et al. Sandwiched image compression: wrapping neural networks around a standard codec. In: 2021 IEEE International Conference on Image Processing. Anchorage, AK, USA; 2021:3757-3761.doi: 10.1109/ICIP42928.2021.9506256
- Son H, Kim T, Lee H, Lee S. Enhanced standard compatible image compression framework based on auxiliary codec networks. IEEE Trans Image Process. 2021;31:664-677. doi: 10.1109/TIP.2021.3134473
- Talebi H, Kelly D, Luo X, et al. Better compression with deep pre-editing. IEEE Trans Image Process. 2021;30:6673-6685.doi: 10.1109/TIP.2021.3096085
- Chadha A, Andreopoulos Y. Deep perceptual preprocessing for video coding. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Nashville, TN, USA; 2021:14852-14861. doi: 10.1109/CVPR46437.2021.01461
- Xiang G, Jia H, Liu J, Cai B, Li Y, Xie X. Adaptive perceptual preprocessing for video coding. In: 2016 IEEE International Symposium on Circuits and Systems. Montreal, QC, Canada; 2016:2535-2538. doi: 10.1109/ISCAS.2016.7539109
- Vidal E, Sturmel N, Guillemot C, Corlay P, Coudoux F-X. New adaptive filters as perceptual preprocessing for rate-quality performance optimization of video coding. Signal Process Image Commun. 2017;52:124-137. doi: 10.1016/j.image.2016.12.003
- Campos J, Meierhans S, Djelouah A, Schroers C. Content Adaptive Optimization for Neural Image Compression. arXiv. Preprint posted online 2019. doi: 10.48550/arXiv.1906.01223
- Tsubota K, Akutsu H, Aizawa K. Universal deep image compression via content-adaptive optimization with adapters. In: Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision. Waikoloa, HI, USA; 2023:2529-2538. doi: 10.1109/WACV56688.2023.00256
- Shen S, Yue H, Yang J. Dec-Adapter: exploring efficient decoder-side adapter for bridging screen content and natural image compression. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Paris, France; 2023:12887-12896. doi: 10.1109/ICCV51070.2023.01184
- Ballé J, Laparra V, Simoncelli EP. End-to-end Optimized Image Compression. arXiv. Preprint posted online 2016. doi: 10.48550/arXiv.1611.01704
- Ballé J, Minnen D, Singh S, Hwang SJ, Johnston N. Variational image compression with a scale hyperprior. arXiv. Preprint posted online 2018. doi: 10.48550/arXiv.1802.01436
- Minnen D, Ballé J, Toderici G. Joint Autoregressive and Hierarchical Priors for Learned Image Compression. arXiv. Preprint posted online 2018. doi: 10.48550/arXiv.1809.02736
- Cheng Z, Sun H, Takeuchi M, Katto J. Learned image compression with discretized Gaussian mixture likelihoods and attention modules. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Seattle, WA, USA; 2020:7939-7948. doi: 10.1109/CVPR42600.2020.00796
- Lee J, Cho S, Beack SK. Context-adaptive Entropy Model for End-to-end Optimized Image Compression. arXiv. Preprint posted online 2018. doi: 10.48550/arXiv.1809.10452
- Zou R, Song C, Zhang Z. The devil is in the details: window-based attention for image compression. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. New Orleans, LA, USA; 2022:17492-17501. doi: 10.1109/CVPR52688.2022.01697
- Bjontegaard G. Calculation of average PSNR differences between RD curves. VCEG-M33. 2001. https://api.semanticscholar.org/CorpusID:61598325
- Li H, Li S, Ding S, et al. Image compression for machine and human vision with spatial-frequency adaptation. In: Computer Vision ¨C ECCV 2024 (Lecture Notes in Computer Science). Switzerland: Springer Nature; 2024:382-399. doi: 10.1007/978-3-031-72983-6_22
- Hu Y, Yang S, Yang W, Duan L-Y, Liu J. Towards coding for human and machine vision: a scalable image coding approach. In: 2020 IEEE International Conference on Multimedia and Expo. London, UK; 2020:1-6. doi: 10.1109/ICME46284.2020.9102750
- Choi H, Bajić IV. Scalable image coding for humans and machines. IEEE Trans Image Process. 2022;31:2739-2754. doi: 10.1109/TIP.2022.3160602
- Liu L, Hu Z, Chen Z, Xu D. ICMH-Net: neural image compression towards both machine vision and human vision. In: Proceedings of the 31st ACM International Conference on Multimedia. New York, NY, USA; 2023:8047-8056. doi: 10.1145/3581783.3612041
- Bai Y, Yang X, Liu X, et al. Towards end-to-end image compression and analysis with transformers. In: Proceedings of the AAAI Conference on Artificial Intelligence. Singapore; 2022;36(1):104-112. doi: 10.1609/aaai.v36i1.19884
- Chen Y-H, Weng Y-C, Kao C-H, Chien C, Chiu W-C, Peng W-H. TransTIC: transferring transformer-based image compression from human perception to machine perception. In: Proceedings of the IEEE/CVF International Conference on Computer Vision. Paris, France; 2023:23297-23307. doi: 10.1109/ICCV51070.2023.02129
- Fischer K, Brand F, Kaup A. Boosting neural image compression for machines using latent space masking. IEEE Trans Circuits Syst Video Technol. 2022. doi: 10.1109/TCSVT.2022.3195322
- Feng R, Jin X, Guo Z, et al. Image coding for machines with omnipotent feature learning. In: Computer Vision – ECCV. Cham: Springer; 2022:510-528. doi: 10.1007/978-3-031-19836-6_29
- Feng R, Gao Y, Jin X, Feng R, Chen Z. Semantically structured image compression via irregular group-based decoupling. In: Proceedings of the IEEE/CVF International Conference on Computer Vision. Paris, France; 2023:17237-17247. doi: 10.1109/ICCV51070.2023.01581
- Feng R, Liu J, Jin X, Pan X, Sun H, Chen Z. Prompt-ICM: a unified framework towards image coding for machines with task-driven prompts. arXiv. Preprint posted online 2023. doi: 10.48550/arXiv.2305.02578
- Liu J, Jin X, Feng R, Chen Z, Zeng W. Composable image coding for machine via task-oriented internal adaptor and external prior. In: 2023 IEEE International Conference on Visual Communications and Image Processing. Jeju, Republic of Korea; 2023:1-5. doi: 10.1109/VCIP59821.2023.10402659
- Liu J, Sun H, Katto J. Improving multiple machine vision tasks in the compressed domain. In: 2022 26th International Conference on Pattern Recognition. Montreal, QC, Canada; 2022:331-337. doi: 10.1109/ICPR56361.2022.9956532
- Li X, Shi J, Chen Z. Task-driven semantic coding via reinforcement learning. IEEE Trans Image Process. 2021;30:6307-6320. doi: 10.1109/TIP.2021.3091909
- Yan N, Gao C, Liu D, Li H, Li L, Wu F. SSSIC: semantics-to-signal scalable image coding with learned structural representations. IEEE Trans Image Process. 2021;30:8939-8954. doi: 10.1109/TIP.2021.3121131
- Liu K, Liu D, Li L, Yan N, Li H. Semantics-to-signal scalable image compression with learned revertible representations. Int J Comput Vis. 2021;129(9):2605-2621. doi: 10.1007/s11263-021-01491-7
- Duan L, Liu J, Yang W, Huang T, Gao W. Video coding for machines: a paradigm of collaborative compression and intelligent analytics. IEEE Trans Image Process. 2020;29:8680-8695. doi: 10.1109/TIP.2020.3016485
- Yang W, Huang H, Hu Y, Duan L-Y, Liu J. Video coding for machines: compact visual representation compression for intelligent collaborative analytics. IEEE Trans Pattern Anal Mach Intell. 2024;46(7):5174-5191. doi: 10.1109/TPAMI.2024.3367293
- Wang S, Wang S, Yang W, et al. Towards analysis-friendly face representation with scalable feature and texture compression. IEEE Trans Multimed. 2021;23:2957-2971. doi: 10.1109/TMM.2021.3094300
- Codevilla F, Simard JG, Goroshin R, Pal C. Learned image compression for machine perception. arXiv. Preprint posted online 2021. doi: 10.48550/arXiv.2111.02249
- Charbonnier P, Blanc-Feraud L, Aubert G, Barlaud M. Two deterministic half-quadratic regularization algorithms for computed imaging. In: Proceedings of the IEEE International Conference on Image Processing. Austin, TX, USA; 1994:168-172. doi: 10.1109/ICIP.1994.413553
- Loshchilov I, Hutter F. Decoupled Weight Decay Regularization. arXiv. Preprint posted online 2017. doi: 10.48550/arXiv.1711.05101
- Dong C, Deng Y, Loy CC, Tang X. Compression artifacts reduction by a deep convolutional network. In: Proceedings of the IEEE International Conference on Computer Vision. Santiago, Chile; 2015:576-584. doi: 10.1109/ICCV.2015.73
