Video frame interpolation based on invertible neural network for Internet of things

Lingfan Wu , Keyuan Ye , Huan Gao , Huchen Jiang , Hong Zhang

›› 2026, Vol. 12 ›› Issue (3) : 540 -549.

PDF (2280KB)
›› 2026, Vol. 12 ›› Issue (3) :540 -549. DOI: 10.1016/j.dcan.2025.01.005
Research article
research-article
Video frame interpolation based on invertible neural network for Internet of things
Author information +
History +
PDF (2280KB)

Abstract

Video frame interpolation focuses on directly synthesizing intermediate frames by utilizing inter-frame changes, relying heavily on large, high-frame-rate video datasets for supervised training, which imposes significant demands on bandwidth and computational resources in the Internet of Everything (IoE) environments. Since video frame rate down and up conversion are inverse processes, an Invertible Neural Network (INN) provides an efficient solution by ensuring lossless and symmetrical information transfer in forward and backward processes. This paper introduces a self-supervised video frame rate conversion method based on an INN to reconstruct missing intermediate frames. By leveraging an invertible coupling structure, the model encodes the spatio-temporal features of sparse input frames into a Gaussian distribution, which effectively simulates the frame rate downsampling process. Through the network’s invertibility and lossless processing capabilities, the intermediate frames are then reconstructed through reverse inference. This approach captures missing information from a Gaussian prior, ensuring stability and realism in the generated frames. Extensive experiments on public video datasets show that the proposed method surpasses existing state-of-the-art algorithms in accuracy and efficiency, offering superior visual quality, faster processing speeds, and reduced model parameters, especially for high-frame-rate recovery.

Keywords

Video frame interpolation / Deep learning / Invertible neural network

Cite this article

Download citation ▾
Lingfan Wu, Keyuan Ye, Huan Gao, Huchen Jiang, Hong Zhang. Video frame interpolation based on invertible neural network for Internet of things. , 2026, 12 (3) : 540-549 DOI:10.1016/j.dcan.2025.01.005

登录浏览全文

4963

注册一个新账户 忘记密码

CRediT authorship contribution statement

Lingfan Wu: Funding acquisition, Formal analysis, Data curation, Conceptualization. Keyuan Ye: Writing -- original draft, Visualization, Validation, Software, Resources, Project administration, Methodology, Investigation. Huan Gao: Writing -- review & editing, Writing -- original draft, Visualization, Validation. Huchen Jiang: Resources. Hong Zhang: Writing -- review & editing, Funding acquisition.

Declaration of competing interest

The authors declare the following financial interests/personal relationships which may be considered as potential competing interests: Keyuan Ye reports financial support was provided by National Natural Science Foundation of China (No. 62476207). Keyuan Ye reports financial support was provided by the Chongqing Natural Science Foundation Innovation and Development Joint Fund Project under Grant CSTB2023NSCQ-LZX0085. Keyuan Ye reports financial support was provided by the Key Industrial Innovation Chain Project in Industrial Domain of Shaanxi Province (Grant No. 2020ZDLGY05-01). Hong Zhang reports financial support was provided by National Natural Science Foundation of China (No. 62433003). Hong Zhang reports financial support was provided by National Natural Science Foundation of China (No. 62476017). If there are other authors, they declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper.

Acknowledgements

This research was supported partially by the National Natural Science Foundation of China (No. 62476207, 62433003, 62476017), the Chongqing Natural Science Foundation Innovation and Development Joint Fund Project under Grant CSTB2023NSCQ-LZX0085 and the Key Industrial Innovation Chain Project in Industrial Domain of Shaanxi Province (Grant No. 2020ZDLGY05-01).

References

[1]

Y. Li, W. Ma, Y. Han, A spatial prediction—based motion—compensated frame rate up—conversion, Future Internet 11 (2) (2019) 26.

[2]

K. Ntalianis, N. Mastorakis, Increasing the accuracy of video abstraction from multiple sources in the Internet of things, 2017.

[3]

A. Andreas, C.X. Mavromoustakis, G. Mastorakis, D.—T. Do, J.M. Batalla, E. Pallis, E.K. Markakis, Towards an optimized security approach to iot devices with confidential healthcare data exchange, Multimed. Tools Appl. 80 (2021) 31435-31449.

[4]

K. Hilman, H.W. Park, Y. Kim, Using motion—compensated frame—rate conversion for the correction of 3:2 pulldown artifacts in video sequences, IEEE Trans. Circuits Syst. Video Technol. 10 (2000) 869-877.

[5]

S. Hong, B. Berkeley, S.S. Kim, Motion image enhancement of lcds, in: IEEE International Conference on Image Processing 2005, vol. 2, 2005, II—17.

[6]

R. Castagno, P. Haavisto, G. Ramponi, A method for motion adaptive frame rate up—conversion, IEEE Trans. Circuits Syst. Video Technol. 6 (1996) 436-446.

[7]

S.C. Tai, Y.—R. Chen, Z.—B. Huang, C.—C. Wang, A multi—pass true motion estimation scheme with motion vector propagation for frame rate up—conversion applications, J. Disp. Technol. 4 (2008) 188-197.

[8]

T. Brox, A. Bruhn, N. Papenberg, J. Weickert, High accuracy optical flow estimation based on a theory for warping, in: European Conference on Computer Vision, 2004.

[9]

W. Weng, Y. Zhang, Z. Xiong, Event—based blurry frame interpolation under blind exposure, in: 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2023, pp. 1588-1598.

[10]

L. Sun, C. Sakaridis, J. Liang, P. Sun, J. Cao, K. Zhang, Q. Jiang, K. Wang, L.V. Gool, Event—based frame interpolation with ad—hoc deblurring, in: 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2023, pp. 18043-18052.

[11]

T. Kim, Y. Chae, H. Jang, K.—J. Yoon, Event—based video frame interpolation with cross—modal asymmetric bidirectional motion fields, in: 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2023, pp. 18032-18042.

[12]

H. Cho, T. Kim, Y. Jeong, K.—J. Yoon, Tta—evf: test—time adaptation for event—based video frame interpolation via reliable pixel and sample estimation, in: 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2024, pp. 25701-25711.

[13]

R. Jiang, F. Tu, Y. Long, A. Vaish, B. Zhou, Q. Wang, W. Zhang, Y. Fang, L.E.G. Capel, B. Mu, T. Dai, A. Suess, Evs—assisted joint deblurring, rolling—shutter correction and video frame interpolation through sensor inverse modeling, in: 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2024, pp. 25172-25181.

[14]

G. de Haan, P.W.A.C. Biezen, H. Huijgen, O.A. Ojo, True—motion estimation with 3—d recursive search block matching, IEEE Trans. Circuits Syst. Video Technol. 3 (1993) 368-379.

[15]

J. Ho Park, J. Kim, C.—S. Kim, Biformer: learning bilateral motion estimation via bilateral transformer for 4k video frame interpolation, in: 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2023, pp. 1568-1577.

[16]

S. Niklaus, F. Liu, Softmax splatting for video frame interpolation, in: 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2020, pp. 5436-5445.

[17]

J. Ho Park, K. Ko, C. Lee, C.—S. Kim, BMBC: bilateral motion estimation with bilateral cost volume for video interpolation, arXiv:2007.12622.

[18]

H. Jiang, D. Sun, V. Jampani, M.—H. Yang, E.G. Learned—Miller, J. Kautz, Super slomo: high quality estimation of multiple intermediate frames for video interpolation, in: 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2017, pp. 9000-9008.

[19]

X. Jin, L. Wu, J. Chen, Y. Chen, J. Koo, C. Hee Hahm, A unified pyramid recurrent network for video frame interpolation, in: 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022, pp. 1578-1587.

[20]

Z. Li, Z.—L. Zhu, L. Han, Q. Hou, C. Guo, M.—M. Cheng, Amt: all—pairs multi—field transforms for efficient frame interpolation, in: 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2023, pp. 9801-9810.

[21]

X. Xu, L. Siyao, W. Sun, Q. Yin, M.—H. Yang, Quadratic video interpolation, arXiv:1911.00627.

[22]

W. Bao, W.—S. Lai, C. Ma, X. Zhang, Z. Gao, M.—H. Yang, Depth—aware video frame interpolation, in: 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2019, pp. 3698-3707.

[23]

K. Zhou, W. Li, X. Han, J. Lu, Exploring motion ambiguity and alignment for high—quality video frame interpolation, in: 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022, pp. 22169-22179.

[24]

S. Niklaus, L. Mai, F. Liu, Video frame interpolation via adaptive convolution, in: 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017, pp. 2270-2279.

[25]

H. Lee, T. Kim, T.—Y. Chung, D. Pak, Y. Ban, S. Lee, Adacof: adaptive collaboration of flows for video frame interpolation, in: 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2019, pp. 5315-5324.

[26]

X. Cheng, Z. Chen, Video frame interpolation via deformable separable convolution, in: Proceedings of the AAAI Conference on Artificial Intelligence, vol. 34, 2020, pp. 10607-10614.

[27]

S. Schaub—Meyer, O. Wang, H. Zimmer, M. Grosse, A. Sorkine—Hornung, Phase—based frame interpolation for video, in: 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2015, pp. 1410-1418.

[28]

D.Y. Lee, H. Ko, J. Kim, A.C. Bovik, Space—time video regularity and visual fidelity: Compression, resolution and frame rate adaptation, 2021.

[29]

Y. Lu, G. Liang, L. Wang, Self—supervised learning of event—guided video frame interpolation for rolling shutter frames, arXiv:2306.15507.

[30]

T. Kalluri, D. Pathak, M. Chandraker, D. Tran, Flavr: flow—agnostic video representations for fast frame interpolation, in: Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, 2023, pp. 2071-2082.

[31]

J. Liu, L. Kong, B. Li, Z. Wang, H. Gu, J. Chen, Mono—vifi: a unified learning framework for self—supervised single— and multi—frame monocular depth estimation, arXiv preprint, arXiv:2407.14126.

[32]

W. Bao, W.—S. Lai, X. Zhang, Z. Gao, M.—H. Yang, Memc—net: motion estimation and motion compensation driven neural network for video interpolation and enhancement, IEEE Trans. Pattern Anal. Mach. Intell. 43 (3) (2019) 933-948.

[33]

Z. Huang, T. Zhang, W. Heng, B. Shi, S. Zhou, Rife: Real—Time Intermediate Flow Estimation for Video Frame Interpolation, 2020 (2011).

[34]

L. Dinh, D. Krueger, Y. Bengio, Nice: non—linear independent components estimation, arXiv preprint, arXiv:1410.8516.

[35]

L. Dinh, J. Sohl—Dickstein, S. Bengio, Density estimation using real nvp, arXiv preprint, arXiv:1605.08803.

[36]

L. Ardizzone, J. Kruse, C. Lüth, N. Bracher, C. Rother, U. Köthe, Conditional invertible neural networks for diverse image—to—image translation, in: Pattern Recognition: 42nd DAGM German Conference, DAGM GCPR 2020, Tübingen, Germany, September 28—October 1, 2020, Proceedings 42, Springer, 2021, pp. 373-387.

[37]

P. Wang, B. Li, T. Zhang, J. Wang, An image super—resolution reconstruction algorithm based on invertible neural network, Comput. Eng. Sci. 45 (03) (2023) 478.

[38]

M. Xiao, S. Zheng, C. Liu, Y. Wang, D. He, G. Ke, J. Bian, Z. Lin, T.—Y. Liu, Invertible Image Rescaling, in: Computer Vision—ECCV 2020: 16th European Conference, Glasgow, UK, August 23—28, 2020, Proceedings, Part I 16, Springer, 2020, pp. 126-144.

[39]

T. Xue, B. Chen, J. Wu, D. Wei, W.T. Freeman, Video enhancement with task—oriented flow, Int. J. Comput. Vis. 127 (2019) 1106-1125.

[40]

J. Park, K. Ko, C. Lee, C.—S. Kim, Bmbc: bilateral motion estimation with bilateral cost volume for video interpolation, in: Computer Vision—ECCV 2020: 16th European Conference, Glasgow, UK, August 23—28, 2020, Proceedings, Part XIV 16, Springer, 2020, pp. 109-125.

[41]

Z. Liu, R.A. Yeh, X. Tang, Y. Liu, A. Agarwala, Video frame synthesis using deep voxel flow, in: Proceedings of the IEEE International Conference on Computer Vision, 2017, pp. 4463-4471.

[42]

J. Park, C. Lee, C.—S. Kim, Asymmetric bilateral motion estimation for video frame interpolation, in: Proceedings of the IEEE/CVF International Conference on Computer Vision, 2021, pp. 14539-14548.

[43]

F. Perazzi, J. Pont—Tuset, B. McWilliams, L. Van Gool, M. Gross, A. Sorkine—Hornung, A benchmark dataset and evaluation methodology for video object segmentation, in: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2016, pp. 724-732.

[44]

D.P. Kingma, J. Ba, Adam: a method for stochastic optimization, CoRR, arXiv:1412.6980.

[45]

M. Choi, H. Kim, B. Han, N. Xu, K.M. Lee, Channel attention is all you need for video frame interpolation, in: Proceedings of the AAAI Conference on Artificial Intelligence, vol. 34, 2020, pp. 10663-10671.

[46]

S. Lee, N. Choi, W.I. Choi, Enhanced correlation matching based video frame interpolation, in: Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, 2022, pp. 2839-2847.

[47]

S. Niklaus, F. Liu, Softmax splatting for video frame interpolation, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2020, pp. 5437-5446.

[48]

H. Sim, J. Oh, M. Kim, Xvfi: extreme video frame interpolation, in: Proceedings of the IEEE/CVF International Conference on Computer Vision, 2021, pp. 14489-14498.

[49]

L. Kong, B. Jiang, D. Luo, W. Chu, X. Huang, Y. Tai, C. Wang, J. Yang, IFRnet: intermediate feature refine network for efficient frame interpolation, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022, pp. 1969-1978.

[50]

X. Jin, L. Wu, J. Chen, Y. Chen, J. Koo, C.—h. Hahm, A unified pyramid recurrent network for video frame interpolation, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 1578-1587.

[51]

X. Jin, L. Wu, G. Shen, Y. Chen, J. Chen, J. Koo, C.—h. Hahm, Enhanced bi—directional motion estimation for video frame interpolation, in: Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, 2023, pp. 5049-5057.

[52]

Z. Li, Z.—L. Zhu, L.—H. Han, Q. Hou, C.—L. Guo, M.—M. Cheng, Amt: all—pairs multi—field transforms for efficient frame interpolation, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 9801-9810.

[53]

K.L. Cheng, Y. Xie, Q. Chen, Iicnet: a generic framework for reversible image conversion, CoRR, arXiv:2109.04242.

PDF (2280KB)

2

Accesses

0

Citation

Detail

Sections
Recommended

/