• Home
  • Search
  • Tuning-free High-Resolution Video Diffusion with Spatial-Temporal Latent Grouping
  • https://doi.org/10.1109/tmm.2025.3618540Copy DOI Icon

Tuning-free High-Resolution Video Diffusion with Spatial-Temporal Latent Grouping

Show More
  • Abstract
  • Literature Map
  • References
  • Similar Papers
Abstract

Recent advances in text-to-video generation have demonstrated the substantial superiority of diffusion models. Nevertheless, generating high-resolution videos based on text description still faces a great challenge due to the enormous computation overhead for video diffusion model training. In this paper, we present a tuning-free video diffusion approach with Spatial-Temporal LAtent Grouping (ST-LAG), for highresolution video generation. ST-LAG exploits the prior knowledge of a pre-trained low-resolution video diffusion model for regionwise video latent denoising, and then combines all the denoised regions of video latent as a whole one to achieve global-wise spatial-temporal coherence. Specifically, ST-LAG denoises the whole video latents via two deliberately designed modules, e.g., Spatial Latent Grouping (SLG) and Temporal Latent Grouping (TLG), at spatial and temporal level, respectively. SLG spatially slices the latent of each frame into different local patches, and then feeds them into the low-resolution video diffusion model for local-region latent denoising. A text re-weighting scheme is devised in SLG to strength the cross-attention between features of text tokens and spatial regions to facilitate spatial-level finegrained details generation. TLG capitalizes on the segmentlevel latent grouping to match the length of each denoised local segment with the frame number in the training stage. The well aligned temporal receptive field facilitates better preservation of motion patterns. In each denoising step, all groups of video latent at spatial and temporal levels are fused together for highresolution video generation. Extensive experiments conducted on the ECTV-Prompt dataset demonstrate the effectiveness of our approach quantitatively and qualitatively.

Similar Papers
  • Research Article
  • Citations22

Stimulating Diffusion Model for Image Denoising via Adaptive Embedding and Ensembling.

  • Dec 01, 2024
  • IEEE transactions on pattern analysis and machine intelligence
  • Tong Li +5
  • PDF
  • Research Article
  • Citations1

A Study on Webtoon Generation Using CLIP and Diffusion Models

  • Sep 21, 2023
  • Electronics
  • Kyungho Yu +4
  • Preprint Article

DiffSD: Diffusion models for seismic denoising

  • May 15, 2023
  • Daniele Trappolini +5
  • Conference Article

RePaint-Enhanced Conditional Diffusion Model for Generating Designs Under Performance Constraints

  • Aug 17, 2025
  • Ke Wang +4
  • Research Article
  • Citations1

Hand1000: Generating Realistic Hands from Text with Only 1,000 Images

  • Apr 11, 2025
  • Proceedings of the AAAI Conference on Artificial Intelligence
  • Haozhuo Zhang +3
  • Research Article
  • Citations1

Pelatihan Menulis “Descriptive Text” di Kelurahan Manguharjo Kecamatan Mayangan Kota Probolinggo

  • Nov 14, 2020
  • Jurnal Abdi Panca Mara
  • Utami Ratna Swari
  • Conference Article
  • Citations56

Space-Time Video Super-Resolution Using Temporal Profiles

  • Oct 12, 2020
  • Zeyu Xiao +4
  • Research Article

Talking-DiSSM: Enhancing Temporal Consistency in Talking Face Video Generation with Bidirectional SSMs

  • Jul 22, 2025
  • ACM Transactions on Intelligent Systems and Technology
  • Zhen Xiao +5
  • Research Article
  • Citations152

A novel end-to-end 1D-ResCNN model to remove artifact from EEG signals

  • Apr 23, 2020
  • Neurocomputing
  • Weitong Sun +3
  • PDF
  • Research Article
  • Citations70

TiVGAN: Text to Image to Video Generation With Step-by-Step Evolutionary Generator

  • Jan 01, 2020
  • IEEE Access
  • Doyeon Kim +2
  • Research Article
  • Citations9

Diffusion Models for 3D Generation: A Survey

  • Feb 01, 2025
  • Computational Visual Media
  • Chen Wang +4
  • Research Article
  • Citations1

Continuous Facial Motion Deblurring

  • Jan 01, 2022
  • IEEE Access
  • Tae Bok Lee +2
  • Conference Article
  • Citations2

Hyperspectral image segmentation using active contours

  • Aug 12, 2004
  • Proceedings of SPIE, the International Society for Optical Engineering/Proceedings of SPIE
  • Cheolha P Lee +1
  • Research Article
  • Citations8

Accelerating Text-to-Image Editing via Cache-Enabled Sparse Diffusion Inference

  • Mar 24, 2024
  • Proceedings of the AAAI Conference on Artificial Intelligence
  • Zihao Yu +4
  • Research Article
  • Citations3

Text to Image Generation: A Literature Review Focus on the Diffusion Model

  • Jan 01, 2025
  • ITM Web of Conferences
  • Jingxi Zhou
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.