• Home
  • Search
  • A Study on Webtoon Generation Using CLIP and Diffusion Models
  • Cite Icon1
  • https://doi.org/10.3390/electronics12183983Copy DOI Icon

A Study on Webtoon Generation Using CLIP and Diffusion Models

Show More
  • Abstract
  • Highlights & Summary
  • PDF
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

This study focuses on harnessing deep-learning-based text-to-image transformation techniques to help webtoon creators’ creative outputs. We converted publicly available datasets (e.g., MSCOCO) into a multimodal webtoon dataset using CartoonGAN. First, the dataset was leveraged for training contrastive language image pre-training (CLIP), a model composed of multi-lingual BERT and a Vision Transformer that learnt to associate text with images. Second, a pre-trained diffusion model was employed to generate webtoons through text and text-similar image input. The webtoon dataset comprised treatments (i.e., textual descriptions) paired with their corresponding webtoon illustrations. CLIP (operating through contrastive learning) extracted features from different data modalities and aligned similar data more closely within the same feature space while pushing dissimilar data apart. This model learnt the relationships between various modalities in multimodal data. To generate webtoons using the diffusion model, the process involved providing the CLIP features of the desired webtoon’s text with those of the most text-similar image to a pre-trained diffusion model. Experiments were conducted using both single- and continuous-text inputs to generate webtoons. In the experiments, both single-text and continuous-text inputs were used to generate webtoons, and the results showed an inception score of 7.14 when using continuous-text inputs. The text-to-image technology developed here could streamline the webtoon creation process for artists by enabling the efficient generation of webtoons based on the provided text. However, the current inability to generate webtoons from multiple sentences or images while maintaining a consistent artistic style was noted. Therefore, further research is imperative to develop a text-to-image model capable of handling multi-sentence and -lingual input while ensuring coherence in the artistic style across the generated webtoon images.

Loading PDF

Similar Papers
  • Conference Article

CrossEM: A Prompt Tuning Framework for Cross-Modal Entity Matching

  • May 19, 2025
  • Yuan Qin +4
  • Dissertation

Human-centric 3D representation learning

  • Jan 01, 2025
  • Fangzhou Hong
  • Conference Article
  • Citations2

Analysing and Evaluating Complementarity of Multi-Modal Data Fusion in AD Diagnosis

  • Dec 01, 2022
  • Zhaodong Chen +4
  • Book Chapter

A New Method to Address Singularity Problem in Multimodal Data Analysis

  • Jan 01, 2017
  • Ankita Mandal +1
  • Research Article
  • Citations2

A Novel Hybrid Attention-Based Dilated Network for Depression Classification Model from Multimodal Data Using Improved Heuristic Approach

  • Jul 10, 2024
  • International Journal of Image and Graphics
  • B Manjulatha +1
  • Conference Article
  • Citations13

DCMN: Double Core Memory Network for Patient Outcome Prediction with Multimodal Data

  • Nov 01, 2019
  • Yujuan Feng +6
  • Research Article
  • Citations9

Multimodal Visual Data Registration for Web-Based Visualization in Media Production

  • Apr 01, 2018
  • IEEE Transactions on Circuits and Systems for Video Technology
  • Hansung Kim +3
  • Conference Article

RePaint-Enhanced Conditional Diffusion Model for Generating Designs Under Performance Constraints

  • Aug 17, 2025
  • Ke Wang +4
  • Research Article

Partitioning the data space before applying hashingusing clustering algorithms

  • Apr 04, 2025
  • Herald of Advanced Information Technology
  • Sergey А Subbotin +1
  • PDF
  • Research Article
  • Citations2

Robust and Fast Sensing of Urban Flood Depth with Social Media Images Using Pre-Trained Large Models and Simple Edge Training

  • Nov 17, 2025
  • Hydrology
  • Lin Lin +4
  • Research Article
  • Citations8

Accelerating Text-to-Image Editing via Cache-Enabled Sparse Diffusion Inference

  • Mar 24, 2024
  • Proceedings of the AAAI Conference on Artificial Intelligence
  • Zihao Yu +4
  • PDF
  • Research Article
  • Citations88

Multimodal data indicators for capturing cognitive, motivational, and emotional learning processes: A systematic literature review

  • May 30, 2020
  • Education and Information Technologies
  • Omid Noroozi +5
  • Research Article
  • Citations22

Integration of Multi-Modal Biomedical Data to Predict Cancer Grade and Patient Survival.

  • Feb 01, 2016
  • ... IEEE-EMBS International Conference on Biomedical and Health Informatics. IEEE-EMBS International Conference on Biomedical and Health Informatics
  • John H Phan +4
  • Research Article

Investigation of Features Influencing on the Artworks of African American Female Artists Between 1960s and 2000s

  • Oct 01, 2024
  • Art and Society
  • Duofei Cui +1
  • Research Article
  • Citations13

GeoPredict-LLM: Intelligent tunnel advanced geological prediction by reprogramming large language models

  • Nov 02, 2024
  • Intelligent Geoengineering
  • Zhenhao Xu +4
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.