• Home
  • Search
  • Multimodal Llms For Image-Based Fabric Composition Identification
  • https://doi.org/10.29284/n8jk2q41Copy DOI Icon

Multimodal Llms For Image-Based Fabric Composition Identification

  • Abstract
  • Literature Map
  • Similar Papers
Abstract

With e-commerce shaping the way people buy clothes today, correctly determining what fabrics are made of is now a key requirement. Conventional methods are often constrained by hardware dependencies (e.g, NIR sensors), limited accessibility and poor fabric classification. We propose a parameter-efficient multimodal system that uses instruction-tuned LLMs to determine the proportion of fibers directly from fabric images. We compare two different multimodal models. The first is Meta’s LLaMA 3.2-Vision 11B Instruct model, which builds on a separately trained image encoder and a vision adapter that integrates image-feature representations via cross-attention layers into a core language model. The second is Salesforce’s BLIP2, which connects a ViT-G/14 image encoder to an OPT language model. In both the models, a light-weight regression head is appended to the final pooled representations and trained alongside LoRA injected adapters targeting the query, key and value projections within the transformer layers. The LoRA configuration is implemented using the PEFT (Parameter-Efficient Fine-Tuning), reducing trainable parameters to a few million for efficient adaptation. Experiments span six fiber categories -cotton, wool, polyester, linen, silk, and other fibers using a custom dataset of labelled textile images. The BLIP2-based model achieved a mean absolute error of 13.40 percentage points, outperforming the LLaMA 3.2 model, which recorded 14.87 percentage points on the test set. The evaluation results validate the use of vision-language models for estimating fiber composition and point toward their potential adoption in both consumer and industrial applications.

Similar Papers
  • Research Article

Multimodal foundation model for evaluating cardiovascular disease from electrocardiography and echocardiography

  • Nov 01, 2025
  • European Heart Journal
  • Arash Aghajani Nargesi +8
  • Research Article
  • Citations5

Diagnostic efficiency of multi-modal MRI based deep learning with Sobel operator in differentiating benign and malignant breast mass lesions-a retrospective study.

  • Jul 17, 2023
  • PeerJ Computer Science
  • Weixia Tang +7
  • Research Article
  • Citations9

Postoperative Karnofsky performance status prediction in patients with IDH wild-type glioblastoma: A multimodal approach integrating clinical and deep imaging features.

  • Nov 11, 2024
  • PloS one
  • Tomoki Sasagasako +10
  • Research Article
  • Citations2

AI-based multimodal prediction of lymph node metastasis and capsular invasion in cT1N0M0 papillary thyroid carcinoma

  • May 27, 2025
  • Frontiers in Endocrinology
  • Xiaowei Peng +8
  • Research Article
  • Citations1

Multimodal prediction models integrating radiomics and three-dimensional deep learning for acute respiratory distress syndrome in acute pancreatitis patients.

  • Jan 01, 2026
  • The Journal of international medical research
  • Jielu Zhou +11
  • PDF
  • Research Article
  • Citations11

Preoperative prediction of extrathyroidal extension: radiomics signature based on multimodal ultrasound to papillary thyroid carcinoma

  • Jul 20, 2023
  • BMC Medical Imaging
  • Fang Wan +5
  • Research Article
  • Citations1

Combining radiomics of X-rays with patient functional rating scales for predicting satisfaction after radial fracture fixation: a multimodal machine learning predictive model.

  • Sep 30, 2025
  • BMC musculoskeletal disorders
  • Changsen Yang +5
  • Research Article
  • Citations9

Multimodal ultrasound deep learning to detect fibrosis in early chronic kidney disease

  • Oct 22, 2024
  • Renal Failure
  • Xiachuan Qin +4
  • Research Article

Training language models to be warm can reduce accuracy and increase sycophancy.

  • Apr 01, 2026
  • Nature
  • Lujain Ibrahim +2
  • Conference Article
  • Citations23

Log-linear model combination with word-dependent scaling factors

  • Sep 06, 2009
  • Björn Hoffmeister +3
  • PDF
  • Conference Article
  • Citations80

Simple Fusion: Return of the Language Model

  • Jan 01, 2018
  • Felix Stahlberg +2
  • Research Article

The Power of Multimodality in Multimodal Large Language Models, Unimodal ChatGPT 5.0, and Human Clinical Experts on a Wound Care Certification Examination: Cross-Sectional Comparative Study.

  • Apr 27, 2026
  • JMIR formative research
  • Mete Ucdal +5
  • Research Article

Large-scale evaluation of multimodal large language models for pneumothorax detection.

  • Feb 12, 2026
  • Diagnostic and interventional radiology (Ankara, Turkey)
  • Hamza Eren Güzel +2
  • Research Article
  • Citations64

Image captioning for effective use of language models in knowledge-based visual question answering

  • Aug 28, 2022
  • Expert Systems with Applications
  • Ander Salaberria +4
  • Research Article

Prediction of lymphovascular invasion in rectal cancer based on multimodal magnetic resonance imaging radiomics model

  • Feb 27, 2026
  • World Journal of Gastrointestinal Surgery
  • Zheng-Hong Zhu +2
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.