Skip to main navigation Skip to search Skip to main content

MultiModFuseNet: Advancing multimodal text classification for low-resource languages through textual-visual feature fusion

  • Md Rajib Hossain
  • , Sadia Afroze
  • , Asif Ekbal
  • , Mohammed Moshiul Hoque
  • , Nazmul Siddique

Research output: Contribution to journalArticlepeer-review

Abstract

The rapid growth of social media activity and the widespread availability of electronic devices have led to an overwhelming influx of multimodal contents on the World Wide Web (WWW), much of which are unstructured, nonfactual, or toxic. Classification of such multimodal web contents (particularly, image-text) is challenging due to image-only instances with ambiguous meanings, short text-only instances lacking sufficient context and semantic clarity, shortage of annotated data and above all lack of tools for low resource languages. Manual classification of these contents is both time-consuming and costly. This paper presents MultiModFuseNet, a low-resource multimodal image-text classification system that leverages the fusion of textual and visual embeddings to overcome image ambiguity and enhance textual interpretation. MultiModFuseNet employs a systematic approach to develop a low-resource multimodal image-text corpus and identifies the best-performing image and text classification models through empirical analysis of Vision-Language Models (VLMs), Large Language Models (LLMs), Multilingual Language Models (MLMs), and Vision Transformers. Based on this analysis, MultiModFuseNet integrates pre-trained Vision Transformers for visual encoding and Multilingual Language Models for textual encoding, combining them through a trained fusion layer. This fusion layer is optimized via extensive ablation studies on the best-performing models, including tuning the learning rate, loss function, fusion technique, and embedding dimensions. MultiModFuseNet outperforms all baseline models, achieving an accuracy improvement of 4.32 % over text-only models, 12.89 % over image-only models, and 12.93 % over VLMs. The corpus and proposed solution are publicly available at: https://github.com/mrhossain/MultiModFuseNeT.
Original languageEnglish
Article number114085
Pages (from-to)1-17
Number of pages17
JournalKnowledge-Based Systems
Volume328
Early online date23 Jul 2025
DOIs
Publication statusPublished (in print/issue) - 25 Oct 2025

Bibliographical note

Publisher Copyright:
© 2025 Elsevier B.V.

Funding

This work was supported by the Directorate of Research and Extension (DRE), Chittagong University of Engineering & Technology (CUET), Chittagong, Bangladesh.

Funders
Chittagong University of Engineering and Technology

    Keywords

    • Natural language processing
    • Multimodal classification
    • Zero-short classification
    • Modality fusion
    • Ablation study
    • Fine-tuning

    Fingerprint

    Dive into the research topics of 'MultiModFuseNet: Advancing multimodal text classification for low-resource languages through textual-visual feature fusion'. Together they form a unique fingerprint.

    Cite this