Abstract
The rapid growth of social media activity and the widespread availability of electronic devices have led to an overwhelming influx of multimodal contents on the World Wide Web (WWW), much of which are unstructured, nonfactual, or toxic. Classification of such multimodal web contents (particularly, image-text) is challenging due to image-only instances with ambiguous meanings, short text-only instances lacking sufficient context and semantic clarity, shortage of annotated data and above all lack of tools for low resource languages. Manual classification of these contents is both time-consuming and costly. This paper presents MultiModFuseNet, a low-resource multimodal image-text classification system that leverages the fusion of textual and visual embeddings to overcome image ambiguity and enhance textual interpretation. MultiModFuseNet employs a systematic approach to develop a low-resource multimodal image-text corpus and identifies the best-performing image and text classification models through empirical analysis of Vision-Language Models (VLMs), Large Language Models (LLMs), Multilingual Language Models (MLMs), and Vision Transformers. Based on this analysis, MultiModFuseNet integrates pre-trained Vision Transformers for visual encoding and Multilingual Language Models for textual encoding, combining them through a trained fusion layer. This fusion layer is optimized via extensive ablation studies on the best-performing models, including tuning the learning rate, loss function, fusion technique, and embedding dimensions. MultiModFuseNet outperforms all baseline models, achieving an accuracy improvement of 4.32 % over text-only models, 12.89 % over image-only models, and 12.93 % over VLMs. The corpus and proposed solution are publicly available at: https://github.com/mrhossain/MultiModFuseNeT.
| Original language | English |
|---|---|
| Article number | 114085 |
| Pages (from-to) | 1-17 |
| Number of pages | 17 |
| Journal | Knowledge-Based Systems |
| Volume | 328 |
| Early online date | 23 Jul 2025 |
| DOIs | |
| Publication status | Published (in print/issue) - 25 Oct 2025 |
Bibliographical note
Publisher Copyright:© 2025 Elsevier B.V.
Funding
This work was supported by the Directorate of Research and Extension (DRE), Chittagong University of Engineering & Technology (CUET), Chittagong, Bangladesh.
| Funders |
|---|
| Chittagong University of Engineering and Technology |
Keywords
- Natural language processing
- Multimodal classification
- Zero-short classification
- Modality fusion
- Ablation study
- Fine-tuning
Fingerprint
Dive into the research topics of 'MultiModFuseNet: Advancing multimodal text classification for low-resource languages through textual-visual feature fusion'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver