Skip to main navigation Skip to search Skip to main content

Improved complex words simplification in non-annotated text using a deep-learning-based lexical complexity prediction model

Research output: Contribution to journalArticlepeer-review

Abstract

Replacing complex words with simpler expressions is an effective way to improve readability and language understanding. We propose combining multiple levels of distributional semantics to measure word complexity beyond superficial lexical features. Without relying on predefined complexity lexicons, our lexical complexity prediction (LCP) model achieves performance comparable to human annotations and supports substitution retrieval for text simplification (TS). Unlike machine-translation-driven TS models trained on annotated corpora, our framework contains two modules: a supervised LCP module and an unsupervised simplification module. In the 2024 Multilingual Lexical Simplification Pipeline English Shared Task, the proposed LCP model is competitive with other state-of-the-art models. Additionally, experiments on TurkCorpus show that our model performs competitively with other popular supervised models while satisfying the requirements of TS in terms of adequacy, fluency, and simplicity.
Original languageEnglish
Article number357
Pages (from-to)1-20
Number of pages20
JournalInternational Journal of Machine Learning and Cybernetics
Volume17
Issue number357
Early online date23 Jun 2026
DOIs
Publication statusPublished (in print/issue) - 23 Jun 2026

Bibliographical note

© The Author(s), under exclusive licence to Springer-Verlag GmbH Germany, part of Springer Nature 2026

Data Availability Statement

This study used publicly available datasets: TurkCorpus (https://huggingface.co/datasets/waboucay/turk_corpus) and CompLex (https://github.com/MMU-TDMLab/CompLex). No new datasets were generated or analyzed during the current study.

Funding

We gratefully acknowledge the valuable support provided by Ulster University and Shandong Jianzhu University, which has been instrumental in facilitating this work. Their contribution has significantly enhanced the progress and outcomes of our research.

Funders
Shandong Jianzhu University

    UN SDGs

    This output contributes to the following UN Sustainable Development Goals (SDGs)

    1. SDG 9 - Industry, Innovation, and Infrastructure
      SDG 9 Industry, Innovation, and Infrastructure

    Keywords

    • Text simplification
    • Lexical simplification
    • Lexical complexity prediction
    • Complex word identification

    Fingerprint

    Dive into the research topics of 'Improved complex words simplification in non-annotated text using a deep-learning-based lexical complexity prediction model'. Together they form a unique fingerprint.

    Cite this