Skip to main navigation Skip to search Skip to main content

Automatic Classification of National Health Service Feedback

  • Christopher Haynes
  • , Marco A Palomino
  • , Liz Stuart
  • , David Viira
  • , Frances Hannon
  • , Gemma Crossingham
  • , Kate Tantam

Research output: Contribution to journalArticlepeer-review

2 Downloads (Pure)

Abstract

Text datasets come in an abundance of shapes, sizes and styles. However, determining what factors limit classification accuracy remains a difficult task which is still the subject of intensive research. Using a challenging UK National Health Service (NHS) dataset, which contains many characteristics known to increase the complexity of classification, we propose an innovative classification pipeline. This pipeline switches between different text pre-processing, scoring and classification techniques during execution. Using this flexible pipeline, a high level of accuracy has been achieved in the classification of a range of datasets, attaining a micro-averaged F1 score of 93.30% on the Reuters-21578 “ApteMod” corpus. An evaluation of this flexible pipeline was carried out using a variety of complex datasets compared against an unsupervised clustering approach. The paper describes how classification accuracy is impacted by an unbalanced category distribution, the rare use of generic terms and the subjective nature of manual human classification.
Original languageEnglish
Article number983
Pages (from-to)1-23
Number of pages23
JournalMathematics
Volume10
Issue number6
Early online date18 Mar 2022
DOIs
Publication statusPublished online - 18 Mar 2022

Data Availability Statement

“Reuters-21578 (ApteMod)”: Used in our software package via Python
NLTKplatform,directdownloadavailableathttp://kdd.ics.uci.edu/databases/reuters21578/reuters21578.html. Retrieved 25 April 2021 “Amazon Hierarchical Reviews”: Yury Kashnitsky. Hierarchical text classification. (April 2020). Version 1. Retrieved 29 April 2021 from https://www.kaggle.com/kashnitsky/hierarchical-text-classification/version/1. “Twitter COVID Sentiment”: Aman Miglani.
Coronavirus tweet NLP—Text Classification. (September 2020). Version 1. Retrieved 27 April 2021 from https://www.kaggle.com/datatattle/covid-19-nlp-text-classification/version/1. “Twitter Tweet Genre”: Pradeep. Text (Tweet) Classification (January 2020). Version 1. Retrieved 28 April 2021 from
https://www.kaggle.com/pradeeptrical/text-tweet-classification/version/1

Funding

This research received no external funding.

UN SDGs

This output contributes to the following UN Sustainable Development Goals (SDGs)

  1. SDG 3 - Good Health and Well-being
    SDG 3 Good Health and Well-being

Keywords

  • NLP
  • classification
  • clustering
  • text pre-processing
  • machine learning
  • National Health Service (NHS)

Fingerprint

Dive into the research topics of 'Automatic Classification of National Health Service Feedback'. Together they form a unique fingerprint.

Cite this