Skip to main navigation Skip to search Skip to main content

Meaning from text: a sentiment classification framework

  • Luis Trindade

    Student thesis: Doctoral Thesis

    Abstract

    Whether conducting research or making everyday decisions we often look for other people’s opinions. The Web 2.0 and the rise of social media have fuelled an increase in reviews, ratings, recommendations and other forms of online expression and online opinion. However, the manual opinion mining process can be lengthy and frustrating, as there is potentially a huge amount of data and noise to navigate through. This problem gave rise to a new area in Text Analytics, Sentiment Analysis (SA). SA, also known as opinion mining, expanded the subject of study from a traditionally fact and information centric view of text to a sentiment/opinion aware view of text.

    SA makes use of Information Retrieval, Machine Learning (ML), and Computational Linguistics (CL) techniques to detect and extract opinionated information contained within a piece of text. In general, ML involves the construction and study of algorithms which can automatically learn from and make predictions on data, as opposed to simply following static instructions. An important aspect in the application of ML methods to text classification (and sentiment classification) is how to represent the text. For the purposes of Natural Language Processing and ML tasks, mathematical data structures (e.g. vectors, sequences and graphs) are commonly employed to represent the text. Each text representation has its advantages and disadvantages.

    This Thesis proposes novel and effective sentiment classification approaches, which incorporate syntactic, semantic and sentiment information in commonly used text representations. It identifies linguistic features that contain syntactic, semantic and sentiment information relevant to sentiment classification. Text representation schemes that can represent these linguistic features are studied, and ML methods that can make use of such text representations are studied, adapted and extended to incorporate the extra linguistic features considered. In particular, a substructure-based text representation is studied, which is effectively an extended bag-of-words text representation. The proposed approaches are evaluated on common benchmark datasets, and state-of-the-art performance is demonstrated.
    Date of AwardMay 2015
    Original languageEnglish
    SupervisorHui Wang (Supervisor), William Blackburn (Supervisor), Hui Wang (Supervisor), William Blackburn (Supervisor) & NIALL ROONEY (Supervisor)

    Keywords

    • sentiment analysis
    • classification
    • opinion mining
    • kernel methods

    Cite this

    '