The DiGreC Treebank

Morgan Macleod, Elena Anagnostopoulou, Dionysios Mertyris, Christina Sevdali

Research output: Contribution to journalArticlepeer-review

15 Downloads (Pure)

Abstract

The DiGreC (DIachrony of GREek Case) treebank is a corpus of selected sentences from Greek texts, ranging from Homer to Modern Greek, which have been annotated morphosyntactically and semantically. The corpus comprises excerpts from 655 texts, for a total of 3385 sentences and 56,440 word tokens; automated tagging and lemmatisation has been supplemented with manual review to ensure accuracy. The data exist in XML and CSV formats, which can be manipulated and converted automatically to other schemata. A web site has also been created to allow users to interact with the data more easily, and to provide specialised functionality for searching and visualisation. This corpus was created to inform theoretical debates regarding the role of case in grammar, and may be of use to researchers searching for specific attestations of a range of different constructions in Greek.
Original languageEnglish
JournalResearch Data Journal for the Humanities and Social Sciences
Publication statusAccepted/In press - 6 Oct 2021

Keywords

  • Greek
  • corpus
  • linguistics
  • classics
  • syntax
  • semantics

Fingerprint

Dive into the research topics of 'The DiGreC Treebank'. Together they form a unique fingerprint.

Cite this