Machine Translation and Automated Analysis of Cuneiform Languages

Digging into Data Challenge project (2017–2021) that produced the first machine translation system for Sumerian and a syntactically annotated Ur III corpus.

Machine Translation and Automated Analysis of Cuneiform Languages
  • Role: Project Coordinator; manager of the UCLA team and coordinator of the Toronto team
  • Funding: NEH, DFG and SSHRC through the Trans-Atlantic Platform Digging into Data Challenge, about US$550,000, 2017–2021
  • Investigators: Heather D. Baker (University of Toronto, Principal Investigator), with Robert K. Englund (UCLA) and Christian Chiarcos (Goethe University Frankfurt) as Co-Principal Investigators

MTAAC built a natural language processing pipeline for the Sumerian administrative texts of the Ur III period (about 21st century BCE), the largest body of texts in the language. Building on the lemmatised corpora already available through Oracc, it produced a manually annotated gold corpus with dependency syntax, morphological and syntactic pre-annotation tools, part-of-speech tagging and named-entity recognition, and the first machine translation system from Sumerian transliterations to English, all released as open code and linked open data through the CDLI.

Sumerian translation pipeline · Project results (DFG)