How well does the current state of Irish Gaelic NLP align with the National Plan for Irish Language Public Services?

As part of its 6 year National Plan for Irish Language Public Services [1], the Irish government is actively implementing several strategies to integrate and develop tools and resources for the Irish Gaelic language in the public service sector. One of the major strategic themes of this plan is Technology and the active integration of the Irish language in the development of language technologies [2]. However, research in the field of Natural Language Processing (NLP) does not always address use-cases on the ground, especially for low-resourced languages [3,4]. Further, little is known about how well the progress of NLP systems aligns with government policies. In this project, we will investigate the extent to which the current research in NLP for the Irish Gaelic language aligns with the aims and specific problem statements in the National Plan. You will, 1) identify language technology priorities in the National plan, 2) collate relevant literature in NLP on the Irish Gaelic language, 3) map existing research and reported performance against these priorities, and 4) propose focus areas for future work in Gaelic NLP. This project involves collecting and analysing academic literature from sources such as the ACL Anthology and relevant Celtic-language workshops. Depending on the size of the dataset, automated methods such as topic modelling may be used to identify research themes.

 

  1. National Plan for Irish Language Public Services. https://assets.gov.ie/static/documents/cdb3dc1c/National_Plan_for_Irish_Language_Public_Services.pdf
  2. Harshita Diddee, Kalika Bali, Monojit Choudhury, and Namrata Mukhija. 2022. The Six Conundrums of Building and Deploying Language Technologies for Social Good. In Proceedings of the 5th ACM SIGCAS/SIGCHI Conference on Computing and Sustainable Societies (COMPASS ’22). Association for Computing Machinery, New York, NY, USA, 12–19. https://doi.org/10.1145/3530190.3534792
  3. Hellina Hailu Nigatu, Atnafu Lambebo Tonja, Benjamin Rosman, Thamar Solorio, and Monojit Choudhury. 2024. The Zeno’s Paradox of ‘Low-Resource’ Languages. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 17753–17774, Miami, Florida, USA. Association for Computational Linguistics.
  4. ACL Anthology. https://aclanthology.org/