Towards an Automatic Classification of Illustrative Examples in a Large Japanese-French Dictionary Obtained by OCR

Autor: Tomokiyo, Mutsuko, Boitet, Christian, Mangeot, Mathieu
Přispěvatelé: Groupe d’Étude en Traduction Automatique/Traitement Automatisé des Langues et de la Parole (GETALP ), Laboratoire d'Informatique de Grenoble (LIG ), Institut polytechnique de Grenoble - Grenoble Institute of Technology (Grenoble INP )-Centre National de la Recherche Scientifique (CNRS)-Université Grenoble Alpes [2016-2019] (UGA [2016-2019])-Institut polytechnique de Grenoble - Grenoble Institute of Technology (Grenoble INP )-Centre National de la Recherche Scientifique (CNRS)-Université Grenoble Alpes [2016-2019] (UGA [2016-2019]), Institut polytechnique de Grenoble - Grenoble Institute of Technology (Grenoble INP )-Centre National de la Recherche Scientifique (CNRS)-Université Grenoble Alpes [2016-2019] (UGA [2016-2019]), Université Savoie Mont Blanc (USMB [Université de Savoie] [Université de Chambéry])
Jazyk: angličtina
Rok vydání: 2018
Předmět:
Zdroj: Proceedings of the First Workshop on Linguistic Resources for Natural Language Processing
First Workshop on Linguistic Resources for Natural Language Processing
First Workshop on Linguistic Resources for Natural Language Processing, Aug 2018, Santa Fe, United States. pp.112-121
Popis: International audience; This paper focuses on improving the Cesselin, a large, open source Japanese-French bilingual dictionary digitalized by OCR, available on the web, and contributively improvable online. Labelling its examples (about 226,000) would significantly enhance their usefulness for language learners. Examples are proverbs, idiomatic constructions, normal usage examples, and, for nouns, phrases containing a quantifier. Proverbs are easy to spot, but not the other types. To find a method for automatically or at least semi-automatically annotating them, we have studied many entries, and hypothesized that the degree of lexical similarity between results of MT into a third language might give good cues. To confirm that hypothesis, we sampled 500 examples and used Google Translate to translate into English the Cesslin Japanese expressions and their French translations. The hypothesis holds well, in particular for distinguishing examples of normal usage from idiomatic examples. Finally, we propose a detailed annotation procedure and discuss its future automatization.
Databáze: OpenAIRE