Principles of Building Syntactic and Semantic Corpus for Large Language Models (LLM) of Low-Resource Languages in the Era of Digitalization
Abstract. In the modern era, the rapid development of large language models in the ecosystem of artificial intelligence and natural language processing makes the issue of preserving and promoting low-resource languages in the digital environment urgent. Against the background of global technological trends, the lack of high-quality syntactic and semantic corpora for low-resource languages significantly limits the integration of these languages into advanced neural network architectures. The main goal of this article is to develop theoretical and methodological principles for building syntactic and semantic corpora for low-resource languages, including the Azerbaijani language, in the context of digital transformation. The stages of data collection, cleaning, annotation, and preparation for model training are covered in detail in the research work. As a methodological basis, the classical principles of corpus linguistics and the requirements of modern deep learning and transformer-based models are synthesized. The results of the study show that building syntactic tree banks and semantic relationship graphs for low-resource languages with morphological agglutinative properties increases the generation quality and contextual understanding ability of large language models by at least 35-40%. At the same time, the proposed models serve to preserve the purity of language in the digital environment and minimize the hallucinations that arise in artificial intelligence systems. The discussions show that the application of semi-automatic annotation methods based on human-computer collaboration (human-in-the-loop) in the corpus construction process significantly increases the accuracy of the results.
An analysis of the international digital experience of Basque, Irish, Welsh and other low-resource languages from the Turkic group proves that the prevention of syntactic interference and the application of special morphological segmentation algorithms are among the main factors ensuring the effectiveness of LLMs on a global scale. The scientific and methodological recommendations put forward at the end of the article create a broad basis for expanding open scientific resources for local researchers, adapting international best practices to the national level, and increasing the efficiency of future NLP projects.
Keywords: large language models (LLM), low-resource languages, syntactic corpus, semantic annotation, digital transformation