Impact of Digitalization on Language: Artificial Intelligence’s Level of "Comprehending" Semantic Integrity and Syntactic Adequacy in Language
Yegana Gahramanova
Abstract. In the era of digitalization, the development of Natural Language Processing (NLP) and Large Language Models (LLMs) has become a central research object of digital linguistics. This research analyzes, on a comparative scale, the degree to which Artificial Intelligence (AI) systems and Neural Machine Translation (NMT) comprehend semantic integrity and syntactic adequacy across languages with different typological structures. The main objective of the study is to explore the interaction between the morphosyntactic characteristics of analytic (English), inflected (Russian, German), and agglutinative (Azerbaijani, Turkish, Finnish) languages with AI's Transformer architecture, as well as to uncover specific semantic hallucinations and syntactic deviations arising in the case of the Azerbaijani language. Structured according, the study employs qualitative and quantitative content analysis, typological comparison, and error typology methods. Empirical analyses demonstrate that while AI models' subword tokenization (BPE/WordPiece) algorithms exhibit high precision in analytic languages, they cause lexical-semantic fragmentation in agglutinative languages such as Azerbaijani due to polyaffixation (polymorphism). Simultaneously, the Subject-Object-Verb (SOV) structure of the Azerbaijani language leads to the disruption of person and tense agreement of sentence-final verbs in long-range dependencies. As a result of the study, the necessity of enriching language-internal parallel corpora and developing morphology-aware tokenizers for low-resource agglutinative languages is substantiated.
Keywords: digital linguistics, artificial ıntelligence, large language models (LLMs), agglutinative languages, Azerbaijani language, semantic integrity, syntactic adequacy, Neural Machine Translation (NMT)