267 AMERICAN Journal of Language, Literacy and Learning in STEM Education www. grnjournal.us AMERICAN Journal of Language, Literacy and Learning in STEM Education Volume 3, Issue 7, 2025 ISSN (E): 2993-2769 Artificial Intelligence and Linguistics: Challenges and Opportunities in Natural Language Processing Nisreen Khalid Abbas College of Education, University of Samarra, Iraq Abstract. Natural Language Processing (NLP), situated at the intersection of artificial intelligence and linguistics, has witnessed unprecedented growth in recent years. Advances in deep learning, large language models, and computational linguistics have transformed how machines understand and generate human language. This review explores the challenges and opportunities arising from this interaction. Linguistic diversity, ambiguity, and context sensitivity remain major obstacles in achieving human-like comprehension, while ethical issues such as bias, cultural representation, and misuse of AI systems present additional hurdles. On the other hand, opportunities emerge in areas such as cross-linguistic communication, intelligent tutoring systems, sentiment analysis, healthcare applications, and automated translation. By bridging theoretical linguistics with applied AI methodologies, future research can foster more robust, fair, and contextually aware NLP systems that advance both linguistic theory and real-world applications. Key words: Artificial Intelligence; Linguistics; Natural Language Processing; Deep Learning; Computational Linguistics; Language Models. Introduction The convergence of artificial intelligence (AI) and linguistics has led to remarkable progress in Natural Language Processing (NLP)[1], enabling machines to perform tasks once thought uniquely human, such as speech recognition, sentiment analysis, and automated translation. Linguistics provides the theoretical foundation for understanding phonology, morphology[2], syntax, semantics, and pragmatics, while AI offers computational methods to operationalize these concepts. Recent breakthroughs in large-scale pre-trained models, such as GPT, BERT, and multilingual systems, have accelerated research and applications. However, this integration presents critical challenges[3], including linguistic diversity, low-resource languages, semantic ambiguity, and ethical implications. This review article critically examines these challenges and highlights future opportunities for research and practice[4]. Artificial Intelligence and Linguistics: A Historical Overview of Convergence in NLP The relationship between artificial intelligence (AI) and linguistics has evolved over several decades, shaped by both theoretical developments in linguistic science and technological breakthroughs in computing. Natural Language Processing (NLP), as a subfield at this intersection, has progressed through multiple distinct stages. Each period brought new methods, challenges[5], and opportunities that redefined how machines process and “understand” human language. The following sections outline the chronological evolution of AI in relation to linguistics, from symbolic systems to neural architectures and beyond[5]. 1. The Symbolic Era (1950s – 1970s): Rule-Based Systems and Early Linguistics The earliest phase of AI-linguistics interaction was rooted in symbolic AI and formal linguistics. Inspired by Noam Chomsky’s generative grammar (1957), researchers sought to model language using explicit grammatical rules. NLP systems during this period relied on hand-crafted rules for 268 AMERICAN Journal of Language, Literacy and Learning in STEM Education www. grnjournal.us parsing and translation. One of the most notable examples was the Georgetown-IBM experiment (1954), where a machine translated over 60 Russian sentences into English using rule-based algorithms[6]. ➢ Linguistic foundation: Chomskyan syntax, phrase structure grammar[7]. ➢ Technological approach: Deterministic rule engines and pattern matching[8]. ➢ Limitations: Inability to handle ambiguity, idioms, or context; poor scalability[9]. This period established the theoretical backbone of computational linguistics but failed to achieve robust real-world performance due to the brittleness of symbolic systems[10]. 2. The Statistical Revolution (1980s – 1990s): Data-Driven NLP With the rise of machine learning and increased access to large text corpora (e.g., Brown Corpus, Penn Treebank), the field shifted toward statistical NLP. Instead of manually encoding linguistic rules, systems began to learn from data using probabilistic models such as Hidden Markov Models (HMMs) and n-gram language models[11]. ➢ Linguistic integration: Focus on syntax and part-of-speech tagging, informed by annotated corpora. ➢ Technological shift: From symbolic logic to probabilistic inference and maximum-likelihood estimation. ➢ Notable applications: Speech recognition, POS tagging, named entity recognition. ➢ Limitations: Difficulty modeling long-range dependencies; poor semantic understanding. This era marked the first real fusion of linguistics with statistical AI, enabling scalable solutions but still lacking deep semantic awareness[12]. 269 AMERICAN Journal of Language, Literacy and Learning in STEM Education www. grnjournal.us 3. The Neural Turn (2000s – 2010s): Deep Learning and Distributed Representations The 2000s introduced a paradigm shift with the emergence of neural networks and distributed representations (word embeddings). Seminal work like word2vec allowed words to be represented as continuous vectors capturing semantic similarity. Recurrent Neural Networks (RNNs) and Long Short-Term Memory (LSTM) models became dominant for sequence modeling tasks[13]. ➢ Linguistic application: Semantic similarity, syntactic parsing, sentiment analysis. ➢ Technological leap: Use of gradient-based optimization and backpropagation in large datasets. ➢ Notable systems: Google’s Neural Machine Translation (GNMT), deep sentiment classifiers. ➢ Limitations: Data-hungry models, limited interpretability, lack of contextual nuance. This phase brought semantic richness and improved accuracy in NLP tasks but introduced new challenges like bias in embeddings and opaque model behavior[14]. 270 AMERICAN Journal of Language, Literacy and Learning in STEM Education www. grnjournal.us 271 AMERICAN Journal of Language, Literacy and Learning in STEM Education www. grnjournal.us 4. The Transformer Era (2018 – Present): Pre-Trained Language Models A major milestone occurred with the introduction of the Transformer architecture (Vaswani et al., 2017), which revolutionized NLP by enabling models to process sequences in parallel using attention mechanisms. Pre-trained models such as BERT (2018), GPT series (2018 – 2024), and T5 have since dominated the field[15]. ➢ Linguistic integration: Contextualized embeddings capture syntax, semantics, and pragmatics. ➢ Technological advancements: Transfer learning, self-supervised pre-training, fine-tuning. ➢ Applications: Chatbots, summarization, translation, question answering, dialogue systems. ➢ Limitations: Ethical concerns (bias, misinformation), high computational cost, hallucinations. These models are capable of handling complex linguistic phenomena such as anaphora resolution, discourse coherence, and cross-lingual tasks. Yet, they raise profound ethical and epistemological questions about meaning, intent, and truth in machine-generated language[16]. 272 AMERICAN Journal of Language, Literacy and Learning in STEM Education www. grnjournal.us 5. Current Trends and Beyond: Towards Cognitive and Ethical NLP In the current phase, researchers are exploring how to incorporate cognitive and psycholinguistic theories into AI models to enhance interpretability, robustness, and human-likeness. At the same time, ethical NLP is emerging as a crucial subfield, focusing on fairness, inclusivity, and accountability[17]. ➢ Emerging directions: Explainable NLP, multimodal language models, zero-shot learning. ➢ Linguistic opportunities: Revitalization of endangered languages, cultural adaptation, discourse modeling. ➢ Ethical priorities: Mitigating bias, ensuring transparency, aligning models with human values. This forward-looking stage underscores the need for cross-disciplinary collaboration between linguists, cognitive scientists, and AI engineers to ensure that future NLP systems are not only powerful but also responsible, inclusive, and linguistically aware[18]. 273 AMERICAN Journal of Language, Literacy and Learning in STEM Education www. grnjournal.us 2. Challenges in Natural Language Processing 2.1 Linguistic Ambiguity Ambiguity is one of the most persistent challenges in NLP because language is inherently flexible. Words often carry multiple meanings (polysemy) or share identical forms with unrelated meanings (homonymy). For instance, “bank” may refer to money storage or a river’s edge, depending on context. Large language models have improved in semantic disambiguation, yet they still struggle with figurative language, sarcasm, and pragmatic cues. These limitations highlight the difficulty of replicating the human brain’s ability to seamlessly interpret context[19]. 2.2 Low-Resource Languages The majority of NLP progress benefits high-resource languages, leaving thousands of world languages underrepresented. Low-resource languages suffer from scarce digital corpora, limited linguistic documentation, and a lack of annotated datasets. Consequently, speakers of these languages are excluded from technological advancements in machine translation, voice assistants, and text processing. This digital gap not only widens global inequality but also threatens cultural diversity. Addressing this challenge requires novel transfer learning, unsupervised approaches, and collaborations with native communities to build inclusive AI systems[20]. 2.3 Cross-Cultural Semantics Language is deeply intertwined with culture, making semantic transfer across languages particularly complex. Idioms, metaphors, and culturally bound expressions often lose meaning in translation because AI systems focus on literal forms. For example, translating “spill the beans” word-for-word fails to convey its idiomatic meaning of revealing a secret. Cross-cultural semantics requires NLP 274 AMERICAN Journal of Language, Literacy and Learning in STEM Education www. grnjournal.us models to not only process words but also to infer shared cultural knowledge, pragmatic intentions, and symbolic associations. Achieving this goal is still one of the greatest hurdles for modern NLP research[21]. 2.4 Bias and Ethics Bias in AI reflects the prejudices embedded in its training data. Language models may reinforce gender stereotypes by associating certain professions with one gender, or they may reproduce racial and cultural misrepresentations. Beyond bias, ethical concerns include misinformation, surveillance misuse, and environmental costs of training massive models. Without safeguards, these systems risk perpetuating inequality and eroding public trust. Thus, bias mitigation, fairness auditing, and ethical guidelines are essential for creating socially responsible NLP technologies that promote inclusivity and protect human rights[22]. 2.5 Pragmatic Understanding While models excel at syntax and semantics, pragmatics remains a major challenge. Pragmatics deals with how meaning changes in different contexts, involving implicatures, indirect speech acts, and politeness strategies. For instance, the phrase “It’s cold here” may be a request to close a window rather than a mere statement. Capturing such nuance requires integration of linguistic knowledge with broader world knowledge and social awareness. Current AI lacks the ability to consistently infer these hidden intentions, underscoring the gap between computational models and human communication[23]. 3. Opportunities and Applications 3.1 Machine Translation Recent advances in transformer-based models have significantly improved the accuracy and fluency of machine translation. Unlike early systems, which relied on word-for-word substitutions, modern systems account for syntax, semantics, and context, enabling real-time multilingual communication. These technologies bridge gaps in education, international business, diplomacy, and healthcare, allowing smoother collaboration across cultures. The integration of multilingual pre-trained models now supports dozens of languages simultaneously, opening the door to global accessibility and reducing communication barriers worldwide[24]. 3.2 Healthcare and Biomedicine NLP is transforming healthcare by extracting vital insights from unstructured clinical texts such as patient records, diagnostic notes, and radiology reports. It supports early disease detection through linguistic biomarkers, helps identify adverse drug reactions, and streamlines administrative tasks. AI chatbots assist patients by answering medical questions, while advanced models link textual data with genomics and pharmacology to advance personalized medicine. By automating knowledge discovery, NLP accelerates healthcare innovation, reduces clinician workload, and improves patient outcomes, though ethical and privacy concerns remain central challenges[25]. 3.3 Education Education is one of the most promising domains for NLP applications. AI-powered tutoring systems provide individualized feedback, helping students acquire language skills through interactive practice. Tools such as grammar checkers, automated essay scorers, and virtual teaching assistants make language learning more engaging and accessible. By adapting to learners’ proficiency levels, these systems support personalized curricula and continuous assessment. NLP thereby democratizes access to high-quality education and supports second-language acquisition globally, especially for learners in remote or under-resourced environments[26]. 3.4 Sentiment and Discourse Analysis Sentiment and discourse analysis have become essential tools in business, marketing, and politics. Companies analyze customer reviews to understand consumer preferences, while governments monitor social media discourse to gauge public opinion. In political science, NLP uncovers 275 AMERICAN Journal of Language, Literacy and Learning in STEM Education www. grnjournal.us ideological framing, rhetorical strategies, and polarization trends. However, these applications also raise ethical questions when used for surveillance or manipulation. When responsibly applied, sentiment and discourse analysis provide powerful insights into human behavior, enabling better decision-making and more responsive communication strategies[27]. 3.5 Language Preservation AI offers innovative ways to document and revitalize endangered languages. NLP can process oral recordings, generate dictionaries, and create digital learning resources for communities at risk of language extinction. By automating transcription and translation, AI helps preserve cultural heritage and maintain linguistic diversity. Projects leveraging NLP for minority and indigenous languages are vital not only for academic research but also for empowering communities to reclaim their identities. In this sense, AI becomes a cultural ally, safeguarding traditions that might otherwise be lost[28]. 4. Future Directions 4.1 Integration of Cognitive and Psycholinguistic Insights Future NLP models will need to draw from cognitive science and psycholinguistics to emulate human-like understanding. This includes incorporating theories of attention, memory, and discourse processing into machine learning frameworks. By modeling how humans acquire and process language, AI systems could achieve more natural dialogue and better contextual reasoning. Such integration would bridge the gap between computational efficiency and cognitive plausibility, producing systems that not only generate text but also interpret meaning in ways aligned with human cognition[29]. 4.2 Development of Explainable NLP Systems The complexity of deep learning models makes them difficult to interpret, raising concerns in sensitive areas like healthcare, law, and finance. Explainable NLP seeks to address this by designing systems that can justify their predictions and highlight decision pathways. This transparency enhances trust, allows researchers to identify biases, and helps users understand limitations. Explainable systems are essential for accountability, particularly when AI influences high-stakes decisions, ensuring that technological advancement aligns with societal and regulatory expectations[30]. 4.3 Cross-Disciplinary Collaboration The future of NLP will depend on close collaboration between linguists, computer scientists, ethicists, and psychologists. Linguists provide theoretical grounding, while computer scientists contribute computational techniques. Ethicists ensure fairness, and psychologists add insights into human cognition. By integrating these perspectives, NLP can evolve beyond purely technical achievements toward socially and linguistically responsible systems. Cross-disciplinary collaboration fosters innovation that is both technologically robust and culturally sensitive, ensuring that AI benefits a broader segment of society[31]. 4.4 Focus on Low-Resource Language Technologies A critical future direction is the development of NLP tools for low-resource languages. Advances in transfer learning, unsupervised learning, and multilingual pre-training enable AI models to generalize across languages with minimal data. Expanding research in this area ensures linguistic inclusivity and reduces the digital divide[32]. By empowering underrepresented communities with language technologies, AI can preserve cultural diversity while democratizing access to knowledge. This inclusive approach not only advances science but also strengthens global equity in digital communication[33]. Conclusion The integration of Artificial Intelligence (AI) and Linguistics has significantly reshaped the landscape of Natural Language Processing (NLP). This dynamic relationship has evolved through several technological eras—from symbolic and rule-based approaches to data-driven statistical methods, and from deep learning to today's transformer-based architectures. Each stage has brought 276 AMERICAN Journal of Language, Literacy and Learning in STEM Education www. grnjournal.us both breakthroughs and new challenges, particularly in addressing the complexity of human language, cultural context, and ethical concerns. While modern NLP models such as GPT, BERT, and T5 demonstrate impressive capabilities in language understanding and generation, they still struggle with linguistic ambiguity, low-resource languages, and pragmatic interpretation. Additionally, concerns about bias, transparency, and the ethical deployment of language technologies remain critical barriers to responsible innovation. However, the opportunities are vast. NLP is now central to multilingual communication, healthcare innovation, education technologies, sentiment analysis, and even language preservation. By leveraging interdisciplinary collaboration, integrating cognitive and psycholinguistic insights, and prioritizing inclusivity, the field can move toward building AI systems that not only perform well but also respect and reflect the rich diversity of human language and culture. In conclusion, the future of NLP lies in its ability to balance computational efficiency with linguistic depth, cultural sensitivity, and ethical responsibility—ensuring AI serves all of humanity in a fair and meaningful way. References 1. “From Algorithms to Intelligence: The Historical Perspective of AI in Software Development: Computer Science & IT Book Chapter | IGI Global Scientific Publishing.” Accessed: Aug. 20, 2025. [Online]. Available: https://www.igi-global.com/chapter/from-algorithms-to- intelligence/383147 2. S. Chakraborty, P. Das, S. Mahmud Dipto, Md. A. Pramanik, and J. Noor, “An Analytical Review of Preprocessing Techniques in Bengali Natural Language Processing,” IEEE Access, vol. 13, pp. 112428–112445, 2025, doi: 10.1109/ACCESS.2025.3574234. 3. “Foundation and large language models: fundamentals, challenges, opportunities, and social impacts | Cluster Computing.” Accessed: Aug. 20, 2025. [Online]. Available: https://link.springer.com/article/10.1007/s10586-023-04203-7 4. T. Zhong et al., “Opportunities and Challenges of Large Language Models for Low-Resource Languages in Humanities Research,” Dec. 09, 2024, arXiv: arXiv:2412.04497. doi: 10.48550/arXiv.2412.04497. 5. K. S. Jones, “Natural Language Processing: A Historical Review,” in Current Issues in Computational Linguistics: In Honour of Don Walker, A. Zampolli, N. Calzolari, and M. Palmer, Eds., Dordrecht: Springer Netherlands, 1994, pp. 3–16. doi: 10.1007/978-0-585-35958-8_1. 6. “Full article: Artificial intelligence: reflecting on the past and looking towards the next paradigm shift.” Accessed: Aug. 20, 2025. [Online]. Available: https://www.tandfonline.com/doi/full/10.1080/0952813X.2024.2323042 7. B. C. Pierce, “Linguistic foundations for bidirectional transformations: invited tutorial,” in Proceedings of the 31st ACM SIGMOD-SIGACT-SIGAI symposium on Principles of Database Systems, in PODS ’12. New York, NY, USA: Association for Computing Machinery, May 2012, pp. 61–64. doi: 10.1145/2213556.2213568. 8. H. Tran-Dang, J.-W. Kim, J.-M. Lee, and D.-S. Kim, “Shaping the Future of Logistics: Data- driven Technology Approaches and Strategic Management,” IETE Tech. Rev., vol. 42, no. 1, pp. 44–79, Jan. 2025, doi: 10.1080/02564602.2024.2445513. 9. “Full article: A Systematic Review of the Limitations and Associated Opportunities of ChatGPT.” Accessed: Aug. 20, 2025. [Online]. Available: https://www.tandfonline.com/doi/full/10.1080/10447318.2024.2344142 10. A. K. Oyebamiji, S. A. Akintelu, S. O. Afolabi, O. Ebenezer, E. T. Akintayo, and C. O. Akintayo, “A Comprehensive Review on Mycosynthesis of Nanoparticles, Characteristics, Applications, and Limitations,” Plasmonics, Jan. 2025, doi: 10.1007/s11468-024-02755-x. 277 AMERICAN Journal of Language, Literacy and Learning in STEM Education www. grnjournal.us 11. H. Y. O. Sum, “A Chronological Narrative Review of AI Evolution in Dentistry,” Pak. J. Life Soc. Sci. PJLSS, vol. 23, no. 1, 2025, doi: 10.57239/PJLSS-2025-23.1.00597. 12. D. Zha et al., “Data-centric Artificial Intelligence: A Survey,” ACM Comput Surv, vol. 57, no. 5, p. 129:1-129:42, Jan. 2025, doi: 10.1145/3711118. 13. J. Zhang et al., “New trend on chemical structure representation learning in toxicology: In reviews of machine learning model methodology,” Crit. Rev. Environ. Sci. Technol., vol. 55, no. 13, pp. 951–976, July 2025, doi: 10.1080/10643389.2025.2469868. 14. L. Zhao, “Advances in functional magnetic resonance imaging-based brain function mapping: a deep learning perspective,” Psychoradiology, vol. 5, p. kkaf007, Mar. 2025, doi: 10.1093/psyrad/kkaf007. 15. H. Gautam, A. Gaur, and D. K. Yadav, “A Survey on the Impact of Pre-Trained Language Models in Sentiment Classification Task,” Int. J. Data Sci. Anal., May 2025, doi: 10.1007/s41060-025- 00805-z. 16. Md. A. Haque and H. R. Siddique, “Generative artificial intelligence and large language models in smart healthcare applications: Current status and future perspectives,” Comput. Biol. Chem., vol. 120, p. 108611, Feb. 2026, doi: 10.1016/j.compbiolchem.2025.108611. 17. L. Orynbay, G. Bekmanova, B. Yergesh, A. Omarbekova, A. Sairanbekova, and A. Sharipbay, “The role of cognitive computing in NLP,” Front. Comput. Sci., vol. 6, Jan. 2025, doi: 10.3389/fcomp.2024.1486581. 18. C. Fenerci, Z. Cheng, D. R. Addis, B. Bellana, and S. Sheldon, “Studying memory narratives with natural language processing,” Trends Cogn. Sci., vol. 29, no. 6, pp. 516–525, June 2025, doi: 10.1016/j.tics.2025.02.003. 19. F. Gaim and J. C. Park, “Natural Language Processing for Tigrinya: Current State and Future Directions,” July 25, 2025, arXiv: arXiv:2507.17974. doi: 10.48550/arXiv.2507.17974. 20. T. O. Tafa et al., “Machine Translation Performance for Low-Resource Languages: A Systematic Literature Review,” IEEE Access, vol. 13, pp. 72486–72505, 2025, doi: 10.1109/ACCESS.2025.3562918. 21. L. N. Toan, “Contrastive Semantics in Cross-Linguistic Analysis: A Comprehensive Review and Synthesis of Five Crucial Models,” REiLA J. Res. Innov. Lang., vol. 7, no. 1, pp. 111–123, May 2025, doi: 10.31849/dew37f39. 22. M. G. Hanna et al., “Ethical and Bias Considerations in Artificial Intelligence/Machine Learning,” Mod. Pathol., vol. 38, no. 3, p. 100686, Mar. 2025, doi: 10.1016/j.modpat.2024.100686. 23. M. Taumoepeau, “Pragmatics and Theory of Mind across cultures,” Philos. Trans. R. Soc. B Biol. Sci., vol. 380, no. 1932, p. 20230500, Aug. 2025, doi: 10.1098/rstb.2023.0500. 24. “Frontiers | Exploring ChatGPT’s potential for augmenting post-editing in machine translation across multiple domains: challenges and opportunities.” Accessed: Aug. 20, 2025. [Online]. Available: https://www.frontiersin.org/journals/artificial- intelligence/articles/10.3389/frai.2025.1526293/full 25. “Open challenges and opportunities in federated foundation models towards biomedical healthcare | BioData Mining.” Accessed: Aug. 20, 2025. [Online]. Available: https://link.springer.com/article/10.1186/s13040-024-00414-9 26. “A systematic review of AI, VR, and LLM applications in special education: Opportunities, challenges, and future directions | Education and Information Technologies.” Accessed: Aug. 20, 2025. [Online]. Available: https://link.springer.com/article/10.1007/s10639-025-13550-4 27. S. M. Maci and P. Anesa, “The impact of AI on discourse analysis: Challenges and opportunities. |EBSCOhost.” Accessed: Aug. 20, 2025. [Online]. Available: 278 AMERICAN Journal of Language, Literacy and Learning in STEM Education www. grnjournal.us https://openurl.ebsco.com/contentitem/doi:10.5281%2Fzenodo.15250758?sid=ebsco:plink:craw ler&id=ebsco:doi:10.5281%2Fzenodo.15250758 28. V. Koc, “Generative AI and Large Language Models in Language Preservation: Opportunities and Challenges,” May 19, 2025, arXiv: arXiv:2501.11496. doi: 10.48550/arXiv.2501.11496. 29. “Gender Stereotypes and Language Processing: Cognitive and Social Insights from a Decade of Research (2012–2023) - Beroíza‐Valenzuela - 2025 - European Journal of Education - Wiley Online Library.” Accessed: Aug. 20, 2025. [Online]. Available: https://onlinelibrary.wiley.com/doi/abs/10.1111/ejed.70063 30. “Explainable AI for Medical Data: Current Methods, Limitations, and Future Directions | ACM Computing Surveys.” Accessed: Aug. 20, 2025. [Online]. Available: https://dl.acm.org/doi/abs/10.1145/3637487 31. Q. Chen et al., “Hypertension-gut microbiota research trends: a bibliometric and visualization analysis (2000–2025),” Front. Microbiol., vol. 16, Aug. 2025, doi: 10.3389/fmicb.2025.1543258. 32. “Grammatical error correction for low-resource languages: a review of challenges, strategies, computational and future directions [PeerJ].” Accessed: Aug. 20, 2025. [Online]. Available: https://peerj.com/articles/cs-3044/ 33. J. McGiff and N. S. Nikolov, “Overcoming Data Scarcity in Generative Language Modelling for Low-Resource Languages: A Systematic Review,” July 08, 2025, arXiv: arXiv:2505.04531. doi: 10.48550/arXiv.2505.04531.