The rapid expansion of digital news ecosystems has transformed the study of conflict discourse by creating unprecedented volumes of textual data that reflect political positions, public attitudes, and strategic narratives. This study develops a digital corpus architecture for investigating war-related terminology and narrative structures in Ukrainian news media, focusing on the extraction and analysis of opinion and action narratives. The research integrates corpus linguistics principles, web-based data collection approaches, and computational text analysis techniques to construct a systematic framework for identifying lexical patterns associated with conflict representation. Drawing upon corpus-driven methodologies, the study examines how war terminology functions as a linguistic mechanism for framing military events, political decisions, humanitarian concerns, and social responses. The proposed architecture combines automated data acquisition, corpus annotation, concordance analysis, collocation measurement, and multidimensional interpretation to analyze large-scale news discourse. The theoretical foundation is informed by corpus-based discourse analysis, where context and pragmatic functions determine the meaning of linguistic choices (Adolphs, 2008). The findings demonstrate that a structured digital corpus enables researchers to identify recurring war lexicons, ideological positioning, and narrative transformations across news sources. The research contributes a methodological framework for digital humanities, media linguistics, and conflict discourse studies by demonstrating how computational corpus approaches can reveal patterns of meaning construction within contemporary war reporting.
Building A Digital Corpus Architecture for War Terminology Research: Extracting Opinion and Action Narratives from Ukrainian News Media
Keywords:
Abstract
Downloads
References
Adolphs, S. (2008). Corpus and context: Investigating pragmatic functions in spoken discourse. John Benjamins Publishing.
Andrés, H., & Velasco, M. (2025, May 29). Guerra Ucrania –Rusia, últimas noticias. El Mundo.
Anthony, L. (2005). AntConc: Design and development of a freeware corpus analysis toolkit for the technical writing classroom. Proceedings of the IEEE International Professional Communication Conference, 729–737.
AntConc. (2021). SAGE Publications, Ltd.
Automated data collection: Web scraping example, part 1. (2017).
Baker, P. (2006). Using corpora in discourse analysis. Continuum.
Bassets, M. (2025, May 28). Alemania amplía la ayuda a Ucrania para fabricar armas que alcancen territorio ruso. Ediciones EL PAÍS S.L.
Biber, D. (2014). Opening. Multi-Dimensional analysis. In Studies in Corpus Linguistics (pp. xxix–xxxviii). John Benjamins Publishing Company.
Carter, R. (2004). Language and creativity: The art of common talk. Routledge.
Collins, P., & Yao, X. (2016). Douglas Biber and Randi Reppen (eds.), The Cambridge handbook of English corpus linguistics. Cambridge: Cambridge University Press, 2015. Pp. 639. ISBN 978-1-107-03738-0 (hardback). English Language and Linguistics, 21(1), 161–168.
Cook, G. (2004). Genetically modified language: The discourse of arguments for GM crops and food. Routledge.
Davies, M. (2004). Student use of large, annotated corpora to analyze syntactic variation. In Studies in Corpus Linguistics (pp. 259–269). John Benjamins Publishing Company.
Davies, M., & Parodi, G. (2022). Constitución de corpus crecientes del español. In Lingüística de corpus en español (pp. 13–32). Routledge.
Editorial, T. (2025, May 28). Zelenskyy says what the Ukrainian military needs most. ТСН.
Evert, S. (2005). The statistics of word cooccurrences: Word pairs and collocations (Doctoral dissertation, University of Stuttgart).
Fendel, V. B. (2024). The i.sicily Sketch engine corpus. Journal of Open Humanities Data, 10.
Gupta, S. (2024). Getting started with Web Scraping. In Web Scraping with Python. Apress.
Hanks, P. (2016, April 5). Create and search a text corpus. Sketch Engine.
Johnson, M. (2024). Web-scraping. Litres.
Kilgarriff, A., Baisa, V., Bušta, J., Jakubíček, M., Kovář, V., Michelfeit, J., Rychlý, P., & Suchomel, V. (2014). The Sketch Engine: Ten years on. Lexicography, 1(1), 7–36.
Kilgarriff, A., Rychly, P., Smrž, P., & Tugwell, D. (2008). The Sketch Engine. In Practical Lexicography (pp. 297–306). Oxford University PressOxford.
Kinkartz, S. (2025b, May 28). Migration: Deutschland will es Flüchtlingen schwerer machen. Deutsche Welle.
LancsBox X. (2025). Retrieved June 3, 2025, from https://lancsbox.lancs.ac.uk/
Laurence Anthony’s Website. (2025). Retrieved June 3, 2025, from https://www.laurenceanthony.net/software/antconc/
Leech, G. (1997). Introducing corpus annotation. In R. Garside, G. Leech, & T. McEnery (Eds.), Corpus annotation: Linguistic information from computer text corpora (pp. 1–18). Longman.