Developing a Corpus-Based Immigration Keyword and Key Cluster List for ESP Settings
Main Article Content
Abstract
Vocabulary knowledge plays a central role in English for Specific Purposes, yet immigration remains a professional domain that has received little attention in specialised word list research. This study developed a corpus-based Immigration Keyword and Key Cluster List (IKKCL) to support ESP learners working with immigration forms. Data were collected from 1,670 official immigration forms downloaded from government websites in different countries. Of these, 1,500 forms formed the primary Global Immigration Forms Corpus (2,939,837 tokens), and the remaining 170 formed a secondary validation corpus (179,897 tokens). Candidate items were extracted through keyword analysis, applying Log-likelihood, Log Ratio, and Simple Maths together against the British National Corpus, followed by lexical profiling and, for two-word clusters, compounding identification and Mutual Information filtering. Surviving candidates were then validated against five institutional immigration glossaries and, where needed, examined through expert concordance analysis before being classified into technical, semi-technical, and supportive profiles. The final IKKCL contained 589 items, including 21 technical, 196 semi-technical, and 372 supportive words and clusters. The list covered an average of 19.06% of running words across both corpora, with near-identical coverage on the unseen validation corpus, indicating stable performance within the same sampling frame. Combined with the General Service List and the Academic Word List, total coverage reached approximately 90%, falling short of the threshold typically associated with independent reading comprehension. The IKKCL can be used as a resource for teaching and learning in courses such as English for Immigration and Tourism, as well as related curricula, through approaches such as data-driven learning.
Article Details

This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.
Authors who publish with this journal agree to the following terms: Authors retain copyright and grant the journal right of first publication with the work simultaneously licensed under a Creative Commons Attribution License that allows others to share the work with an acknowledgement of the work's authorship and initial publication in this journal. Authors are able to enter into separate, additional contractual arrangements for the non-exclusive distribution of the journal's published version of the work (e.g., post it to an institutional repository or publish it in a book), with an acknowledgement of its initial publication in this journal. Authors are permitted and encouraged to post their work online (e.g., in institutional repositories or on their website) prior to and during the submission process, as it can lead to productive exchanges, as well as earlier and greater citation of published work (See The Effect of Open Access).References
Anthony, L. (2026a). AntWordProfiler (Version 2.2.1) [Computer software]. Waseda University. https://www.laurenceanthony.net/software/AntWordProfiler
Anthony, L. (2026b). AntConc (Version 4.2.0) [Computer software]. Waseda University. https://www.laurenceanthony.net/software/AntConc
Biber, D. (1993). Representativeness in corpus design. Literary and Linguistic Computing, 8(4), 243–257. https://doi.org/10.1093/llc/8.4.243
Boulton, A. (2009). Testing the limits of data-driven learning: language proficiency and training. ReCALL, 21(1), 37–54. https://doi.org/10.1017/S0958344009000068
Brezina, V., & Platt, W. (2026). #LancsBox X [Software]. Lancaster University. http://lancsbox.lancs.ac.uk
Chung, T. M., & Nation, P. (2003). Technical vocabulary in specialised texts. Reading in a Foreign Language, 15(2), 103–116. https://doi.org/10.64152/10125/66770
Cobb, T. (2025). Wordlists and data-driven learning. In D. Tafazoli (Ed.), The Palgrave Encyclopedia of Computer-Assisted Language Learning (pp. 1–8). Palgrave Macmillan. https://doi.org/10.1007/978-3-031-51447-0_155-1
Coxhead, A. (2000). A new academic word list. TESOL Quarterly, 34(2), 213–238. https://doi.org/10.2307/3587951
Coxhead, A., McLaughlin, E., & Reid, A. (2019). The development and application of a specialised word list: The case of fabrication. Journal of Vocational Education & Training, 71(2), 175–200. https://doi.org/10.1080/13636820.2018.1471094
Dang, T. N. Y. (2020). Corpus-based word lists in second language vocabulary research, learning, and teaching. In S. Webb (Ed.), The Routledge Handbook of Vocabulary Studies (pp. 288–304). Routledge. https://doi.org/10.4324/9780429291586
Dang, T. N. Y., & Webb, S. (2025). Applications of word lists in second language learning and teaching. Language Teaching, 58(3), 291–311. https://doi.org/10.1017/S0261444825000059
Daskalovska, N. (2015). Corpus-based versus traditional learning of collocations. Computer Assisted Language Learning, 28(2), 130–144. https://doi.org/10.1080/09588221.2013.803982
Drašler, A., & Kavalir, M. (2026). Compiling a custom corpus and word list for ESAP: The case of English for Geographers. English for Specific Purposes, 81, 92–102. https://doi.org/10.1016/j.esp.2025.09.004
Gablasova, D., Brezina, V., & McEnery, T. (2017). Collocations in corpus-based language learning research: Identifying, comparing, and interpreting the evidence. Language Learning, 67(S1), 155–179. https://doi.org/10.1111/lang.12225
Gabrielatos, C., & Marchi, A. (2011). Keyness: Matching metrics to definitions. University of Portsmouth.
Gardner, D., & Davies, M. (2014). A new academic vocabulary list. Applied Linguistics, 35(3), 305–327. https://doi.org/10.1093/applin/amt015
Gries, S. Th. (2008). Dispersions and adjusted frequencies in corpora. International Journal of Corpus Linguistics, 13(4), 403–437. https://doi.org/10.1075/ijcl.13.4.02gri
Hanks, E., Hashimoto, B., & Egbert, J. (2024). The contracts word list: Integral vocabulary for reading and writing English contracts. English for Specific Purposes, 75, 37–48. https://doi.org/10.1016/j.esp.2024.03.002
Hardie, A. (2014). Log Ratio – an informal introduction. Centre for Corpus Approaches to Social Science (CASS), Lancaster University. https://cass.lancs.ac.uk/log-ratio-an-informal-introduction/
Hyland, K. (2019). English for specific purposes: Some influences and impacts. In X. Gao (Ed.), Second handbook of English language teaching (pp. 337–353). Springer. https://doi.org/10.1007/978-3-030-02899-2_19
Hyland, K., & Tse, P. (2007). Is there an “academic vocabulary”? TESOL Quarterly, 41(2), 235–253. https://doi.org/10.1002/j.1545-7249.2007.tb00058.x
IOM. (2019). Glossary on migration. International Organization for Migration.
Kilgarriff, A. (2009). Simple maths for keywords. In M. Mahlberg, V. González-Díaz, & C. Smith (Eds.), Proceedings of the Corpus Linguistics Conference (CL2009) (Article 171). University of Liverpool.
Landis, J. R., & Koch, G. G. (1977). The measurement of observer agreement for categorical data. Biometrics, 33(1), 159–174. https://doi.org/10.2307/2529310
Laosrirattanachai, P., & Laosrirattanachai, P. (2021). Applying lexical profiling to construct technical word lists for Thai tourist guides. PASAA, 62(1), 61–91. https://doi.org/10.58837/CHULA.PASAA.62.1.3
Laosrirattanachai, P., & Laosrirattanachai, P. (2025). Tracing tourism business research trends in Scopus-indexed journals using corpus-based and judgement-based approaches. Humanities, Arts and Social Sciences Studies, 25(1), 32–53. https://doi.org/10.69598/hasss.25.1.268122
Laosrirattanachai, P., Tangnaitammakun, Y., Tinhuatoey, K., Sutthongkong, A., & Laosrirattanachai, P. (2026a). Developing a global stock market terminology list for ESP learners in business English: A corpus-based approach. The New English Teacher, 20(2), 125–140. https://doi.org/10.59865/t.v20i2.9688
Laosrirattanachai, P., Rittikote, N., & Laosrirattanachai, P. (2026b). Developing a nature tourism word list for L2 learners in the tourism sector. Australian Review of Applied Linguistics, Advance online publication. https://doi.org/10.1075/aral.25109.lao
Lee, H., Warschauer, M., & Lee, J. H. (2017). The effects of concordance-based electronic glosses on L2 vocabulary learning. Language Learning & Technology, 21(2), 32–51. https://doi.org/10.64152/10125/44610
Lessard-Clouston, M. (2013). Word lists for vocabulary learning and teaching. The CATESOL Journal, 24(1), 287–304. https://doi.org/10.5070/B5.36167
Lusta, A., Demirel, Ö., & Mohammadzadeh, B. (2025). Breaking the stereotype of data driven learning: Students’ experiences of mobile data-driven learning in a developing country. SAGE Open, 15(4), 1–20. https://doi.org/10.1177/21582440251386904
Martinez, R., & Schmitt, N. (2012). A phrasal expressions list. Applied Linguistics, 33(3), 299–320. https://doi.org/10.1093/applin/ams010
Meebangsai, D., Pongtin, P., Kitipoontanakorn, P., & Laosrirattanachai, P. (2023). Investigating proficiency of academic English in student writing: A comparative case study on vocabulary utilization in student research article writing vis–à–vis national and international research. PASAA, 67, 66–100. https://doi.org/10.58837/CHULA.PASAA.67.1.3
Nation, I. S. P. (2016). Making and using word lists for language learning and testing. John Benjamins.
Nation, I. S. P. (2022). Learning vocabulary in another language (3rd ed.). Cambridge University Press.
Otto, P. (2021). Choosing specialized vocabulary to teach with data-driven learning: An example from civil engineering. English for Specific Purposes, 61, 32–46. https://doi.org/10.1016/j.esp.2020.08.003
Phoocharoensil, S. (2026). Constructing the LERWL: A corpus-based approach to identifying technical vocabulary in language education research. rEFLections, 33(2), 739–759. https://doi.org/10.61508/refl.v33i2.290763
Rayson, P., & Garside, R. (2000). Comparing corpora using frequency profiling. In Proceedings of the Workshop on Comparing Corpora (pp. 1–6). Association for Computational Linguistics. https://doi.org/10.3115/1117729.1117730
Saeedakhtar, A., Bagerin, M., & Abdi, R. (2020). The effect of hands-on and hands-off data-driven learning on low-intermediate learners’ verb-preposition collocations. System, 91, Article 102268. https://doi.org/10.1016/j.system.2020.102268
Schmitt, N. (2000). Vocabulary in language teaching. Cambridge University Press.
Schmitt, N., Jiang, X., & Grabe, W. (2011). The percentage of words known in a text and reading comprehension. The Modern Language Journal, 95(1), 26–43. https://doi.org/10.1111/j.1540-4781.2011.01146.x
Schmitt, N., & Schmitt, D. (2014). A reassessment of frequency and vocabulary size in L2 vocabulary teaching. Language Teaching, 47(4), 484–503. https://doi.org/10.1017/S0261444812000018
Watson Todd, R. (2017). An opaque engineering word list: Which words should a teacher focus on? English for Specific Purposes, 45, 31–39. https://doi.org/10.1016/j.esp.2016.08.003
West, M. (1953). A general service list of English words. Longman.
Wilkins, D. (1972). Linguistics in language teaching. Edward Arnold.
Yalcinkaya, F., Muluk, N., & Sahin, S. (2009). Effects of listening ability on speaking, writing and reading skills of children who were suspected of auditory processing difficulty. International Journal of Pediatric Otorhinolaryngology, 73(8), 1137–1142. https://doi.org/10.1016/j.ijporl.2009.04.022