examining the potential challenges | 10.1163/22131035-14020001 9 unsupervised learning finds patterns or mapping functions where the output variable is not present.39 In statelessness determination, the goal is to determine whether a person is not a national of any state under the operation of its law. For labelled data, variables that guide the outcome would include personal history; information concerning laws and other circumstances of different countries, for example the presence of nationality, and nationality legislation (including amended and repealed laws); the possession of any documents proving nationality (identity documents, birth certificates, passport, certificate of naturalisation, certificate of renunciation); and the possession of any other documents that can be used to investigate if a person is a national of any state (for example, marriage certificate, employment contracts/history, residence permits of the country(ies) of habitual residence, identity and travel documents of parents, spouse and children or other official documents that indicate citizenship or recognise statelessness).40 The output variable can be categorised into a binary: stateless or a national. Each will be readily code-able by a human to create training data.41 All of the above output variables informs the decision-making process of the algorithms. If a system is to analyse documents, then this will involve the use of natural language processing (nlp), which can be both supervised and unsupervised learning for analysis of official documents issued by states (e.g., passport, visas, immigration records, identity documents, birth certificates, government correspondence, case law).42 With the use of these variables, algorithm classifiers (e.g., decision trees, random forests, or neural networks) could provide findings into whether a person is a national of any state or stateless.43 39 40 41 42 43 Ibid, 2–3: ‘Subcategories under unsupervised learning are clustering, which involves grouping observations based on similarities, and association, which involves discovering rules that describe large parts of the data. In unsupervised learning, under clustering, for example, a machine would group customers based on their purchasing behaviour, while under association, the machine defines the rules that describe clustering: customers who purchase eggs also purchase bacon.’ basis of input training data. Handbook on Statelessness, (n 13) paras. 83–84. Lehr and Ohm, (n 38) at 674. Amazon, ‘What is Natural Language Processing (nlp)?’<https://aws.amazon.com /what-is/nlp/> accessed 30 September 2025: nlp is a machine learning technology that gives computers the ability to interpret, manipulate, and comprehend human language. nlp software is used to automatically, analyze the intent or sentiment in the message, process, analyze, and summarize complex texts. It can be used to analyze legal documents like court decisions. ibm, ‘What is a Decision Tree?’ <https://www.ibm.com/think/topics/decision-trees> accessed 30 September 2025. International Human Rights Law Review (2025) 1–31

Select target paragraph3