examining the potential challenges | 10.1163/22131035-14020001
9
unsupervised learning finds patterns or mapping functions where the output
variable is not present.39
In statelessness determination, the goal is to determine whether a person
is not a national of any state under the operation of its law. For labelled data,
variables that guide the outcome would include personal history; information
concerning laws and other circumstances of different countries, for example
the presence of nationality, and nationality legislation (including amended
and repealed laws); the possession of any documents proving nationality
(identity documents, birth certificates, passport, certificate of naturalisation,
certificate of renunciation); and the possession of any other documents that
can be used to investigate if a person is a national of any state (for example,
marriage certificate, employment contracts/history, residence permits of the
country(ies) of habitual residence, identity and travel documents of parents,
spouse and children or other official documents that indicate citizenship or
recognise statelessness).40
The output variable can be categorised into a binary: stateless or a national.
Each will be readily code-able by a human to create training data.41 All of the
above output variables informs the decision-making process of the algorithms.
If a system is to analyse documents, then this will involve the use of natural
language processing (nlp), which can be both supervised and unsupervised
learning for analysis of official documents issued by states (e.g., passport,
visas, immigration records, identity documents, birth certificates, government
correspondence, case law).42 With the use of these variables, algorithm
classifiers (e.g., decision trees, random forests, or neural networks) could
provide findings into whether a person is a national of any state or stateless.43
39
40
41
42
43
Ibid, 2–3: ‘Subcategories under unsupervised learning are clustering, which involves
grouping observations based on similarities, and association, which involves discovering
rules that describe large parts of the data. In unsupervised learning, under clustering, for
example, a machine would group customers based on their purchasing behaviour, while
under association, the machine defines the rules that describe clustering: customers who
purchase eggs also purchase bacon.’ basis of input training data.
Handbook on Statelessness, (n 13) paras. 83–84.
Lehr and Ohm, (n 38) at 674.
Amazon, ‘What is Natural Language Processing (nlp)?’<https://aws.amazon.com
/what-is/nlp/> accessed 30 September 2025: nlp is a machine learning technology that
gives computers the ability to interpret, manipulate, and comprehend human language.
nlp software is used to automatically, analyze the intent or sentiment in the message,
process, analyze, and summarize complex texts. It can be used to analyze legal documents
like court decisions.
ibm, ‘What is a Decision Tree?’ <https://www.ibm.com/think/topics/decision-trees>
accessed 30 September 2025.
International Human Rights Law Review (2025) 1–31