#IDETECT
In January 2017, IDMC and the UN launched the Internal Displacement Event Tagging, Extraction
and Clustering Tool (#IDETECT) challenge on the UN’s data science crowdsourcing platform,
Unite Ideas. It has brought together teams representing dozens of data scientists from around
the world to develop a new tool that we will use to monitor displacement associated with
disasters, conflict, violence and development projects.
#IDETECT will expand and diversify the sources we use for monitoring significantly, helping to
address – though not eliminate – some of the factors that impede our painting a comprehensive global picture. The tool will cast a wide net so we can obtain information about, analyse
and shed light on far more displacement situations than we currently do (see figure 3.1). That
said, #IDETECT’s scope will still be limited to events reported in the media or by partners in
the field. To overcome this reporting bias, we have also begun exploring further approaches
to detect displacement using other types of data and means of analysis.
The tool will make our monitoring more efficient and comprehensive, and it will also provide
the humanitarian community with an easy way to extract and analyse facts from any type of
documents, be they news, field reports, social media or other sources.
How it works: Filtering and tagging
The first step is to mine huge datasets of news, such as the GDELT Project, the European Media
Monitor and social media platforms, and extract records that relate to displacement. The next
is to tag the events as being related to conflict, violence, disasters or other cause or trigger.
Natural language processing
The tool will use natural language processing (NLP) to extract certain facts from the source
material including, but not limited to:
|| The publication date of the document
|| The place where the displacement reportedly occurred
|| The number of people displaced
Data visualisation, human validation and machine learning
It will then visualise the data for us and our partners to review. The results of this human validation process will inform the NLP so that it performs more accurately in the future, a process
known as supervised machine learning (see figure 3.11).
84
GRID
2017
Select target paragraph3
Connect to a paragraph
Connect to an entity
Disable highlights
Add to table of contents