GDELT
From AltData.wiki, The Alternative Data Encyclopedia
GDELT (the Global Database of Events, Language, and Tone) is an open data project that monitors the world's broadcast, print and web news in over 100 languages and converts it into a free, computable database of global events, people, organizations, locations, themes, emotions and imagery.[1] The project is supported by Google Jigsaw and describes itself as a realtime network diagram and database of global human society for open research.[1]
Created by Kalev Leetaru, formerly associated with Yahoo! and Georgetown University, together with political scientist Philip Schrodt and other collaborators, GDELT maintains historical archives stretching back to January 1, 1979 that update every 15 minutes.[1][2] Although designed primarily as a research platform for conflict forecasting and social science, its event and tone datasets are widely used in the alternative data industry as a free source of geopolitical and news sentiment signals.[2]
What It Does
GDELT continuously ingests news media from nearly every country and applies natural language processing and data mining to extract structured records of what is happening worldwide, who is involved, where, and how the world is feeling about it.[1] The project frames this as translating textual descriptions of world events into codified entries in a global spreadsheet updated around the clock.[1]
Data And Methodology
The GDELT Event Database records over 300 categories of physical activity, from protests to diplomatic exchanges, coded using the Conflict and Mediation Event Observations (CAMEO) framework, with nearly 60 attributes per event including georeferenced location.[1][2] The Global Knowledge Graph compiles every person, organization, company and location mentioned alongside millions of themes and thousands of emotion measures, while the Global Content Analysis Measures (GCAM) suite assesses over 2,200 emotions and themes per article.[1] Since January 2016, the Visual Global Knowledge Graph has processed samples of up to one million news images daily through Google's Vision API, and the Translingual platform machine-translates monitored news in 65 languages in realtime.[1]
Products And Access
All GDELT outputs are free and open.[1] Data can be downloaded as raw CSV files, queried through the GDELT Analysis Service visualization tools, or analyzed at scale in Google BigQuery, where the full quarter-billion-record event database is available as a public dataset, a distribution arrangement announced on the Google Cloud Platform blog in May 2014.[1][2] The project also publishes free derived products such as the Global Conflict Dashboard, the Daily Trend Report and the World Leaders Index, which ranks heads of state by the average tone of their news coverage each day.[1]
Buyers And Use Cases
As a free dataset rather than a commercial vendor, GDELT publishes suggested solution areas including situational awareness, influencer networks, risk assessment and global trends, policy reaction, and humanitarian crisis response.[1] In quantitative finance, academic work has combined GDELT with market data for stock movement prediction experiments, and hedge funds commonly consume its event and tone streams as inputs, though the project itself does not sell data or services.[2]
History
Precursors to GDELT were described by co-creator Philip Schrodt in a January 2011 conference paper on automated near-real-time political event data.[2] The project was created by Kalev Leetaru along with Schrodt and others, launched publicly in 2013 era coverage including a Foreign Policy feature mapping every recorded protest since 1979, and became available as a public BigQuery dataset the following year.[2] Its historical coverage continues to be extended, with collaborations aiming to push codified global event history back toward the year 1800.[1]
Landscape
GDELT sits in the open-data layer of the news analytics landscape, complementing commercial vendors such as RavenPack and Dataminr. It is frequently compared with government-sponsored event datasets such as the Integrated Crisis Early Warning System (ICEWS) and ACLED, and scholars have critiqued its media-composition biases, which quantitative users factor into signal design.[2]