Elite polarization,
measured from parliamentary speech
An interactive map of how legislators talk about their own and rival parties, built from millions of political mentions extracted across national parliaments. It includes the world map, country trajectories, cross-national comparisons and the underlying data.
How the map is built
This project aggregates parliamentary data worldwide to measure elite polarization by analyzing what politicians say about one another inside parliaments. It employs an agent-AI-driven framework for parliamentary data collection, homologation, and analysis, orchestrated through the OpenClaw framework.
The system collects transcripts from academic archives across many countries. When corpora exist only on parliamentary websites, it uses web-scraping techniques to acquire them. Because many parliaments publish proceedings as PDF scans, the pipeline separates born-digital PDFs, where text is extracted directly, from scanned-image PDFs, where optical character recognition is applied. For video-only records, it extracts official captions or subtitles; otherwise, it tests automated speech recognition on a sample before transcription.
OpenClaw coordinates the workflow across collection, normalization, affiliation recovery, extraction, validation, and escalation. Collected transcripts are normalized into a unified format in which each row is one speaker’s turn, with country, date, speaker name, party, and speech text. Since speaker–party affiliation is often missing, a recovery algorithm uses registry and metadata joins, deterministic date-window rules for known affiliations over time, and fuzzy matching against biographical and parliamentary records. Unresolved cases then fall back to language-model lookup.
The extraction stage uses the Google Gemma 4 26B LLM to identify passages where a speaker evaluates another named politician, distinguishing attacks or praise from neutral and procedural references. Each mention is coded as in-group or out-group using recovered coalition affiliation. The output is a cross-national dataset of elite-hostility relationships, aggregated by party-pair and time period, covering over 120 million speeches across 139 countries.
How the corpus was acquired
139 corpora · 119.6 million speeches (+ 3.9 million broadcast-caption segments), grouped by how each parliament’s record was obtained.
World map
Choropleth of any indicator, pooled across all years (mention-weighted) or for a single year.
Trends over time
One indicator over a year window — line per country, or switch to mention-weighted regional averages.
Country profile
Each party's trajectory within a country — out-party evaluation by default. Countries without a party breakdown show the national aggregate.
Compare indicators
Scatter two indicators against each other for a given year. Bubble size = corpus volume; colour = region.
Heat map
Country × year grid for one indicator. Grey cells have no data.
Data table
The underlying panel. Switch between country-level and party-level rows; filter, search and sort.
Validation against survey polarization
Does speech negativity track what voters feel? The parliamentary out-party measure against Orhan's (2022) survey-based Affective Polarization Index (API), built from CSES like-dislike scores across 52 democracies and 178 elections. Each survey observation is paired with the parliamentary speech of the legislature that preceded that election.
x: out-party negativity in parliamentary speech, the mention-weighted mean of the negated out-party evaluation over the legislature preceding the election (previous election + 1 to election year − 1; five-year fallback when the previous election is more than six years back; cells with at least 50 out-party mentions). y: Orhan's API for that election. Dashed line: least-squares fit. Hover for the country and year.
Countries
Country means, sorted by residual: the survey index minus the value the fitted line predicts from speech negativity. Positive residuals are countries whose voters are more polarized than their parliament's speech suggests.
| Country | Elections | Speech negativity | Survey API | Residual |
|---|
Robustness to the pairing window
The preceding-legislature rule is the paper's specification: post-election speech cannot have shaped pre-election attitudes. Adding the election year itself changes the fit by +0.003 at the election level and -0.025 on country means in this build; the full grid is below.
| Window | n elections | r | n countries | r, country means |
|---|