This repository contains a modular simulation pipeline for studying electoral systems, particularly the single transferable vote (STV) and plurality voting.
The pipeline integrates:
- district plan generation (GerryChain)
- synthetic ballot generation
- election simulation (VoteKit)
- result analysis and visualization
It enables researchers and practitioners to compare how district magnitude, voter behavior, and demographic structure affect representation outcomes. The goal is to make ranked-choice voting simulations easier to run without requiring extensive custom code.
This repository was developed using the UV build system. This build system is generally available through chocolatey on Windows (choco install uv), homebrew on MacOS brew install uv, and through direct installation on Linux (e.g. apt install uv). You can also install directly from source using the instructions at UV's homepage or you can install the system into a conda environment (conda install conda-forge::uv).
After installing UV, you can build the necessary virtual environment for this repository by invoking the command
uv sync
from the terminal while in the base directory.
There are four ways in, all run from the repository root.
run.py sets PYTHONHASHSEED=0 for itself, so GerryChain's chains are
reproducible without passing anything extra on the command line. (Python fixes
hash randomization at interpreter startup, so this cannot be done from inside a
running process — run.py re-launches itself once with the seed set. The
py_env file is no longer needed, but uv run --env-file py_env … still works
and skips that relaunch.)
Pass the config by name. The name resolves with or without the directory and extension, so these are equivalent:
uv run run.py basic
uv run run.py basic.json
uv run run.py configs/basic.json
Each of those runs the basic config end to end: it generates the geodata if
data/san_diego.geojson is missing, then district plans, VoteKit settings,
ballots, elections, and summaries.
Outputs are keyed by the config's run_name, not its filename — basic.json
has "run_name": "Basic - 3 X 3", so its results land in
outputs/Basic - 3 X 3/.
Re-running is cheap. Every stage checks whether its outputs already exist and
are still valid for the current config, so a second run.py basic picks up
where the first left off instead of redoing finished work.
uv run run.py --run-all
Runs each config in configs/ in turn, then draws one cross-run comparison
figure over all of their summaries. Project-wide settings are not a run, so
project-settings.json at the repo root is not picked up here.
uv run run.py
With no arguments you are prompted to use an existing configuration file or create a new one.
- Choose yes to give the path to an existing config (e.g.
configs/basic.json). - Choose no to be walked through building one. It is saved to
configs/<run_name>.jsonand then run. Only run-specific parameters are asked for — see Configuration for the split.
Note: The first time you run this command, it may take a moment before the prompt appears. This is because some imports take time to load. Subsequent runs will start much faster.
The data-generation stage is also its own entry point, which is useful after
changing anything under geometry_data and before committing to a full run:
uv run python -m pipeline.data_generator --config configs/basic.json
This writes geodata_path plus its adjacency graph and a metadata sidecar, and
nothing else. It reads the same layered configuration as run.py, so the
geometry can live in project-settings.json.
However it is started, the pipeline executes the whole workflow sequentially, each stage reading the previous stage's outputs:
| Stage | Script | Summary |
|---|---|---|
| 0 | data_generator.py |
Builds the block-level geodata for the city — geometry, VAP by race, elections, and district labels — and is skipped when geodata_path already holds valid data |
| 1 | districts_generator.py |
Generates district plans using GerryChain by converting geographical data into a graph |
| 2 | settings_generator.py |
Creates VoteKit settings JSONs by aggregating population data and computing turnout-adjusted bloc proportions for subsampled district plans |
| 3 | profile_generator.py |
Generates voter preference profiles (simulated ballots) for each settings file under three voting behavior models (impulsive, deliberate, and Cambridge) |
| 4 | simulate_elections.py |
Runs the election simulation (FastSTV) on the generated voter profiles to determine and record the winners |
| 5 | summarize_results.py |
Post-processes the election results into a dataframe, generates the matplotlib figures, and writes the run's JSON artifacts for the report |
| 6 | report_generator.py |
Collects every completed run's artifacts into docs/ and renders docs/index.html, the published report |
The report is a pipeline stage, not a document anyone edits. run.py rebuilds it
at the end of every run, so the ordinary way to publish new results is simply to
run the config.
To rebuild it from the artifacts already in outputs/ without re-simulating:
uv run run.py --report-only
uv run python -m pipeline.report_generator does the same thing directly.
Neither reaches the network. --vendor re-downloads the web fonts into docs/
and is only needed if they are missing; Bootstrap and D3 load from a CDN, pinned
by version and SRI hash.
docs/index.html cannot be opened from the file system. The page loads its
charts as ES modules and its data with fetch, and browsers block both over
file:// as cross-origin requests — the page comes up empty. Serve the
directory instead:
uv run python -m http.server -d docs 8000
then open http://localhost:8000. (Opening the file directly shows a message saying this rather than failing silently.)
| Path | |
|---|---|
docs/index.html |
Generated. Every rebuild overwrites it. |
docs/data/ |
Generated. One directory per run, mirrored from outputs/<run>/summaries/report/, plus manifest.json. A run whose outputs are gone has its directory removed, so what is published always corresponds to a run that exists. |
docs/prose/*.md |
Written by hand — abstract.md, background.md, methodology.md, conclusion.md. Markdown, rendered into the page. Empty ones show a placeholder, so the site is publishable before the writing is done. |
docs/prose/runs/<slug>.md |
Written by hand — one optional description per run, rendered under that run's heading. generate_report touches an empty file for every run that doesn't have one yet, so the exact path to edit is always there rather than only documented. Empty shows the same kind of placeholder as the top-level sections. |
docs/templates/, docs/css/, docs/js/ |
The page itself: structure, styles, and the D3 charts. Edited when changing how the report looks or behaves, not when results change. |
Adding a run therefore changes manifest.json, not any HTML.
docs/ is what GitHub Pages serves (Settings → Pages → branch main, folder
/docs), so it is committed rather than ignored.
Configuration is split in two. project-settings.json holds the settings that
describe the project rather than any one simulation, so they are written once
instead of being copied into every run:
| Key | Description |
|---|---|
geodata_path |
Path to the geographic dataset every run reads (.geojson or .gpkg) |
geometry_data |
How that dataset is built: state/county, CRSs, block source, Districtr plan, council district layer |
gerrychain_output_dir |
Chain output location |
population_column, population_vap_column, pop_of_interest_column |
Columns carrying total population, VAP, and the focal group |
seed |
Random seed |
chain_length |
Total steps in the Markov chain |
It lives at the repo root, alongside run.py. Every file in configs/ is one
run. A run config is layered over the project settings at load time, and any
key it sets wins — nested objects such as geometry_data merge key by key, so a
run can override a single geometry setting without restating the rest.
The interactive setup prompts only for run-specific parameters:
| Prompt | Type | Description |
|---|---|---|
| Run name | string | Identifier used for output directories and logs |
| Total number of seats | integer | Total number of representatives elected |
| Number of districts | integer | Number of districts in a district configuration. This value must evenly divide the total number of seats so that the number of winners per district is an integer. Users may specify multiple district configurations. |
| Number of simulated elections per district plan | integer | Number of simulated elections per district plan |
| Group names | string | Names of the bloc groups, comma-separated (e.g. A, B). Specify focal group first. |
| Candidate names | string | Names of the bloc group's candidates, comma-separated (e.g. A1, A2, A3). |
| Cohesion parameters | float (0-1) | Probability that voters from a group vote for candidates from their own group. Higher values indicate stronger within-group voting cohesion. |
| Candidate strength parameters | positive float | Shape parameters of the Dirichlet distribution that control how voters within a group distribute their preferences across candidate slates. α = 0 models perfect consensus among voters, α = 1 neutral preferences, and α → ∞ indifference. |
| Turnout | float (0-1) | Turnout rate for each voter bloc |
Each entry in district_configs controls how many candidates its districts put
on the ballot. The pool is drawn per district as a binomial over the range
[floor, candidate_pool_max], with the mean placed at candidate_pool_mean:
"district_configs": [
{ "num_districts": 9, "winners": 1,
"candidate_pool_max": 9, "candidate_pool_mean": 5.67 }
]| Key | Meaning |
|---|---|
candidate_pool_max |
Ceiling on the pool. Required. |
candidate_pool_mean |
Where the average pool size sits. Required. |
The floor is not configurable: it is the smallest ballot every configured
voting rule can actually run on (minimum_candidates), one more than the most
demanding requirement in the run.
Neither key is validated — keep the mean inside [floor, candidate_pool_max].
A mean outside that range gives numpy a probability outside [0, 1] and it
raises ValueError: p < 0, p > 1 or p is NaN; a ceiling equal to the floor
divides by zero. Because the pool size feeds the per-slate apportionment, it is
the main lever on how often a small slate fields anyone at all.
Both keys are part of the profile signature, so changing either regenerates the profiles for that magnitude.
A voter model determines the kind of ballot it produces, and a voting rule can only read one kind:
| Voter model | Ballots | Rules it can run |
|---|---|---|
slate_pl, slate_bt |
RankProfile |
STV, IRV, Plurality, Borda, the two-round rules |
name_cumulative |
ScoreProfile |
Cumulative, Limited |
Configure both families in one run and each rule is simulated under the models whose ballots it supports, with the rest skipped and reported:
[simulate_elections] slate_pl produces RankProfile ballots; skipping ['Cumulative', 'Limited'] (they need ScoreProfile).
[simulate_elections] name_cumulative produces ScoreProfile ballots; skipping ['FastSTV'] (they need RankProfile).
Each model writes its own results file, so the summary carries one row per
(model, rule) pair that actually ran. A rule accepting either type (e.g.
BlockPlurality) runs under both.
Score budgets. A score ballot is only valid for one budget: Cumulative
rejects any ballot spending more than n_seats points and Limited more than
its budget. Rules with different budgets therefore need different ballots, so
the profile stage reads the budgets off the configured rules and generates one
set per distinct budget, stored under its own subdirectory:
profiles.zip
├── slate_pl/3/…csv ranked ballots, no budget
├── name_cumulative/2/3/…csv every ballot spends exactly 2 points
└── name_cumulative/3/3/…csv every ballot spends exactly 3 points
So Cumulative (budget = n_seats = 3) and Limited (budget: 2) run in the
same simulation, each reading the ballots it can accept. Changing any budget
changes the profile signature and regenerates the score profiles, leaving the
ranked ones alone.
slate_pl and slate_bt ballots normally rank every candidate. A run can
instead truncate them to lengths drawn from Cambridge, MA's historical
2009-2017 RCV data — split by whether a ballot's first choice was the
historical majority or minority slate — via a top-level cambridge_truncation
block:
"cambridge_truncation": {
"enabled": true,
"method": "bounded",
"majority_slates": ["WAIO"],
"minority_slates": ["HIS", "AAPI", "BLK"],
"min_length": 3,
"max_length": 6
}| Key | Meaning |
|---|---|
enabled |
Must be true (not just present) to turn truncation on. |
majority_slates, minority_slates |
Every slate on the ballot must be pooled into exactly one group. majority_slates maps to Cambridge's historical White-majority ballots, minority_slates to its minority ballots. Omitting both falls back to putting WHI alone in majority and everything else in minority, which only suits the two-bloc model — a four-bloc config (e.g. WAIO/HIS/AAPI/BLK) must set both explicitly. |
method |
"historical" (default) samples straight from Cambridge's raw length distribution. "bounded" restricts it to [min_length, max_length] and renormalizes, dropping any historical lengths outside that range. |
min_length, max_length |
Required when method is "bounded"; ignored otherwise. |
Each ballot is classified by its own first choice, not by its voter's bloc, matching how the historical data itself is split. A grouped ballot (many voters sharing one ranking) has each voter draw its own truncation length independently rather than truncating the whole group to one length.
For a hybrid run (multiple district_configs tiers), tier_overrides lets one
tier use a different method/min_length/max_length than the rest, keyed by
that tier's num_districts (as a string):
"cambridge_truncation": {
"enabled": true,
"majority_slates": ["WAIO"],
"minority_slates": ["HIS", "AAPI", "BLK"],
"tier_overrides": {
"1": { "method": "bounded", "min_length": 6, "max_length": 10 }
}
}An override replaces method/min_length/max_length wholesale rather than
merging field by field — a bounded override with no min_length of its own
does not inherit an unrelated tier's bound. majority_slates/minority_slates/
enabled always come from the top-level block; every tier truncates the same
majority/minority split. A tier with no entry in tier_overrides uses the
top-level method/min_length/max_length.
cambridge_truncation is part of the profile signature, so changing any of
these keys regenerates profiles.
Truncating an existing run without regenerating profiles. Rebuilding
profiles is the most expensive stage to repeat, so
pipeline/truncate_profiles.py applies the same truncation as a second pass
over an already-generated profiles.zip (or general_profiles.zip /
primary_profiles.zip), reading the run's own cambridge_truncation config:
uv run python -m pipeline.truncate_profiles configs/basic-truncation.json
uv run python -m pipeline.truncate_profiles configs/basic-truncation.json --in-place
uv run python -m pipeline.truncate_profiles configs/basic-truncation.json --all-archives --seed 0
By default it writes <archive>_truncated.zip beside the original; --in-place
replaces it (via a temp file, so an interrupted run leaves the original
intact). Score profiles (name_cumulative) are always copied through
untouched, since they have no ranking to shorten.
A voting_configs entry that names a general_class is a two-round rule. The
two stages are named for what they do: the primary narrows the field by
Plurality to m_1 finalists, then a freshly-sampled profile restricted to those
finalists decides the general (see pipeline/two_round_election.py).
"AlaskaTwoProfile": {
"m_1": 4,
"tiebreak": "random",
"general_class": "STV",
"general_kwargs": { "n_seats": 1, "tiebreak": "random" }
}round2_class / round2_kwargs are still accepted as the former names of
general_class / general_kwargs, so existing configs keep running. Note that
renaming them in a config changes its election-results signature, which forces
that run to re-simulate.
Both stages are recorded. Alongside the usual
outputs/<run>/election_results/<mode>/<file>.json, which holds the general's
winners, a run with a two-round rule also writes
outputs/<run>/primary_results/<mode>/<same file>.json giving the finalists each
primary advanced. The two files carry the same signature and the same
profile_files list in the same order, so they line up row for row:
The general's ballots are kept too, in outputs/<run>/general_profiles.zip — one
resampled, finalists-only profile per (rule, profile file).
By default both rounds draw on the same electorate — the one turnout describes.
Add an optional top-level primary_turnout block to model a narrowing round with
lower participation, as a primary typically has:
"turnout": { "WAIO": 1, "POC": 0.7 },
"primary_turnout": { "POC": 0.4 }
"turnout": { "AAPI": 0.75, "HIS": 0.75, "WAIO": 1, "BLK": 1 },
"primary_turnout": { "AAPI": 0.4, "HIS": 0.4 }It is a partial override: blocs left out keep their turnout rate, so only
the ones whose participation drops need naming. It applies to round 1 of every
two-round rule; round 2 and every single-round rule keep using turnout
unchanged. VoteKit's built-in Alaska and TopTwo build round 2 by stripping
their round-1 ballots rather than resampling, so they cannot use a separate
primary electorate — use AlaskaTwoProfile / TopTwoTwoProfile instead, and the
simulation stage warns if the built-ins are configured alongside
primary_turnout.
Setting it makes the profile stage generate a second archive,
outputs/<run_name>/primary_profiles.zip, holding one primary-round profile per
entry in profiles.zip. That roughly doubles profile-generation time for the
run: a different electorate means different ballots, and ballots cannot be
reweighted after sampling.
Two parameters are currently being set to a default value:
| Parameter | Value | Description |
|---|---|---|
| Number of subsamples | 100 | Number of district plans to retain for election simulation. |
| Number of voters | 10,000 | Number of voters for each simulation. |
Chain length was previously defaulted here; it is now a project-wide setting in
project-settings.json.