Methods & evidence
Trace every
conclusion.
What we measure.
How we check it. Where it stops.
- Collect
OONI · IODA · CensoredPlanet · Voidly probes
- Compare
Correlate evidence. Classify signals.
- Publish
A country score, an incident and its sources.
Method overview. Each published result still needs its own source, scope and date.
Model statistics in this reference come from the bundle. Network and sample responses show available source dates.
Current model records ↗Evidence ≠ cause
An outage or anomaly is not automatically deliberate censorship.
Score ≠ probability
An index score is different from data confidence or a forecast.
History ≠ timing
Country transfer, future events and calibration need separate tests.
The complete reference
Go as deep as you need.
Overview
The Global Censorship Index combines measurements from multiple networks into a country score, with classification and evidence processing between collection and publication. A score summarizes measured conditions; it does not observe every network or establish the cause of every anomaly.
The published v2.2 description says collection runs continuously and new blocks can appear within hours. Its limitations allow up to 24 hours for events to reach scores. Treat those as pipeline descriptions; use actual source timestamps when citing a result.
Data sources
OONI measurements
- Source-reported samples
- 38,780,449
- Source basis
- all-time: sum of ooni-data.json row totals (not the rolling 30-day scoring window)
- Sample source
- json
- Source timestamp
- Not supplied
- Curated index coverage
- 130 countries in the bundled index
- Tests
- Web, messaging, circumvention
Sample totals and curated index coverage describe different collections. Inspect the sample response.
Voidly sensor network
- API active nodes
- 15
- API total nodes
- 16
- Probes in 24 hours
- 140,235
- Returned regions
- Latin America, Africa, North America, Europe, Oceania, Asia Pacific
- Source updated
- 2026-09-24T00:51:36.915216Z
active_nodes = core Voidly fleet; active_nodes_total = core + community (reconciles with /v1/probe/stats active_nodes).
Inspect network scope and nodes. Missing network data is not replaced with the original page’s approximate count.
External measurement sources
- IODA
- Internet outage detection at ASN level
- CensoredPlanet
- Remote DNS/HTTP blocking; the reference describes coverage of 50 countries
- Citizen Lab
- Domain categorization: 5,154 distinct domains ingested, measured 2026-08-27 at /v1/probe/domains distinct_total
Probe coverage limits
The reference’s configured network footprint is not a live active-node count. The described deployments include the US, UK, Germany, Japan, Singapore, India, Brazil and South Africa. They help answer whether a domain is reachable from outside a country, and do not establish inside-country reachability. The validated source response above supplies the available node counts and region list.
The reference identifies limited inside-country coverage in IR, CN, RU, VE, EG, PK, MM, SY, TR, SD, KP, BY, UZ, TM, SA, VN, TH, LB and AZ. Inside-country evidence relies on volunteer OONI measurements and remote CensoredPlanet DNS/HTTP probing. Community operators can see the probe program.
Request checked: . A page check is not a measurement timestamp.
Multi-source correlation
No single network captures the complete picture. OONI provides active probing with geographic gaps. CensoredPlanet offers remote measurement with limited ground truth. IODA detects outages rather than selective blocking.
The v2.2 pipeline describes Voidly’s network testing VPN accessibility and censorship patterns every five minutes, then correlating those measurements with OONI, CensoredPlanet and IODA. Correlation produces structured incident records and linked evidence; an anomaly by itself remains ambiguous.
Models & historical evaluations
These reference metrics are dated. The bundled metric snapshot was generated . The source’s May 2026 experiments and older forecast panels are identified separately below. Open the model registry preview for currently returned versions, source notes and dates.
Classifier v3.3 · bundled snapshot
Gradient boosting trained on 4,237 labeled samples, including 1,116 positive samples across 131 countries. The reference describes aggregate-data training with no raw user data. The live classifier record also explains that many positive labels are network disruptions; read its label-composition caveat before treating F1 as confirmed-censorship accuracy.
- Algorithm
- GradientBoosting
- LOCO mean F1
- 0.71
- Mean F1 for countries with n≥30
- 0.63 across 61 countries
- LOCO median F1
- 0.87; inflated by 46 of 127 countries at a perfect 1.0 on very small samples
- Stratified five-fold F1 / AUC
- 0.73 / 0.90; sanity checks, not deployment performance
- Hard, high-volume countries
- CN 0.29 · BY 0.21 · AZ 0.11
- Training samples
- 4,237 (1,116 positive) across 131 countries
- Reference retraining schedule
- Weekly, Sundays at 02:00 UTC
The headline LOCO MEDIAN F1 (0.870) is dominated by many small-sample countries scoring a perfect 1.0 on a handful of days; the MEAN F1 (0.711) is the honest single number. Censorship-heavy, high-volume countries score materially lower (CN ~0.29 on n=95, BY ~0.21 on n=84, AZ ~0.11 on n=65). Cite the mean — or the specific per-country number — for hard countries, not the median. The model is CLEAN (no label leakage); this is a distribution caveat, not an accuracy retraction.
The earlier v2 reported 99.8% F1 and 1.000 ROC AUC, and was retired on May 21, 2026 after an audit found that the label-derived country_risk_tier feature contributed 85% of importance. Read the leakage audit.
Classifier feature importance · bundled reference
v3 removed the country-tier shortcut. The original v3 account describes roughly 73% importance in its top three signals and no single dominating feature. The following stored importance list belongs to the later v3.3 reference, with 13 base features plus three regime-similarity-weighted contagion features; do not mix the two generations.
- anomaly_rate
- 22.1%
- month
- 20.2%
- measurement_count
- 17.0%
- neighbor_max_anomaly_7d
- 8.9%
- neighbor_incident_count_7d
- 7.7%
- neighbor_block_rate_7d
- 7.6%
- rate_count_interaction
- 5.6%
v3.3’s reference records stratified F1 0.729, LOCO mean F1 0.711 and inflated median F1 0.870. Inspect current feature importance.
Legacy Sentinel forecast · bundled snapshot
XGBoost with isotonic calibration is a separate model from both classifier v3.3 and honest_forecast_v1. Its legacy target_7day sliding-window label measures the prevailing regime; these scores must not be attached to the newer OONI-event forecast.
- Training-holdout LOCO median AUC / F1
- 0.88 / 0.51
- Snapshot status
- DEGRADED — rolling 30d precision 0.172 < 0.6
- Rolling 30-day precision / recall
- 0.17 / 32%
- Rolling Brier / calibration MAE
- 0.28 / 0.29
- Evaluated country-days
- 898
- Training records
- 14.6K
- Reference retraining schedule
- Weekly, Sundays at 02:00 UTC
The public accuracy endpoint describes daily evaluation over a rolling 30-day window. The values above are the dated bundle, not a fresh endpoint fetch.
Earlier legacy splits, calibration and onset audit
The original reference contains different legacy holdout panels. They are preserved here as separate historical records; none is a current classifier evaluation or proof of future shutdown timing.
- Earlier LOCO median panel
- AUC 0.91 · F1 0.55; train on 18 countries, test on the 19th, median of 19 holdouts
- Earlier time-based panel
- AUC 0.50 · F1 0.00; train before T, test after T
- Earlier stratified panel
- AUC 0.98 · F1 0.79; 15% random holdout, temporal leakage, sanity check only
- Later validation panel: stratified
- AUC 0.98 · F1 0.84; not a deployment claim
- Later validation panel: time-based
- AUC 0.48 · F1 0.47; below-chance AUC
The reference replaced a 0.998 F1 headline after the accuracy endpoint warned that stratified AUC overstated time-based performance by 47.9 percentage points.
The May 20, 2026 calibration audit found Brier 0.59 and calibration MAE 0.60; the 5% predicted-risk bucket had a 65% observed incident rate. Isotonic regression was refit to 810 live predicted/observed pairs, giving in-sample Brier 0.22 and MAE 0.00. A watched-country gate limited extrapolation. These training-fit improvements do not establish new-event forecast skill. Calibration drift, reliability diagram, full refit finding.
The legacy target is 98.9% autocorrelated day to day: a predict-yesterday baseline reaches AUC 0.957. A strict forward-temporal audit found model AUC 0.589 on the raw label and about 0.33 on actual onset transitions, below chance. It reflects a current censored regime, with essentially no demonstrated skill at predicting a new shutdown before it occurs. Read the onset-skill finding.
May 21, 2026 ensemble and experimental models
These are dated model-history entries from the v2.2 reference, not newly verified serving states.
- Multi-horizon forecast: separate 1-, 7- and 30-day XGBoost/isotonic models; reported LOCO AUC 0.91 / 0.88 / 0.84, per-horizon SHAP, 90% conformal intervals and a monotonicity check.
/v1/forecast/{cc}/multi-horizon. Their sliding-window targets share the legacy label-autocorrelation problem; shuffled/LOCO AUC does not establish onset skill. Onset audit. - ACI online conformal: Gibbs & Candès (2021); the reference describes online adaptation, initial α 0.10 → 0.21 and empirical coverage 91.3% in legacy
/v1/forecast/{cc}/7dayresponses. This historical uncertainty description is not an interval for the newer event model. - CenDTect DBSCAN: per-country unsupervised rolling 45-day window, reported AUC 0.6506, promoted as a second opinion.
/v1/anomaly/dbscan/{cc}. - Per-domain HDBSCAN drift: weekly processing of novel domain-level blocking patterns.
/v1/anomaly/domain-drift/leaderboard. - Per-measurement classifier: Niaki KDD23 row-level XGBoost, reported AUC 1.0 with the caveat that the model reconstructs the labeling rule.
POST /v1/measurement/classify. - AS-topology GNN: two-layer GraphSAGE on a 7,060-node CAIDA peering graph; LOOCV AUC 0.80 on six tier-1 ASNs, an underpowered sample.
/v1/forecast/asn-gnn/{asn}.
Scoring system
The published index uses a 0–100 scale: 0 denotes complete freedom in the scale definition, and 100 denotes total censorship. These are descriptive score bands, not forecast probabilities or the separate data-quality confidence bands.
- 0–10 · Free
- Minimal or no censorship
- 11–25 · Low
- Limited content restrictions
- 26–45 · Medium
- Significant restrictions on some platforms
- 46–70 · High
- Widespread blocking of platforms and news
- 71–100 · Severe
- Pervasive censorship / isolated internet
Limitations
- National averages do not capture regional variation.
- VPN detection is underreported in highly restricted environments.
- Sample sizes vary by country and affect confidence.
- Real-time events may take up to 24 hours to reach the scores.
- Content filtering and throttling are harder to detect than blocking.
- Self-censorship and legal restrictions are not measured.
Confidence intervals
The reference describes country-score intervals reflecting measurement certainty. Wider intervals indicate less data or greater variability. The examples below are illustrative only; they are neither current countries nor a forecast validation result.
| Country | Illustrative score | Interval | Confidence | Note |
|---|---|---|---|---|
| Country A | 66% | ±2% | High | Large sample |
| Country B | 42% | ±4% | High | — |
| Country C | 31% | ±3% | High | — |
| Country D | 21% | ±7% | Medium | Smaller sample |
Open actual index scores. A measurement-confidence interval, a model score, a probability and an empirical forecast-coverage statistic answer different questions.
Validation & reproducibility
The published workflow compares scores with external benchmarks, known events and three holdout approaches. The Freedom House Freedom on the Net comparison is a self-reported correlation of r=0.87. The reference gives Iranian shutdowns matching score spikes as an example of event checking.
- Stratified k-fold
- Random splits provide a sanity check but can leak temporal and country patterns. Do not cite them as deployment performance.
- Leave-one-country-out
- Train on all other countries, evaluate on the held-out country. This tests geographic transfer, not future-time generalization.
- Time-based
- Train before T, test after T. This tests temporal transfer and can expose failure on novel events.
Classifier and forecast results remain separated in the model section above. The accuracy source explicitly warns that stratified AUC can overstate the time-based split by 47.9 percentage points, and recommends LOCO or populated rolling-production results. Even a LOCO median needs the classifier’s small-sample caveat; it is not universally the best headline. Read the replication guidance.
Independent replication
The reference states that no independent third-party evaluation has been conducted. Published data and transparency endpoints are available for independent replication. Keep a copy of the source response, version, training/source dates, target, split and caveats when citing performance.
Update pipeline
- OONI
- Ingestion
- Feature engineering
- ML scoring
- Index update
Published v2.2 schedule, rather than a guarantee that the last run completed:
- Classifier retraining
- Weekly, Sundays at 02:00 UTC
- Publication
- Daily at 03:00 UTC
- Score latency
- Approximately six hours for aggregates and five minutes for probes
- Raw ingestion
- Every five minutes for probes; every six hours for OONI, IODA and CensoredPlanet
The reference describes KS-test drift detection on a rolling holdout triggering extra retraining, a D8 promotion gate covering G1–G7 plus calibration, thresholds for AUC/recall/coverage, and a post-promotion watchdog that rolls back regressions. It describes automated publication without a human in that loop. The existing operational page carries the counters and source; the pipeline diagram here does not verify a particular run.
Citation & licensing
Cite the methodology page
Cite the index dataset
Voidly Research. (2026). Global Censorship Index. https://voidly.ai/censorship-index
@misc{voidly_censorship_index,
author = {Voidly Research},
title = {Global Censorship Index},
year = {2026},
url = {https://voidly.ai/censorship-index}
}Voidly-original incidents, scores and annotations are CC BY 4.0, permitting reuse with attribution. Raw upstream measurements keep their own licenses. As the data page specifies, OONI raw data and Citizen Lab test lists use CC BY-NC-SA 4.0; redistribution of those subsets must honor those terms. Source-specific data and licensing.
Data access
- Full index JSON with schema.org markup
- Index CSV for Sheets, Excel and Pandas
- Machine-readable methodology
- Detection, prediction and realtime API documentation
The machine-readable methodology retains some legacy model fields. Use the versioned model source for its features, target and evaluation; do not infer one common model from the general method document.
Contact
- Research
- research@voidly.ai
- Partnerships
- partnerships@voidly.ai
- General
- team@voidly.ai