Research · Synthetic Polling Methodology
← Newsroom
Methodology research · July 2026

Averaging five AI models reduced polling error in our test

We tested five AI models from five labs on the same panel of 6,400 registered-voter profiles, simulating responses in eight past elections with certified results. Across seven races, the panel average had less error than any individual model. The models disagreed most on the eighth race, which all five misread.

By · Chief Technology Officer, Civly ·
AI models tested
5
Simulated voters
6,400
Certified races graded
8
Panel average typical miss
±1 pt
The result

The panel average had the lowest error

Finding

Across seven races, the panel average missed by about 1 point. The best individual model missed by 2.3 points and the weakest by 4.5. Averaging reduced the error from models that leaned in different directions.

Each model simulated the same 6,400 voters — real registrations drawn from the Pennsylvania, North Carolina, Florida, and Kansas voter files — answering the 2024 presidential race in all four states, the 2024 Senate races in Pennsylvania and Florida, the 2024 North Carolina governor’s race, and the 2020 Kansas Senate race. We graded every reading against the officially certified result, scored so that no model is ever graded on a race it was corrected against.

EngineTypical miss vs. certified results
The panel averageMean of the models, one reading per race±1 pt
GPT‑5 miniOpenAI±2.3
Claude Sonnet 5Anthropic — our previous single engine±2.6
Gemini 3.5 FlashGoogle±3.2
GLM‑5.2Z.AI±3.5
DeepSeek v4 FlashDeepSeek±4.5

Typical miss measures the size of the errors across seven races. Each model's estimate was corrected using only the other races, then compared with the certified result. Differences under about a point are treated as ties. The eighth race is examined below.

The stress test

All five models misread North Carolina's governor race

In the 2024 North Carolina governor's race, Mark Robinson faced extensive scandal coverage and Josh Stein won by 14.8 points. The five models produced widely different estimates, and all missed the result.

EngineIts reading of NC governor 2024
Claude Sonnet 5Stein +68
Gemini 3.5 FlashStein +45
GLM‑5.2Stein +30
DeepSeek v4 FlashStein +1
GPT‑5 miniRobinson +3
Certified resultStein +14.8
Model disagreement

The models' estimates were 4 to 11 points apart on the other seven races. In North Carolina, the spread was 71 points. Civly now uses disagreement between models to flag estimates that need closer scrutiny. The flag indicates model uncertainty; it does not establish how volatile a race is.

Method and limitations

Testing against past election results

Certified results. All eight races have official results, including Pennsylvania's 2024 Senate race, decided by 0.2 points. Each model used the same voter profiles, questions, and scoring procedure.

Separate calibration and evaluation. Corrections to each model's estimates were fitted using only the other races in the test.

Possible memorization. Four of the five models could recite the certified results when asked, which may make these scores more favorable than results on future elections. GPT-5 mini, the only model that could not recite them, was the best individual performer. All five still missed the North Carolina governor race. We are recording predictions for 2026 races before election day and will publish them alongside the results.

Adding news did not improve the results. We tested date-verified pre-election headlines as additional context for the simulated voters. They did not resolve the North Carolina error, and their slant affected the estimates. We excluded them from the method.

What this changes

Every Civly candidate poll now runs on a panel of AIs

Starting this week, Civly synthetic candidate polls use multiple AI models from different labs for each simulated voter. They report the panel average and include a volatility flag when the models disagree substantially.

Run it on your race

Use synthetic polling in your race

Civly's earlier method called 9 of 10 winners in the 2026 New York and Maryland primaries before voting began. Candidate polls now use a panel of models.


Run dates July 16–17, 2026 · identical simulated panels of 1,600 registered voters per state, drawn from the Pennsylvania, North Carolina, Florida, and Kansas voter files · eight officially certified races graded leave‑one‑out with per‑state corrections · no names, addresses, or phone numbers are ever shown to any model · prepared by Civly.