The AI Rarely Decides

A Screened Inventory of Artificial Intelligence for Democratic Discourse

Krishan Mathis · Draft for journal submission · 2026-07-16 · derived from the project manuscript (v3). Target: journal in digital democracy / governance of technology (e.g. Policy & InternetDigital Government: Research and PracticeJournal of Deliberative Democracy). Citation style: author–date. All references verified against the sources listed; project dataset catalog_verified_v2.csv.


Abstract

Artificial intelligence is widely believed to be reshaping democratic discourse, and a growing literature catalogues “AI tools for democracy” on the assumption that a substantial AI-native ecosystem exists and increasingly exercises democratic judgment. We test that assumption. We assembled 158 candidate systems from the field’s own directories, grant programmes, and literature, screened them for unit of analysis and liveness, and coded each surviving system on two independent criteria — whether the AI is constitutive of the system’s democratic function, and whether the AI decidesor merely supports a human decision — verifying every attribute against primary documentation rather than promotional material. Three findings contradict the field’s self-description. First, of 123 live tools, fewer than half are AI-constitutive; flagship systems including Pol.is and Community Notes document that their load-bearing algorithms are classical statistics (Small et al., 2021; Wojcik et al., 2022). Second, of 100 live tools containing any AI, the model makes the consequential judgment in only 28; in 72% of cases it scores, ranks, flags or summarises while a human decides — and the most AI-constitutive products are the least autonomous, by their builders’ own documentation (Jigsaw, 2024; Full Fact, 2025). Third, fewer than half of the AI-constitutive systems are addressed to citizens; the remainder serve platforms, newsrooms, researchers, lobbyists, and the state. An adversarial verification pass overturned 18% of our own first-pass coding, which we report as a finding about the reliability of citation-based ecosystem claims. We argue that governance scrutiny is currently calibrated to model sophistication rather than to decision consequence, and propose the principle govern authority wherever it sits — in the model, in the human loop, or in the architecture.

Keywords: artificial intelligence · democratic discourse · deliberation · civic technology · content moderation · algorithmic governance · taxonomy

1. Introduction

Every major communication technology has reshaped democratic practice, and artificial intelligence is routinely described as the next such transformation — one that does not merely carry democratic communication but participates in it: summarising arguments, moderating discussion, identifying common ground, and facilitating deliberation at scales no human moderator could reach (Carnegie Endowment, 2026a; Landemore, 2020; Summerfield et al., in Tessler et al., 2024). Empirical demonstrations exist: an LLM mediator produced group statements that participants preferred to those of human mediators (Tessler et al., 2024); dialogues with a language model durably reduced conspiracy beliefs (Costello, Pennycook, & Rand, 2024); AI-assisted chat interventions improved the quality of online political conversations (Argyle et al., 2023).

That is an accurate description of what AI can do. It is not, we find, an accurate description of what is deployed.

This study began as a conventional survey intending to catalogue the ecosystem of AI systems supporting democratic discourse and to derive a functional taxonomy from it. When the catalogue was actually coded against primary sources, the taxonomy’s empirical foundation dissolved. The systems most often cited as evidence of AI’s democratic transformation are frequently not AI: Pol.is states in its own documentation that it does not use natural language processing, delivering its function through principal component analysis and clustering over a vote matrix (Computational Democracy Project, n.d.; Small et al., 2021). Community Notes adjudicates note visibility through matrix factorisation and logistic regression (Wojcik et al., 2022). Stanford’s widely cited “automated moderator” is rule-based — speaking queues and timers (Stanford Deliberative Democracy Lab, n.d.). And where AI is genuinely present, it overwhelmingly advises rather than decides.

We therefore reframe the research object. Rather than mapping an assumed ecosystem, we ask what remains of the field’s population when an explicit inclusion criterion is applied (RQ1–RQ2), whether the AI that remains exercises democratic judgment or supports it (RQ3), for whom the surviving systems are built (RQ4), and why the field believes a larger, more mature, and more autonomous ecosystem exists than the evidence supports (RQ7). The contribution is threefold: an inclusion criterion the literature lacks; a screened, primary-source-verified inventory with its attrition reported; and an authority-centred reading of democratic AI with consequences for governance research.

2. Background and related work

The intersection of AI and democracy is studied across political science, communication studies, computer science, human–computer interaction, computational social science, and AI ethics. The resulting literature is fragmented — but our analysis suggests fragmentation is the lesser problem. The greater one is a shared blind spot: none of the traditions requires a system to actually use AI before counting it as AI.

Civic technology research catalogues participation infrastructure — petitions, consultations, participatory budgeting — and increasingly describes it as AI-enabled, usually on the strength of a feature (People Powered, 2025; Democracy Technologies, n.d.). The platforms carrying real participatory load are instructive: Decidim, launched in 2016 from Barcelona’s municipalist movement, structures proposal, deliberation, amendment and accountability stages for hundreds of institutions without AI in its critical path (Barandiaran et al., 2024; Borge et al., 2026); its official AI module performs Bayesian spam detection. Consul Democracy’s LLM civic assistant is a funded project that has not shipped (Consul Democracy Foundation, n.d.).

Deliberative democracy theory evaluates systems by their effects on deliberative quality — inclusion, reason-giving, reflection (Habermas, 1996; Fishkin, 2018). This is the right normative test, but it is indifferent to implementation: a system that improves deliberation through interaction design scores identically to one that improves it through a language model. Pol.is enters the AI literature largely through this door, via vTaiwan’s documented policy successes (Participedia, n.d.; vTaiwan, n.d.), without interrogation of its mechanism.

AI governance research asks how AI systems should behave — fairness, accountability, explainability, oversight — and has been indispensable in naming risks (Carnegie Endowment, 2026a; Cogitatio, 2025). But by taking the AI-ness of its objects as given, it generalises governance concerns across systems whose algorithmic content ranges from fine-tuned language models to naive Bayes spam filters.

Computational social science is the tradition most attentive to what algorithms actually do — and, tellingly, the least prone to the error, because its practitioners distinguish PCA from a transformer (Wojcik et al., 2022; Argyle et al., 2023).

Existing classification strategies — technical, domain-based, governance-focused, application-based, institutional — each capture something and miss the same thing: only the technical classification tests for AI at all, and it abandons the democratic question to do so. What is needed is a sequence: screen technically, then classify functionally.

3. Two inclusion criteria

3.1 The constitutive test

A system is AI-constitutive if removing its AI component destroys its democratic function. It is AI-augmented if removing the AI leaves a working democratic system with fewer conveniences. It is AI-absentif there is no AI to remove.

The test is counterfactual and can be applied to documentation rather than marketing copy. The Habermas Machine is AI-constitutive: a fine-tuned language model plus reward model that drafts group statements maximising predicted endorsement; remove the model and nothing remains (Tessler et al., 2024). Pol.is is AI-absent or at the outer edge of AI-augmented: its documentation states that no natural language processing is involved (Computational Democracy Project, n.d.). Decidim is AI-augmented at most. Consul is AI-absent today — a roadmap is not a finding.

The distinction is not pedantic. It changes the population (removing roughly half); it changes the causal story (if Pol.is is vote-matrix clustering, vTaiwan’s outcomes credit a well-designed statistical instrument and a decade of civic institution-building, not AI); and it changes the governance questions (the risks of a spam filter and of a model that drafts the group’s collective statement are not the same risks).

3.2 The authority test

A second, independent criterion asks whether the AI decides — its output is the determination, or it acts autonomously — or supports: it scores, ranks, retrieves, flags or summarises while a human makes the judgment. The two criteria dissociate sharply. ClaimBuster does not exist without its classifier (constitutive) but the classifier only flags check-worthy sentences for human adjudication (supports) (Hassan et al., 2017). Conversely, Discord’s AutoMod blocks messages autonomously (decides) using predominantly regular expressions and wordlists (augmented at best) (Discord, n.d.).

3.3 Unit of analysis

The unit is the artefact: a discrete, deployable system whose AI component, if any, can be inspected in documentation or source. Movements (g0v), organisations (mySociety), programmes, prize competitions and event series are recorded as ecosystem attributes, not rows. This rule alone ejected 19 entries from the candidate pool, several of them load-bearing citations in the literature.

4. Methods

4.1 Design

An inductive screening-and-coding study: (1) landscape construction from the field’s own directories, grant programmes and literature; (2) unit-of-analysis screening; (3) liveness screening; (4) dual-criterion coding; (5) adversarial verification against primary sources; (6) analysis of distributions, attrition, authority and audience; (7) taxonomy only where the data supports it.

4.2 Sources and sample

Candidates (N = 158) were drawn from the Civic Tech Field Guide, the Democracy Technologies database, People Powered’s platform ratings, OECD OPSI, all ten grantees of OpenAI’s Democratic Inputs to AI programme, the Collective Intelligence Project and Plurality Institute, arXiv/CHI/CSCW proceedings 2023–2026, and dedicated sweeps of the fact-checking, content-moderation, civic-learning and horizon-scanning literatures — including non-English (German, Japanese) sweeps after an initial English-only pass produced demonstrably false negatives (see §4.4).

Disclosure. An earlier project draft referred to a curated dataset of “several hundred” tools. That dataset did not exist; the inventory reported here was built from scratch. A paper arguing that the field asserts empirical foundations it does not have cannot itself assert one.

4.3 Coding and rules of evidence

Each system was coded against official documentation, source repositories, peer-reviewed publications, model cards, and algorithmic transparency records; promotional copy was not accepted as evidence of AI use — the single highest-yield rule, as several systems describing themselves as AI-powered document internally that they are not. Coders recorded unknown rather than inferring, flagged apparently defunct or misidentified systems, and attached a mandatory free-text evidence field to every constitutive call so that any classification can be contested without redoing the research. Seven independent coding passes were run.

4.4 Verification and coder error

The initial coding was adversarially re-verified, with coders instructed to distrust the preliminary classification. The pass overturned 29 of 157 entries (18%) — in the fact-checking domain, 10 of 20 (50%). We report this as a finding rather than a limitation: a disciplined pass with an explicit codebook still erred at nearly one in five on first attempt. The field, by contrast, does not code; it cites. Verification also caught our own methodological error: an early draft compared domains coded against different criteria (epistemic vs. artifact), inflating a reported two-ecosystem split from 1.6× to 2×; the corrected figures appear in §5.4.

4.5 Limitations

This is a screened inventory, not a census; the Plurality Institute / Prosocial Design Network open dataset (70+ tools) remains unreconciled. The AI boundary is a judgment — a broader definition admitting classical unsupervised learning would reclassify Pol.is, Community Notes and derivatives, materially changing the headline; evidence is recorded per row so the judgment can be revisited. Coder disagreement concentrates where removing the AI destroys the tool but not the practice (ClaimBuster, Full Fact). Vendor evidence quality is inversely correlated with autonomy claims. The ecosystem changes monthly; all figures are a snapshot (July 2026). Effectiveness is not measured: screening establishes what exists and what it is permitted to do, not whether it improves democratic outcomes.

5. Results

5.1 Attrition

Of 158 candidates, 19 were category errors (programmes, prize competitions, methods, research apparatus), 13 were defunct, dormant or never shipped, and 2 out of scope, leaving 123 live, in-scope tools. Of these, 58 (47%) are AI-constitutive, ~41 (33%) AI-augmented, and ~24 (20%) AI-absent. Roughly half the systems the field counts as AI tools for democracy are not AI tools, and the attrition is driven by flagships: Pol.is (PCA/clustering; Computational Democracy Project, n.d.), Community Notes (matrix factorisation; Wojcik et al., 2022), Stanford’s rule-based “automated moderator,” Ethelo’s mixed-integer solver with hand-coded comments — on which Engaged California runs — and Decidim, Consul and DemocracyOS.

5.2 The central result: the AI almost never decides

Of the 100 live tools containing any AI, the model makes the consequential judgment in 28. Authority varies by domain: research/AI-governance 64%, information integrity/governance 33%, civic learning/government 20%, deliberation/participation 17%. In 72% of deployed cases the AI advises and a human decides.

The dissociation runs opposite to intuition: the most AI-constitutive products are the least autonomous, by their builders’ own documentation. Perspective API’s model card lists “fully automated moderation” under uses to avoid(Jigsaw, 2024). Full Fact states that AI should serve editorial judgment rather than replace it (Full Fact, 2025). Wikipedia’s ORES scores edits but does not act on them (Halfaker & Geiger, 2020). Cofacts assigns the epistemic work to roughly 2,000 volunteer editors, using AI only for deduplication (Rights CoLab, 2023). Community Notes lets humans write and rate all notes, delegating only cross-cluster visibility scoring to the algorithm (Wojcik et al., 2022). Conversely, where AI does act autonomously on a public, it is frequently not sophisticated AI: Discord’s AutoMod is predominantly regular expressions (Discord, n.d.); ClueBot NG reverts Wikipedia vandalism with a 2010s-era neural network.

The pure-AI products sell scores. The systems that take action run on rules and classical methods.

5.3 Audience

Of the 58 AI-constitutive tools, 23 are addressed to citizens; the remainder serve platform moderators (9), institutions (8), journalists (6), researchers (5), civil servants (5), and market research (2). The state’s own filings document the asymmetry: the UK’s Consult records that “no citizens interact with the tool” (i.AI, 2025); Parlex is IP-whitelisted to government networks. The same capability is sold commercially to lobbyists (Quorum, FiscalNote, BillTrack50). After a liveness filter, citizen-facing civic-learning AI amounts to two working tools: wahl.chat, built by students (wahl.chat, n.d.), and CivicChats, built by a university.

5.4 Supporting results

Two ecosystems. Information integrity/governance is 59% AI-constitutive; deliberation/participation 36% — a structural 1.6× difference: deliberation tools help people talk over a human process that functions without them, while information-integrity tools face a volume problem where machine learning is load-bearing by construction. Yet information integrity is still only 33% AI-deciding: the two ecosystems differ in how much they need the model and agree in refusing to let it judge.

Lifecycle coverage. Mapped onto an eight-stage discourse lifecycle, AI-constitutive tools cluster in collective sensemaking, structured deliberation, consensus discovery and moderation — stages that consist of processing large volumes of synchronous text, the native affordance of current models and the shape of the adjacent commercial problems (market research, trust and safety). Collective decision (stage 7) is empty of AI and full of civic software; reflection and institutional learning (stage 8) is empty entirely.

Refusal. Three systems have deliberately declined AI and published their reasoning — mySociety, Citizen OS, Kialo (mySociety, 2025). Existing codebooks cannot distinguish considered refusal from non-adoption; we propose ABSENT-BY-REFUSAL as a coding category.

6. Discussion

6.1 Govern authority, not architecture — wherever authority sits

The dominant anxiety in AI-and-democracy research is that algorithmic systems will accrue democratic authority. The data suggests the authority has largely not been delegated: restraint is the norm, repeated, independent, and documented. The regularity is striking — the more a tool claims to adjudicate truth or legitimate speech, the further its designers push the model away from the decision and toward the pipeline.

Two consequences cut in opposite directions. The reassuring one: practitioners have been considerably more careful than critics assume. The uncomfortable one: where AI does act autonomously on a public, it tends to run on classical methods that attract no governance attention — matrix factorisation decides what Community Notes shows to millions; regular expressions decide what Discord blocks. Governance scrutiny is calibrated to the sophistication of the model rather than to the consequence of the decision. Scrutiny should scale with what the system is permitted to decide, not with how modern its model is.

Analysis of the infrastructure layer beneath these tools (OpenHaven, 2026) forces one refinement: architecture is not the opposite of authority but one of the places it hides. Where a protocol implements audit trails, identity, or federation determines who can verify, who can exclude, and who can leave. Hence the fuller principle: govern authority wherever it sits — in the model, in the human loop, or in the architecture. Two coding variables follow for future passes: steward governance model and dependency-stack capture risk — a tool can keep AI advisory and humans deciding and still be one acquisition or funding cliff away from vanishing.

6.2 A demand-side deficit

The audience asymmetry is not explained by technical difficulty — government, lobbyists and platforms procure the same capabilities citizens lack. It is explained by procurement: every actor in the democratic process with a budget has bought AI; the public has no budget, no procurement function, and no vendor. Democratic AI does not have a supply problem; it has a demand problem. The production gap compounds it: every purpose-built democratic AI tool in the inventory is a prototype, pilot, or unmaintained, while the production-grade systems were built to sell market research or trust-and-safety services — a capital-structure problem, not a maturity problem that time will fix (OpenHaven, 2026).

6.3 Why the field misperceives its own ecosystem

Four mechanisms produce the gap between literature and landscape: algorithmic ambiguity (no explicit AI definition), feature accretion (platforms recategorised wholesale after adding a module), category drift (programmes and prize competitions counted as tools), and autonomy inflation (presence mistaken for authority). None is dishonest; together they produce a literature that systematically overstates how much democratic AI is running and what it is permitted to do. Our own 18% first-pass error rate shows the misperception is easy to acquire and hard to notice.

6.4 Research agenda

(1) Complete the census and report attrition. (2) Measure the decides/supports variable systematically, at scale. (3) Test the affordance hypothesis: does the capability distribution track what models make easy rather than what democracies need? (4) Study documented AI refusal. (5) Build the citizen side — civic learning and institutional memory are tractable, high-value, and unbuilt. (6) Interrogate the consensus monoculture: nothing in the inventory is designed to help a minority sustain a disagreement well. (7) Investigate the demand side, including the legitimacy “AI penalty.”

7. Conclusion

This study set out to map an ecosystem and found that the ecosystem is, in substantial part, a literature. The systems most often cited as evidence that AI is transforming democratic discourse are frequently not AI; where AI is genuinely present it overwhelmingly advises rather than decides — a restraint its builders chose and documented, and which governance discourse has not registered; and the systems that exist are built for the state, platforms, newsrooms and lobbyists, while the public makes do with what volunteers and students build for free. None of this argues against democratic AI. It argues that democratic AI is a project not yet begun rather than a landscape to be surveyed — and that the field’s first task is to stop mistaking its citations for an ecosystem, and its anxieties for a distribution of power that nobody has actually handed over.


References

Journal articles and books

  • Argyle, L. P., Bail, C. A., Busby, E. C., Gubler, J. R., Howe, T., Rytting, C., Sorensen, T., & Wingate, D. (2023). Leveraging AI for democratic discourse: Chat interventions can improve online political conversations at scale. PNAS, 120(41). https://doi.org/10.1073/pnas.2311627120
  • Barandiaran, X. E., Calleja-López, A., Monterde, A., & Romero, C. (2024). Decidim, a Technopolitical Network for Participatory Democracy. Springer (open access). https://library.oapen.org/handle/20.500.12657/87634
  • Borge, R., et al. (2026). Alternative digital platforms and the renewal of the public sphere: Decidim and the democratic governance of participatory infrastructures. Social Sciences, 15(3), 166. https://www.mdpi.com/2076-0760/15/3/166
  • Costello, T. H., Pennycook, G., & Rand, D. G. (2024). Durably reducing conspiracy beliefs through dialogues with AI. Science, 385(6714). https://doi.org/10.1126/science.adq1814 — see also Editorial Expression of Concern (2026), https://doi.org/10.1126/science.aej2383
  • Fishkin, J. S. (2018). Democracy When the People Are Thinking. Oxford University Press.
  • Habermas, J. (1996). Between Facts and Norms. MIT Press.
  • Halfaker, A., & Geiger, R. S. (2020). ORES: Lowering barriers with flexible machine-learning-backed moderation tools for Wikimedia. Proceedings of the ACM on Human-Computer Interaction, 4(CSCW2). arXiv:1909.05189
  • Hassan, N., Arslan, F., Li, C., & Tremayne, M. (2017). Toward automated fact-checking: Detecting check-worthy factual claims by ClaimBuster. Proceedings of KDD ’17.
  • Landemore, H. (2020). Open Democracy: Reinventing Popular Rule for the Twenty-First Century. Princeton University Press.
  • Small, C., Bjorkegren, M., Erkkilä, T., Shaw, L., & Megill, C. (2021). Polis: Scaling deliberation by mapping high dimensional opinion spaces. Recerca: Revista de Pensament i Anàlisi, 26(2). https://doi.org/10.6035/recerca.5516
  • Tessler, M. H., Bakker, M. A., Jarrett, D., Sheahan, H., Chadwick, M. J., Koster, R., Evans, G., Campbell-Gillingham, L., Collins, T., Parkes, D. C., Botvinick, M., & Summerfield, C. (2024). AI can help humans find common ground in democratic deliberation. Science, 386(6719). https://doi.org/10.1126/science.adq2852
  • Wojcik, S., Hilgard, S., Judd, N., Mocanu, D., Ragain, S., Hunzaker, M. B. F., Coleman, K., & Baxter, J. (2022). Birdwatch: Crowd wisdom and bridging algorithms can inform understanding and reduce the spread of misinformation. arXiv:2210.15723. https://arxiv.org/abs/2210.15723

Reports and grey literature

Primary system documentation (accessed July 2026)

Data

  • Project dataset: catalog_verified_v2.csv (158 entries, seven coding passes, per-row evidence fields); codebook in Appendix B of the project manuscript.

Note: in-text citations of the form (Author, year) refer to this list. Some system-documentation URLs are entry points rather than deep links; per-row deep links are recorded in the dataset’s evidence fields.

Scroll to Top