AI Accountability in Politics Requires Continuous Monitoring
Published
Modified
Deepfakes are only the visible layer of political AI risk Chatbots increasingly shape how voters receive political information Continuous monitoring makes changing AI behaviour visible and accountable

A conservative political action committee named Citizens for Sanity released a series of AI-generated videos in June 2026 mocking Texas Senate candidate James Talarico dressed as a governess and singing a parody of a show tune about transgender children, and by the time PolitiFact measured the campaign's reach later that month, the ads had garnered about 1.4 million impressions on Facebook alone, according to Meta's own advertising data. No one who watched the clip mistook it for an actual recording. The trick was the message, meant to be communicated as a spectacle rather than to be believed as a fact. This distinction, between content created to deceive and content created to be identified as fake and still doing political work, has absorbed most of the attention paid to AI and elections this cycle, while a quieter and much larger change, in how tens of millions of voters get their political information from chatbots in the first place, has prompted much less scrutiny. This second level is where AI accountability in politics needs to begin.
Deepfakes Are the Visible Layer of AI Political Risk
The Talarico ad wasn't an isolated trick. Citizens for Sanity posted six AI-generated videos over several days mocking the Democratic Senate candidate, a campaign that PolitiFact journalist Loreben Tuquero spotted in the Meta ad library and found running alongside a Texas law that bans political deepfakes within thirty days of the election and treats violations as a Class A misdemeanor. Adrian Shelley, director of the advocacy group Public Citizen in Texas, told Tuquero that the tools had crossed a line: campaigns no longer needed a movie studio to depict an opponent doing or saying anything, just one subscription and an afternoon. In the race for governor of Massachusetts that same year, Republican Brian Shortsleeve posted a fabricated audio clip mimicking Democrat Maura Healey, a revelation that did nothing to slow down its online circulation.
Two years before the Talarico ad aired, during the 2024 New Hampshire primary, an AI-generated imitation of President Joe Biden’s voice told voters to save their ballots for November. The Federal Communications Commission fined Kramer six million dollars in September 2024, and Loyaan Egal, who heads the law enforcement office, said the incident threatened the foundations of democratic elections. New Hampshire prosecutors separately charged him with felony voter suppression. What Talarico's ads, Shortsleeve clip, and Kramer's Robocall share is not partisanship but a cost curve: synthetic political content is now cheap enough for any campaign, super PAC, or lone agent to produce, which is precisely why the ability to reliably vouch for what's real has become rare and valuable.
In July 2025, someone using an AI-cloned voice of Secretary of State Marco Rubio contacted three foreign ministers, a U.S. governor and a member of Congress through the messaging app Signal, according to a State Department cable quoted by Reuters correspondent Humeyra Pamuk. The telegram estimated that the impersonator was likely fishing for access to information or accounts rather than playing a hoax, and a senior State Department official confirmed that an investigation was underway. The episode had nothing to do with the campaign ad, yet it belongs to the same account as Talarico's parody song and Kramer's robocall, as it shows how quickly voice cloning spread from campaign attack ads in an attempt to manipulate the government's own machinery.
Chatbots Are Becoming a Political Information Layer
Visible fakes are a small share of the exposure most voters now have to AI in a political context. A March 2026 survey by America's Political Pulse found that forty-six percent of Americans used AI to receive news at least occasionally, and thirty-nine percent said they turned to it specifically to understand politics. A separate Data for Progress poll conducted in February found that thirty-nine percent of U.S. voters are likely to use a chatbot to learn about a candidate or voting measure. None of these activities involve fabricated video or cloned voice. It's like millions of ordinary people asking an honest question to ChatGPT, Gemini, or Copilot and trusting whatever comes back.

Whether this trust is well-placed remains an open empirical question, and the evidence gathered so far is neither reassuring nor damning. A 2025 study published in Nature Communications found that political messages written by large language models persuaded readers about as effectively as messages written by ordinary people, with participants crediting the use of facts, logical structure, and the cool tone of the AI-generated text. A companion work published in Science tested nineteen models on more than seven hundred political issues and came to a similar conclusion, adding a less comfortable finding: additional training and direct engineering made the models measurably more convincing while making them measurably less accurate, a trade-off that no vendor advertises.
Lawrence Norden, who leads election work at the Brennan Center for Justice, called the 2026 midterm elections the first chatbot election, noting that the share of Americans using AI assistants has nearly doubled since 2024 and that nearly a quarter now use one daily. His team fed several top chatbots a range of election conspiracy theories before the vote and found, to their surprise, that each model challenged every false claim tested, even after repeated prompting. This result is really encouraging, and it's also exactly the kind of finding that can't be assumed to be true, as nothing about how a chatbot answered a question in August guarantees the same answer in October, after a system prompt has changed or a detail pass has shifted its guardrails without anyone outside the company knowing it happened.
Why One-Off AI Audits Miss Changing Model Behavior
The Carnegie analysis describes the problem through the acronym DEEP: chatbots are Dynamic, Ephemeral, Embedded and Personalized. Chatbots are dynamic, as their responses change with hidden details, modified system prompts, or live web access that a user never sees. They're ephemeral, because an AI overview or chatbot response usually disappears the moment a browser tab is closed, leaving no file for a journalist or regulator to inspect later. They're built into software layers of guardrails that vary by device and interface, even when the underlying model is nominally the same, and they're increasingly personalized, shaped by memory features that OpenAI introduced in 2024 and that Anthropic, Google, and Microsoft have each added since then. Overall, these features mean that a single control captures a frame of a movie that is still playing.
The underlying accuracy problem gives the question of time its stakes. NewsGuard's August 2025 audit of ten top chatbots found that when a fake news claim was presented, models confirmed it as true thirty-five percent of the time. Earlier that year, the British Broadcasting Corporation provided its own reports on ChatGPT, Copilot, Gemini, and Perplexity, asked each system questions about the news, and found that nineteen percent of the responses introduced factual errors that the BBC's own journalism did not contain. Being wrong nineteen or thirty-five percent of the time in ordinary news is obviously no more reliable when the issue turns to a ballot measure, and there is no permanent mechanism to check the number again the following month.
The AI Watchman monitoring project runs the same politically sensitive questions across models each week and has already identified changes that would otherwise not have been recorded. In August 2025, as Israeli operations in Gaza intensified, the project recorded a jump in GPT-4.1's denial rate to Israel-related questions. The following month, while the Texas legislature was debating a bill restricting access to abortion medications that eventually passed, the team measured a sharp increase in GPT-5's denial rate on abortion-related prompts, which fell to the original value within a week. No finding proves that a company intentionally tuned a model in response to current events, and researchers say so clearly, but each is the kind of pattern that an audit would never show at a point in time. This loophole is the practical argument for AI accountability in politics as an ongoing discipline rather than an occasional press release.

AI Companies Apply Different Political-Content Rules
Every major AI developer insists that they do not introduce their own policy into their product, and each has begun publishing evidence to support the claim, although the evidence inevitably reflects a standard chosen by the company itself. Anthropic published an update on its election safeguards on April 24, 2026, ahead of the midterm elections, stating that two of the Claude models scored ninety-five and ninety-six percent on an internal measure of impartiality and responded appropriately to election-related prompts almost every time they were tested. OpenAI's Model Behavior team, released a five-part bias framework last October and reported that newer GPT-5 models showed about thirty percent less measurable political bias than previous versions, with less than one in ten thousand live responses showing any sign of bias. Both companies deserve credit for their publishing methodology rather than a marketing line. Both disclosures are useful, but they remain company-designed and company-reported evaluations
Meta opted for a different framework, saying it wants Llama models to articulate both sides of a controversial issue rather than adopting one, while Google initially narrowed down Gemini's election-related responses in 2024 before shifting to a similar approach to both sides, according to a Washington Post analysis published in mid-2026. An evaluation by Future of Free Speech, a research center at Vanderbilt University that Anthropic itself has cited as an external reviewer, ranked xAI's Grok 4 as the strongest large language model tested for free expression purposes, an assessment released shortly after a previous version of the same chatbot had generated antisemitic content in a widely reported incident at the time. Free expression and harm mitigation are both legitimate goals, and nothing in today's landscape forces any company to reconcile them in the same way, which is the point: these are political and ethical judgments made by product groups, not neutral engineering outcomes.
The fabricated Talarico ad and Rubio's voice clone are, in a narrow sense, the easiest problems, since once discovered they can be labeled, debunked, or prosecuted, as Kramer's indictment showed. A model's decision about what sources to show up or how much weight to give to one side of a questionable policy question is harder to catch in the act, because nothing about it looks like a scam. It looks like an ordinary answer given in the same confident tone that the model uses for everything else, and unless someone asks the same question every week and keeps a record, a turning point in how a chatbot handles abortion, immigration, or an election result can go unnoticed until a pattern is already well established.
What AI Accountability in Politics Requires
The treatment proposed by Metaxa and Engler is not another one-time audit, but permanent infrastructure: research groups, charitable funders, and universities pooling resources to ask important questions on a fixed set of policy issues each week, indefinitely, and maintaining results in the way that social scientists have long maintained panel data on family income or labor markets. Congress has not funded anything resembling this infrastructure, and the Brennan Center's own monitoring of state legislatures found many proposed bills involving AI and elections, but few that had become law by mid-2026. Monitoring creates evidence, while penalties require clear legal obligations and enforcement. Kramer's fine and Texas' own deepfake statute show that penalties already exist for the most blatant violations, and the most difficult task now is to extend the same willingness to punish into the much larger space of ordinary, seemingly scam responses that never look like a scam at all.
The obvious objection is that government-backed control of what chatbots say about politics risks becoming the kind of speech control that the First Amendment exists to prevent, and state lawmakers have mostly avoided outright bans on AI-generated political content for this reason, opting instead for disclosure rules. This caution is reasonable, but it answers a different question than what continuous monitoring poses. Requiring a campaign to flag a synthetic ad, or requiring a company to let outside researchers see how its model answered the same question in March and again in September, doesn't dictate what the model or ad might say. It makes the pattern visible to the public affected by it, closer to a financial disclosure requirement than a speech restriction, and the Campaign Finance Act, where sunlight rather than prohibition has long been the preferred treatment for the risk that money buys undue influence, offers a working standard for how AI oversight could proceed without causing the same objection.
The European Union has already gone further than the United States on this front. Under the AI Act's code of practice, developers of larger models must mitigate the potential for manipulation of their systems, a category that can include political outcomes and propaganda production, giving Brussels a regulatory hook that Washington is currently lacking. State legislators drafting their own AI and election bills, campaign finance regulators who have so far focused only on paid political advertising, and AI companies themselves, several of which already publish internal bias assessments, all have a standard close by: publishing methodology, allowing external replication, and undergoing a surveillance regime that runs continuously rather than once. None of this requires resolving the deepest disagreement about what political neutrality should mean, a debate that Anthropic, OpenAI, Google, and Meta have publicly had with different answers and no obvious way to judge each other.
No one needed a State Department telegram to acknowledge that a synthetic Marco Rubio calling a Secretary of State on Signal was a problem, and no one needed a study to see that Citizens for Sanity wanted their Talarico ad to be shared as a spectacle rather than confused with fact. These cases resolve themselves, eventually, through prosecutors, fact-checkers, and platforms that pull content. The one that is resolved more slowly, if at all, is the question posed by the forty-six percent of Americans who are now asking a chatbot instead of a search engine or newspaper for the people asking for their vote, as no one keeps a record of what these systems told them last month, let alone if the answer would be the same today. Until this record exists, the industry's own scores are the only items anyone has and were rated by the party tested.
This article reflects the analytical judgment of the author and does not constitute policy advice or the official position of any affiliated institution.