RhetoricMap

See how a debate splits between attacking and explaining.

Paste a transcript. RhetoricMap marks every sentence as an attack on an opponent, a case for policy, both, or neither — and shows you the words that decided each call.

Attack on an opponent Case for policy Both at once Neither
An illustrative exchange, coded live by the tool on this page. Tags on the right show which cues decided each call.

Analyze a transcript

Works on any transcript where speakers are marked like NAME: at the start of a line. Presidential debates, parliamentary sessions, school debate rounds. Nothing leaves your browser.

Comma-separated. Helps the classifier spot who is being targeted. 2020 primary candidates are already known.

Pick a sample above or paste your own transcript, then choose Analyze.

What the research found

RhetoricMap grew out of a study of the 2020 Democratic presidential primary, where the party tied debate eligibility to a campaign's unique donor count — 65,000 donors for the first debate, 225,000 by January 2020. Speech converted into money on a clock, which makes the cycle an unusually legible case.

$2Min 24 hours, from 63,277 donors

The fastest documented response in the study followed an attack. After Kamala Harris challenged Joe Biden's record on school busing in June 2019, her campaign reported that sum overnight, 58% of it from first-time givers averaging about $30.

+257%quarter over quarter

But an attack was not the only thing that moved money fast. Andrew Yang's basic-income pilot announcement — a policy proposal, not a criticism — preceded a $10M quarter in which 99% of gifts fell under $200.

Third partycaptured the moment

When Elizabeth Warren attacked Pete Buttigieg over a high-dollar fundraiser, the clearest beneficiary was neither of them. Bernie Sanders surrogates converted the moment into their own donation appeal within the hour.

509,000donors without a spike

Warren's own base grew from $6M to $24.6M per quarter with no single debate credited as the driver. A large share of small-dollar giving was never organized around debate moments at all.

Taken together: the system does not reward attacks over substance so much as it rewards whichever content travels fastest. That is a harder problem, and a measurable one — which is why this tool exists.

How the classifier decides

No model, no black box. Every sentence is scored against the cue list below, and the tool shows you which cues fired. The rules printed here are read directly from the code that runs — if the list changes, this page changes with it.

One cue is data, not pattern. The name cue is built at run time from the list below plus anything you type into Opponent names. Metacharacters are escaped and entries sorted longest-first, so a multi-word name wins over its own parts. The cue is suppressed when the matched name also appears in the speaker's own name, because a speaker referring to themselves is not attacking anyone.

    A sentence scoring on both sides is marked both at once. A sentence scoring on neither — procedural remarks, biography, thanking the moderator — is marked neither, because real transcripts are full of it and pretending otherwise would inflate every other number.

    Classification follows the communicative function of a statement, not its tone or its truth. A calm, accurate criticism of a rival's record counts as an attack. An angry demand for a policy does not.

    What this tool does not do

    Stated plainly, because a measurement tool that hides its error bars is worse than no tool.

    It does not check whether anything is true

    An accurate criticism and a fabricated one are coded identically. RhetoricMap measures what kind of move a speaker is making, not whether the move is honest.

    It does not prove that rhetoric caused anything

    The research behind it reports associations across five well-documented cases, not causal effects. A spike after a debate may reflect the coverage the moment generated, the candidate's existing trajectory, or a fundraising email sent the same night.

    Rule-based coding misses what it was not told about

    Sarcasm, implication, and an attack delivered without a single negative word will slip through as something else. The cue list is transparent precisely so you can see where it will fail.

    No second coder has checked it

    Published content analysis reports inter-coder reliability: two people coding the same text independently and measuring agreement. This tool has not been through that yet. Until it has, treat its output as a structured first pass, not a finding.

    Where it goes next

    The tool exists because the research hit a wall: hand-coding a debate is slow enough that one person can only do a handful of exchanges. These are the steps from here, in order of what would make the output trustworthy soonest.

    1. Code all eleven 2020 debates end to end

      The study examined the moments that news coverage flagged. Running every sanctioned debate in full removes the bias toward exchanges that happened to go viral.

    2. Measure agreement against human coders

      Two independent readers code the same sample, and the tool is scored against them. Without that number, nobody should cite this output.

    3. Join the coding to itemized contribution records

      Federal filings and ActBlue transactions are public. Reconciled to the candidate-day level, they would turn a five-case comparison into something a regression could actually test.

    4. Publish the coded corpus openly

      The most useful output may not be the tool but the dataset: a coded, timestamped record of eleven debates that other people can check and argue with.