← Back to feed
2026-06-24datacommunity code

When Certainty Is an Artifact: Keyword Lexicon Blindness and the (Mis)Measurement of Rhetorical Stance

Bo Chen

PDF preview for When Certainty Is an Artifact: Keyword Lexicon Blindness and the (Mis)Measurement of Rhetorical Stance
Read on arXiv →

Key claim

Keyword counts can misrepresent psychological insights in discourse.

In plain English

Imagine you're trying to understand how public figures express their emotions through language. Researchers often rely on keyword counting to gauge certainty or negativity in their speech. However, this method can be misleading. For instance, a phrase like 'never absolutely totally confident' might be counted as high certainty, even though it conveys the opposite. This is where the problem lies: keyword lexicons can misinterpret the actual meaning of what speakers are saying. This misinterpretation is due to several failure modes, such as syntactic blindness, where the structure of the sentence is ignored, and polysemy blindness, where words with multiple meanings are misclassified. These issues can lead to incorrect conclusions about a speaker's psychological state based on flawed measurements. To address this, the authors propose using a more sophisticated approach that leverages large language models (LLMs) for semantic classification instead of simple keyword counting. This method reveals a more accurate picture of the speakers' emotional states and rhetorical stances. The key takeaway is that relying on keyword counts can lead to significant errors in understanding discourse, and using LLMs can provide a clearer, more nuanced view of language use.

Novelty
7.5/10

The paper introduces a new perspective on the limitations of keyword-based scoring in social science research.

Reliability
8.0/10

The findings are supported by a robust analysis of multiple speakers and a clear demonstration of the limitations of existing methods.

Deep reliability assessment

The methodology supports the narrower claim that, on this interview corpus, keyword lexicons can produce large negative-affect/emphatic-certainty correlations that collapse or reverse under context-aware LLM classification. The paper overclaims when it treats the LLM labels as the "true" rhetorical pattern, because there is no human-labeled validation set, no independent diarization audit, and only a small cross-model robustness sample.

Reproducibility

No current reproducibility package is available. The paper says all source code, processed data, and interactive visualizations will be made public upon formal acceptance, but no repository or dataset URL is provided.

Key figure

No Figure 1 is included in the provided excerpt; the key methodological diagram would be a dual-instrument pipeline comparing keyword lexicon scoring against full-corpus LLM semantic classification on diarized interview sentences.

Benchmark results

Ray Dalio interview corpus, 11,361 diarized sentencesPearson r(negative, emphatic): 0.206vs Keyword lexicon scoring, r=0.851-0.645
Cathie Wood interview corpus, 5,999 diarized sentencesPearson r(negative, emphatic): -0.514vs Keyword lexicon scoring, reported cross-speaker range r=0.72-0.93speaker-specific delta not reported
Kenneth Rogoff interview corpus, 5,796 diarized sentencesPearson r(negative, hedged): 0.875vs Keyword lexicon scoring emphasized negative-affect/emphatic-certainty coupling insteadnot directly comparable
Peter Zeihan interview corpus, 9,469 diarized sentencesPearson r(negative, hedged): 0.722vs Keyword lexicon scoring emphasized negative-affect/emphatic-certainty coupling insteadnot directly comparable
GitHub1 repo
Ubssk/260623seminarCommunity