Search papers, labs, and topics across Lattice.
This study investigates the sources of instability in stance analysis of public discourse by comparing two speaker-diarization pipelines and two measurement methods鈥擫LM annotation and keyword lexicon鈥攁cross 256 YouTube interviews. The findings reveal that while preprocessing pipeline sensitivity is significant for speakers with limited samples, the disagreement between measurement methods is more pronounced, often leading to opposite conclusions about affective valence and epistemic modality. Importantly, aggregate valence proportions remain stable across methods, highlighting the need for researchers to rigorously assess the robustness of their findings against both preprocessing and measurement variations.
Discrepancies between LLM annotations and keyword lexicons can lead to fundamentally different interpretations of public discourse, especially for well-sampled speakers.
Computational social science increasingly relies on automated preprocessing pipelines -- speaker diarization, ASR transcript cleaning, sentence segmentation -- to convert raw media into analyzable text. When these pipelines produce different outputs from the same input, two distinct sources of instability can arise: the preprocessing pipeline itself (diarization method, segmentation rules) and the downstream measurement instrument (LLM annotation vs.\ keyword lexicon). Using 256 YouTube interviews across 41 public figures from five domains, we compare two speaker-diarization pipelines and two measurement methods, all targeting the coupling between affective valence and epistemic modality. We find that (1) preprocessing pipeline sensitivity is concentrated in speakers with limited video samples (N $\leq 5$); for the four best-sampled speakers (N $\geq 16$), the mean absolute pipeline-induced change in $r(\text{neg}, \text{emph})$ is only $0.13$; (2) cross-method disagreement is larger and more systematic -- the LLM and keyword-lexicon methods assign opposite coupling directions to several well-sampled speakers, even within the same preprocessing pipeline; and (3) aggregate valence proportions are highly stable ($|螖p(\text{neg})| < 6$pp) regardless of pipeline or method, masking both sources of instability. The contribution is a diagnostic framework that separates pipeline effects from measurement effects: researchers studying cross-dimensional relationships in interview data should verify that their conclusions are robust to both sources of variation, with particular attention to measurement method choice.