The research investigated the validity of financial sentiment analysis tools. It tested the assumption that human-labelled sentiment data and model-extracted market signals measure the same phenomena. The study utilized a corpus of securities class actions spanning from 2002 to 2025, containing 70,500 X messages and a single-annotator human labelled gold sample. Five models – VADER, Loughran-McDonald, FinBERT, Twitter-RoBERTa, and an LLM annotator – were processed through a common pipeline.
The findings indicated that the relationship between sentiment construct and predictive validity is influenced by the sampling convention and score representation. Under conventional method-specific sampling, human agreement aligned more closely with same-day associations than with one-day leads. However, on a fixed-n panel, agreement exhibited similar graded rank correlations at both horizons, while coarse ordering remained weak.
Furthermore, the study observed that a conversation containing 17.6% spam did not demonstrate message volume’s ability to predict market damage or settlement size. This suggests that raw message volume may not be a reliable indicator of financial risk.
The research highlights the importance of considering evaluation timeframes when assessing the performance of financial sentiment models. It suggests that a nuanced understanding of the data and sampling methods is crucial for determining predictive rankings.
Source: https://arxiv.org/abs/2609.11144