MonoTM addresses limitations in traditional topic modeling by decoupling mixture estimation from semantic interpretation. The framework utilizes sparse autoencoders (SAEs) to extract interpretable features from dense representations. It estimates topic mixtures from the full SAE bag-of-features representation, then learns topic descriptors using a separate vocabulary of corpus-grounded semantic features. This approach aims to represent topics with more meaningful semantic units than individual words, improving their utility for downstream analysis. Across three benchmark corpora, the framework favors specific SAE configurations and feature subsets for document--topic mixture estimation and semantic interpretation. This design preserves global topic structure while representing topics with semantic units more meaningful than individual words, making them more useful for corpus analysis.
Source: https://arxiv.org/abs/2609.09575
