The study investigated how large language models respond to input data that contradicts their perceived understanding. Researchers generated text in English, Czech, Slovak, and Upper Sorbian using RDF triples containing local Czech and Slovak data. The data was categorized as factual (FA), counterfactual (CFA), and fictional (FI). The goal was to assess the impact of context-memory conflict – the degree to which an LLM’s internal knowledge clashes with the provided input.
Surprisingly, the analysis found only a weak context-memory conflict on a human-annotated sample. Kimi K3, an LLM judge, aligned well with human annotations on this sample. Counterfactual inputs received slightly lower faithfulness scores than factual ones, with a difference of -0.05 on a 1-5 scale. This indicates a relatively tolerant response from the model to contradictory information.
Furthermore, the research highlighted that selecting an inappropriate LLM judge could lead to an overestimation of the context-memory conflict. This suggests that the evaluation process itself can influence the observed discrepancies. The findings have implications for the design and deployment of retrieval-augmented generation and data-to-text systems, where LLMs are increasingly relied upon for factual accuracy.