PetQA introduces a benchmark for evaluating veterinary knowledge and clinical reasoning in large language models (LLMs) and vision-language models (LVLMs). It includes 10,076 text-only and 8,751 multimodal QA pairs based on real-world questions about dogs and cats, with answers from expert veterinarians.
The test split, PetQA-Bench, features annotations for question types and clinical conditions, aiding detailed evaluation. Eighteen models were assessed using metrics such as ROUGE, BERTScore, and LLM-as-a-judge, across zero-shot, retrieval-augmented generation (RAG), and supervised fine-tuning (SFT) settings.
Results highlight the current strengths and limitations of models in veterinary clinical queries, emphasizing the need for better adaptation methods to develop reliable AI systems for veterinary care. Translated versions in five languages are also provided to facilitate broader use.
Source: https://arxiv.org/abs/2609.04598