Skip to content

LLMs1 min read

PetQA benchmark evaluates veterinary knowledge in language and vision models

PetQA is a Korean QA benchmark with over 10,000 text and multimodal pairs, assessing veterinary knowledge and clinical reasoning in LLMs and LVLMs across various settings.

By OpenSmartRoute editorial · written through the router by llm-onprem

From arXiv cs.CL - “PetQA: Benchmarking Veterinary Knowledge and Clinical Reasoning

PetQA introduces a benchmark for evaluating veterinary knowledge and clinical reasoning in large language models (LLMs) and vision-language models (LVLMs). It includes 10,076 text-only and 8,751 multimodal QA pairs based on real-world questions about dogs and cats, with answers from expert veterinarians.

The test split, PetQA-Bench, features annotations for question types and clinical conditions, aiding detailed evaluation. Eighteen models were assessed using metrics such as ROUGE, BERTScore, and LLM-as-a-judge, across zero-shot, retrieval-augmented generation (RAG), and supervised fine-tuning (SFT) settings.

Results highlight the current strengths and limitations of models in veterinary clinical queries, emphasizing the need for better adaptation methods to develop reliable AI systems for veterinary care. Translated versions in five languages are also provided to facilitate broader use.

Source: https://arxiv.org/abs/2609.04598

Published Sep 7, 2026 · updated Sep 7, 2026 · 121 words

Keep reading

Related posts

More in LLMs