Existing AI agents struggle with the stylistic fidelity of human-authored academic charts. Current approaches suffer from sparse reward signals, which severely undermine reinforcement learning (RL) when applied to visual reasoning tasks.
ViCo introduces a training framework that employs iterative reflections to align generated chart images progressively with reference examples. The system augments Monte Carlo Tree Search with consistency-based pruning to synthesize high-quality reflection trajectories before applying a multi-step RL algorithm.
Source: https://arxiv.org/abs/2609.16014



