●builderThis highlights a risk in training multimodal models where agents learn to 'hallucinate' correct answers based on luck rather than vision.
●researcherYou can use counterfactual rollouts to better evaluate whether multimodal agents are actually grounding claims in visual evidence.