TrustSwap Reveals Source-Trust Shortcuts in Fact-Checking RL Agents
September 30, 2026
Testing shows that RL-trained fact-checkers often rely on source reliability labels rather than evidence content, with label changes altering up to 50% of verdicts. Standard GRPO fine-tuning worsens this shortcut in 8B models, though trust-swap augmentation (TSA) is proposed as a mitigation.
HOW THIS AFFECTS YOU
●
researcherYou should evaluate if your RL agents are truly reasoning over content or just exploiting metadata shortcuts.
●
policyThis highlights a reliability risk where models may provide incorrect verdicts based solely on source labels.