OSWorld-Science Benchmark for Computer-Use Agents in Scientific Workflows
October 1, 2026
OSWorld-Science evaluates 12 vision-language models across 146 tasks involving specialized scientific software like molecular drawing and physical simulation. The benchmark utilizes execution-based evaluators to verify results within complex, multi-step scientific research workflows.
HOW THIS AFFECTS YOU
●
builderThis provides a framework for testing agents intended for automated lab or research workflows.
●
researcherYou can use this to test how well VLMs handle specialized scientific UI and object manipulation.