StarDojo: Benchmarking Open-Ended Behaviors of Agentic Multimodal LLMs in Production–Living Simulations with Stardew Valley
Abstract
Autonomous agents navigating human society must masterboth production activities and social interactions, yet existing bench-marks rarely evaluate these skills simultaneously. To bridge this gap,we introduce StarDojo, a novel benchmark based on Stardew Valley,designed to assess AI agents in open-ended production–living simulations.In StarDojo, agents are tasked to perform essential livelihood activitiessuch as farming and crafting, while simultaneously engaging in socialinteractions to establish relationships within a vibrant community. Star-Dojo features 1,000 meticulously curated tasks across five key domains:farming, crafting, exploration, combat, and social interactions. Addition-ally, we provide a compact subset of 100 representative tasks for efficientmodel evaluation. The benchmark offers a unified, user-friendly interfacethat eliminates the need for keyboard and mouse control, supports allmajor operating systems, and enables the parallel execution of multipleenvironment instances, making it particularly well-suited for evaluatingthe most capable foundation agents, powered by multimodal large lan-guage models (MLLMs). Extensive evaluations of state-of-the-art MLLMsagents demonstrate substantial limitations, with the best-performingmodel, GPT-4.1, achieving only a 12.7% success rate, primarily due tochallenges in visual understanding, multimodal reasoning and low-levelmanipulation. As a user-friendly environment and benchmark, StarDojoaims to facilitate further research towards robust, open-ended agents incomplex production-living environments.