Benchmarking Embodied Fluid Intelligence

The videos below show rollouts from different LLM agents as they attempt to solve the Flint dataset. Every model attempted the same scenes, three for each task, and each row below is one scene, so the models can be compared on exactly the same room.

Click the videos to see a larger play through. They show the agent's egocentric view, a third-person view of the room, and an LLM-written summary of the agent's reasoning.