IROS 2026: Notes From the Field

IROS 2026: Notes From the Field

Robots and World Models

I skipped a post yesterday to get some extra sleep. Yes, sleep is actually more effective than coffee in promoting long-term concentration!

I have attended enough graduate-student lectures and paper presentations to get a sense of what currently plagues the robotics research community. There seems to be a focus primarily on implementing software to make newly commercialized robotic hardware functional in the real world. For hardware-focused presentations, the emphasis is more on developing unique morphologies: soft, aerial, snake-like, underwater.

This was clear when I roamed the exhibit hall, looked at the competition section, and compared the robots there to those demonstrated in the vendor section. The competitions were quite varied (see below for video), ranging from intensely difficult obstacle courses that would challenge a human to electric RC cars on a small race course. NIST (National Institute of Standards and Technology) staff judged the obstacle courses, which were designed to focus on particular types of difficult terrain and were all constructed from timber. One featured a short segment of stairs, with standard and large (about 2' tall) steps; another with rotating discs mounted at different angles, and a mesh of 2 x 4 planks mounted about 6" apart to require the legged robot to position itself on the narrow edge of the wood to avoid dropping down to the floor below.

The race cars were more focused on completing the course in the minimum time. Human drivers also competed to register their best time. Another competition tested which robot could complete an assembly task in the least time, with the catchy headline, "Can a Robot Build IKEA Furniture?"

Take a few steps to the other side of the hall to visit the vendor exhibit, and you see humanoid robots or their related accessories and software in virtually every exhibit. The humanoid robots were dancing, boxing, and doing high kicks. So why does the vendor side of the hall exhibit superhuman agility, while the other carefully works to achieve a few minutes of stability on challenging terrain?

Though we have managed to create amazing robot hardware, we have only just begun to give it the ability to function in the world. It is easy to assume that because we have taught computers to read, write, think logically, and dance and sing (in the case of the humanoid robot demos), that we would also have an AI with a comprehensive understanding of the world. Let me unpack this a bit.

We take the ability to function as a body in the world for granted. If someone gave you a motorized puppet and gave you the controls that would allow you to move the arms, legs, and fingers, and then told you to pour yourself a cup of coffee using this puppet, you would have a better understanding of the level of difficulty in making a humanoid robot do the same thing. Today, we can teach robots to move in a human-like fashion, just like the puppet, via imitation learning. But when you put the robot into a new environment and ask it to pour the cup of coffee, it will be confronted with an astounding number of variables: what is a cup, where are cups stored, what is a coffee pot, how to know if it has coffee inside, how to hold the oddly shaped cup and the pot, and how to know when the cup is full. The list goes on.

Do we start writing software, albeit with the help of our newly created AI coding agents, that defines how to behave for each of the individual variables defined above, or do we try to create a new type of AI world model that is not focused on reading and writing, but is instead focused on the physics of objects and object relationships (the coffee goes inside the cup)? It would also understand the physics of motion, so that it would know how to grasp and manipulate these objects. The hope is that the AI model can generalize its understanding to include objects it has never seen before, in the same way chatbots (LLMs) can answer questions using knowledge derived from foundational sources.

I would say that this is the primary focus of the research community at this conference. We have a bevy of commercialized robot platforms; we have enormous compute power, in the cloud and at the edge; and now we need something else: an AI world model that will endow the humanoid robot with the intelligence it needs to operate within the human world and generalize to the various households and factories it will inhabit. Will this get us a robot that can do more than pour coffee, that can also do basic house cleaning? Probably not. There is still the matter of executive function, like knowing that the vacuum must be recharged before it can be used.

Establishing a truly useful AI world model will be far more difficult than what we have already achieved with the current crop of foundation models. It is hard for us humans to imagine why, because we often function subconsciously in the world, and we cannot deconstruct all the insights required for even the most basic physical actions.

IROS 2026, Day 3: Robot Competitions · Sep 28 – 29 📸
Shared album · Tap to view!

Keynotes

Even though the keynotes weren't focused on a particular paper like most of the other presentations, they were often the most thought-provoking. One in particular got me thinking about the creative process and how it has varied for me over the years, particularly since I am now a "Forever Engineer" (more descriptive than the diminutive "retired").

The image below was produced by Yigit Menguc, cofounder of Marble Wit Research. In each quadrant, the motivation of an idea is described by different criteria:

  • Fundamental: An idea that is more about establishing a foundational principle in a particular domain, of the kind that we might use for decades hence to establish many other ideas. Think of the laser, which gave rise to so many other ideas and continues to do so today
  • Comprehensive: An idea that fills a particular unmet need of its target user community. Imagine an innovative phone holder that magically floats above the dashboard by magnetic levitation.
  • Curiosity: An idea that must be explored because it might satisfy a desire for greater understanding. How do the tiny monarch butterflies navigate and locate their original breeding grounds in Mexico?
  • Market: An idea that will likely make money because there is a readily identifiable marketplace and clear demand. Imagine a cell phone that works anywhere in the world, indoors and out, using both terrestrial and satellite antennas.
C

I can see now that my creativity is generally focused in the lower-left quadrant, whereas I was almost exclusively in the lower-right quadrant when I was employed. This is enlightening because it reminds me that creativity need not lead to a profit!

The speaker said that we should prepare the mind and body for failure and accept failure with a sense of playfulness and glee. I completely agree, and I think such a perspective helps restore what makes learning so enjoyable, and what children often exhibit so effortlessly.