In late August, humanoid robots gathered in Beijing to run, jump, box, play football, and perform martial arts at the 2026 World Humanoid Robot Games. In an opening heat, one completed the 100-meter sprint in 9.39 seconds—faster than Usain Bolt’s human world record of 9.58 seconds. By the end of the week, several robots had gone under Bolt’s mark, with the winner finishing in 8.64 seconds.
They also fell down. A lot.
Robots crashed into padded barriers, threw sparks from their waists, stumbled during races, and were carried off on stretchers. Earlier in the year, humanoid robots had already danced during China’s Lunar New Year celebrations, while a robot fighting competition produced the surreal spectacle of one machine kicking the head off another.
The machines were becoming, to borrow from Daft Punk, harder, better, faster, stronger.
But not necessarily smarter—not in the sense that matters most in the physical world: the ability to transfer skills, improvise when the room changes, and recover when a plan fails.
These spectacles hint at what many researchers consider the next frontier: embodied AI. Unlike the AI we have come to know through screens, embodied AI combines perception, reasoning, planning, and action in the physical world. It does not merely describe what should be done. It attempts to do it.
Around the World
One of the strongest critics of the idea that ever-larger language models alone will lead to human-like intelligence is Yann LeCun, the pioneering computer scientist who left Meta in late 2025 to establish AMI Labs, where he is pursuing advanced machine intelligence built around world models.
Large language models learn from vast collections of human writing. They can tell us a remarkable amount about gravity, motion, objects, distance, and cause and effect. But knowing how humans describe the world is not the same as understanding how the world works.
A language model can explain what will happen when someone pushes a glass toward the edge of a table. A robot reaching across that table must determine whether its arm will hit the glass, predict what might happen if it does, and alter its movement before the glass falls.
LeCun argues that advanced AI will require world models: internal representations that allow machines to anticipate how an environment may change and what consequences particular actions produce.
Meta’s V-JEPA 2, developed as part of this research program, learns from video to model physical reality. After additional training using less than 62 hours of robot-arm data, it planned movements involving unfamiliar objects in new laboratory environments—reaching, grasping, and placing. The broader ambition is for machines to observe, predict, plan, act, and learn from experience.
This begins to resemble how biological intelligence develops. Long before children can read about gravity, they have dropped things. They learn that objects disappear behind obstacles but continue to exist, that surfaces support weight, and that actions produce consequences.
For embodied AI, the world itself becomes part of the training data.
Doin’ It Right
This helps explain why embodied AI represents another major transition in the Intelligence Age.
Earlier AI became increasingly good at recognizing: faces, objects, patterns, speech. Generative AI learned to create: writing, images, music, video, and code. Agentic AI is beginning to pursue goals, performing sequences of tasks rather than simply answering prompts.
Embodied AI adds another verb: intervene.
The difference is simple. Instead of asking, “Tell me how to do this,” we increasingly ask, “Do this.”
That matters because much of human civilization remains stubbornly physical. Food must be produced, packages moved, buildings maintained, patients assisted, warehouses operated, and older people cared for. Generative AI largely entered the information economy. Embodied AI potentially extends artificial intelligence into the material economy.
Civilization has always tested intelligence at the point of contact with objects and other bodies. Screens allowed artificial intelligence to skip that test. Embodiment puts it back.
Recent advances make this increasingly conceivable. Multimodal models integrate language with vision, spatial information, sound, and sensor feedback. Vision-language-action models attempt to translate perception directly into physical movement.
Simulation lets robots accumulate experience without repeatedly breaking expensive hardware. But reality remains messier: friction changes, sensors misread, lighting shifts, surfaces wear. No simulation can anticipate every physical surprise.
Agentic systems then supply the high-level logic: breaking a broad goal like “clean the kitchen” into smaller tasks, deciding what to do first, and adapting when something goes wrong.
But the physical world imposes a standard that digital AI can sometimes avoid.
A language model can produce an incorrect answer that sounds perfectly convincing. A robot cannot bluff gravity. It either keeps its balance or falls. It grasps the object or drops it. It clears the doorway or crashes into it.
Reality provides its own fact-checking mechanism.
Short Circuit
That is what makes the failures at the Robot Games as revealing as the victories.
A machine capable of sprinting extraordinarily fast may still struggle to manipulate an unfamiliar object. A robot trained to fight may not know how to fold clothes. These reveal the distinction between highly developed physical capabilities and general physical intelligence.
Some of the most transformative achievements in embodied AI may eventually look far less impressive than robot boxing: preparing food without spilling, opening unfamiliar packaging, finding a dropped object, or safely helping an older person stand.
That same week, Beijing also hosted the 2026 World Robot Conference, featuring robots sorting packages, sewing, folding clothes, and fetching objects. More than spectacle, these were attempts to imagine where robotics ultimately goes: into ordinary human environments and useful work.
The strange lesson of Beijing was that a machine could outrun Usain Bolt before it could reliably navigate the ordinary messiness of human life.
First, machines learned to recognize. Then, to generate. Now, to act.
What comes next is making intelligence survive contact with reality.
Read more Stories on Simpol.ph
Ylona Garcia Celebrates a Decade of Heartache, Growth, and Sound
M1ss Jade So Channels Duality in Her Latest Release “BAPHOMETA”
Dr. Louward Allen Zubiri To Teach Filipino Language At Yale University






















