
A core physical AI technology that allows artificial intelligence (AI) to learn human judgment criteria from just a few videos has emerged in Korea. Because AI can learn behaviors aligned with human intent without people having to directly evaluate thousands to tens of thousands of robot actions, the technology is expected to significantly reduce the time and cost required to develop robots, autonomous vehicles, and AI agents.
KAIST said Thursday that a research team led by Professor Yoo Chang-dong of the School of Electrical Engineering has developed the world's first "VOTP (Video-based Optimal TransPort Preference)" technology, which learns human judgment criteria from only a small number of preference videos.
Physical AI refers to AI that moves and acts directly in the real world, such as robots and autonomous vehicles, going beyond AI that generates text and images. Representative examples include robots that perform dangerous tasks in factories, autonomous vehicles that assess complex road situations, and medical robots that carry out precise surgeries.
For physical AI to be used in actual field settings, machines must be able to judge which behaviors better align with human intent. To do this, AI must learn a "reward function," the criterion for distinguishing good behavior from bad behavior. Previously, however, creating a reward function required people to directly watch and evaluate robot actions thousands to tens of thousands of times. This process consumed enormous time and cost, and has been cited as a major barrier to commercializing physical AI.
Professor Yoo's team focused on the fact that people learn new tasks after watching only a few demonstrations. VOTP, developed by the team, is designed to let AI grasp human-preferred behavior patterns using only around 10 good demonstration videos and bad demonstration videos. Video AI visually analyzes subtle differences in robot behavior and uses a mathematical technique called "optimal transport" to automatically infer preferences across numerous videos.
Through this, AI can judge on its own which behaviors are more desirable, even in various situations that people have not individually evaluated, and can learn behaviors aligned with human intent. The team explained that it confirmed VOTP's learning efficiency and generalization performance in experiments targeting various environments and tasks.
The technology can be applied across physical AI as a whole, including robotic arm control, humanoid robots, autonomous vehicles, smart factories, drones, and surgical robots. It also has potential applications in software AI agents that view computer screens and perform tasks on their own. In particular, when introducing a new robot to an industrial site, an expert can select and evaluate just a few field videos, and AI can analyze numerous situations based on them and learn optimal behavior, reducing testing periods and data-building costs.
Lou Minh Tung, a doctoral student in the School of Electrical Engineering, participated as the first author in this research. The research paper has been accepted at "ICML 2026," a world-class AI conference to be held at COEX in Seoul this July, and was selected as an oral presentation paper, given to only 168 papers in the top 0.7% out of a total of 23,918 submitted papers.






