KAIST Develops AI Technology That Automatically Generates Sounds as If a “Jurassic Park” Dinosaur Were Actually Walking Toward You
<(From Left) Hyun-Bin Oh, Takida Yuhta, Uesaka Toshimitsu, Tae-Hyun Oh, Mitsufuji Yuki>
When people watch a scene in the film Jurassic Park where a giant dinosaur walks toward them, they naturally imagine a heavy, rumbling sound, as if the ground were shaking. This is because humans predict sound by considering not only the shape of an object, but also physical properties such as its size, weight, and speed of movement. However, existing video-to-audio generation AI mainly generates sound based on the category of objects or scene information in the video, and has not sufficiently reflected physical properties that vary depending on weight or speed.
KAIST (President Kwang Hyung Lee) announced on the 26th of May that a collaborative research team involving Professor Tae-Hyun Oh of the School of Computing, KAIST, together with joint researchers from POSTECH (President Sung Keun Kim) and Sony AI, has developed “PAVAS (Physics-Aware Video-to-Audio Synthesis),” an artificial intelligence (AI) technology that understands the physical situation in a video and generates more realistic sound.
<Concept Diagram of PAVAS (Physics-Aware Video-to-Audio Synthesis) Technology>
The key feature of this technology is that it is designed so that AI can infer invisible physical information such as the mass and velocity of objects in a video on its own. Ordinary videos do not provide exact numerical values for an object’s weight or speed, but the research team enabled AI to estimate them by analyzing the surrounding environment and movement context, and to reflect the results in the sound generation process.
In other words, the AI was designed to go beyond simply recognizing “what is visible” and to understand the physical cause of “why this sound should occur.”
As a result of technical validation, the research team’s AI generated sounds very similar to real-world environments in scenes involving physical interactions such as collisions or impacts between objects. In particular, it produced more realistic audio in which loudness and tone naturally changed when the mass and velocity of objects varied.
Recently, generative AI technologies that simultaneously generate video and audio have been advancing rapidly. Representative examples include Google’s “Veo 3” and ByteDance’s “Seedance 2.0.” However, in actual film, advertising, and game production sites, there is far greater demand for post-production work that adds sound effects suited to existing video scenes or supplements audio than for generating entirely new videos.
While existing commercial AI models have focused on generating video and audio together, PAVAS is differentiated by its ability to analyze the movement and collision characteristics of objects in a video and generate realistic sound effects that precisely match the scene.
<Comparison of Spectrograms Generated by Conventional Video-to-Audio Models and PAVAS>
The research team explained that this technology presents new possibilities in the field of “Physical AI,” or physically consistent generative AI. Physically consistent generative AI refers to AI that goes beyond simply producing plausible results and understands the laws of physics and causal relationships in the real world.
In the future, this technology is expected to provide more immersive user experiences in a wide range of fields, including the automation of content sound production, augmented reality (AR) and virtual reality (VR) content, the metaverse, and robotics simulation.
Professor Tae-Hyun Oh stated, “While existing generative AI has developed by increasing the scale of data and models, this research is meaningful in that it was designed so that AI directly understands physical quantities and causal relationships,” adding, “In the future, it can be expanded into a core foundational technology for next-generation multimodal AI that simultaneously understands and processes diverse types of information, including text, video, and speech.”
This study was led by POSTECH integrated M.S.-Ph.D. student Hyun-Bin Oh as the first author, with KAIST Professor Tae-Hyun Oh and Sony AI researchers Yuhta Takida, Toshimitsu Uesaka, and Yuki Mitsufuji participating as co-authors. This research was selected as an Oral presentation paper at CVPR 2026 (Computer Vision and Pattern Recognition 2026), the world’s most prestigious academic conference in the field of computer vision (image-based artificial intelligence technology), where only the top 0.88% of all papers are selected for oral presentation, recognizing the excellence of the work. The presentation is scheduled to take place on June 6.
※ Paper title: “PAVAS: Physics-Aware Video-to-Audio Synthesis,” DOI: https://arxiv.org/abs/2512.08282
This research was supported by the Mid-Career Research Program under the Basic Research Program of the Ministry of Science and ICT, the Pioneer Research Program for Future Converging Technology of the Ministry of Science, ICT and Future Planning, the AGI Program of the Ministry of Science and ICT, and the KAIST InnoCORE Program.
Soul-Searching & Odds-Defying Determination: A Commencement Story of Dr. Tae-Hyun Oh
(Dr. Tae-Hyun Oh, one of the 2736 graduates of the 2018)
Each and every one of the 2,736 graduates has come a long way to the 2018 Commencement. Tae-Hyun Oh, who just started his new research career at MIT after completing his Ph.D. at KAIST, is no exception.
Unlike the most KAIST freshmen straight out of the ingenious science academies of Korea, he is among the many who endured very challenging and turbulent adolescent years. Buffeted by family instability and struggling during his time at school, he saw himself trapped by seemingly impenetrable barriers. His mother, who hated to see his struggling, advised him to take a break to reflect on who he is and what he wanted to do.
After dropping out of high school in his first year, ways to make money and support his family occupied his thoughts. He took on odd jobs from a car body shop to a gas station, but the real world was very tough and sometimes even cruel to the high school dropout.
Bias and prejudice stigmatizing dropouts hurt him so much. He often overheard a parent who dropped by the body shop that he worked in saying, “If you do not study hard, you will end up like this guy.” Hearing such things terrified him and awoke his sense of purpose. So he decided to do something meaningful and be a better man than he was.
“I didn’t like the person I was growing up to become. I needed to find myself and get away from the place I was growing up. It was my adventure and it was the best decision I ever made,” says Oh.
After completing his high school diploma national certificate, he planned to apply to an engineering college. On his second try, he gained admission into the Department of Electrical Engineering at Kwang Woon University with a full scholarship. He was so thrilled for this opportunity and hoped he could do well at college. Signal processing and image processing became the interest of his research and he finished his undergraduate degree summa cum laude.
Gaining confidence in his studies, he searched around graduate school department websites in Korea to select the path he was interested in. Among others, the Robotics and Computer Vision Lab of Professor In-So Kweon at the Department of Electrical Engineering at KAIST was attractive to him.
Professor Kweon’s lab is globally renowned for robot vision technology. Their technologies were applied into HUBO, the KAIST-developed bimodal humanoid robot that won the 2015 DARPA Challenges.
“I am so appreciate of Professor Kweon, who accepted and guided me,” he said. Under Professor Kweon’s advising, he could finish his Master’s and Ph.D. courses in seven years. The mathematical modeling on fundamental computer algorithms became his main research topic.
While at KAIST, his academic research has blossomed. He won a total of 13 research prizes sponsored by corporations at home and abroad such as Kolon, Samsung, Hyundai Motors, and Qualcomm.
In 2015, he won the Microsoft Research Asia Fellowship as the sole Korean among 13 Ph.D. candidates in the Asian region. With the MSRA fellowship, he could intern at the MS Research Beijing Office for half a year and then in Redmond, Washington in the US.
“Professor Kweon’s lab filled me up with knowledge. Whenever I presented our team’s paper at an international conference, I was amazed by the strong interest shown by foreign experts, researchers, and professors. Their strong support and interest encouraged me a lot. I was fully charged with the belief that I could go abroad and explore more opportunities,” he said.
Dr. Oh, who completed his dissertation last fall, now works at the Department of Electrical Engineering and Computer Science at MIT under Professor Wojciech Matusik.
“I think the research environment at KAIST is on par with MIT. I have very rich resources for my studies and research at both schools, but at MIT the working culture is a little different and it remains a big challenge for me. I am still not familiar with collaborative work with colleagues from very diverse backgrounds and countries, and to persuade them and communicate with them is very tough. But I think I am getting better and better,” he said.
Oh, who is an avid computer game player as well, said life seems to be a game. The level of the game will be upgraded to the next level after something is accomplished. He feels great joy when he is moving up and he believes such diverse experiences have helped him become a better person day by day. Once he identified what gave him a strong sense of purpose, he wasn’t stressed out by his studies any more. He was so excited to be able to follow his passion and is ready for the next challenge.