UNIST Develops AI That Reads Sarcasm to Reshape Facial Expressions

Professor Kim Tae-hwan's Team Unveils 'C-MET,' a Cross-Modal AI Linking Voice and Expression Synthesizes Subtle Emotional Shifts From Voice Alone, Without Separate Reference Photos Attaches Like a Plug-In to Existing AI… "An Innovation for Virtual Humans and Content Production"

Technology|
|
By Jang Ji-seung, Ulsan
||
Comparison of emotion editing results between C-MET and existing methods. Research image = UNIST - Seoul Economic Daily Technology News from South Korea
Comparison of emotion editing results between C-MET and existing methods. Research image = UNIST

A technology has been developed in which artificial intelligence (AI) recognizes the subtle tonal difference between praise and sarcasm in a short phrase like "well done," and realistically changes the facial expression of a speaker in a video.

A team led by Professor Kim Tae-hwan at the Graduate School of Artificial Intelligence at the Ulsan National Institute of Science and Technology (UNIST) announced Tuesday that it has developed an AI module called 'C-MET (Cross-Modal Emotion Transfer),' which extracts emotion from voice signals and transforms a speaker's facial expression in a video without a separate reference image.

Talking-face synthesis technology has been considered central to the creation of virtual humans and digital avatars. However, existing technologies either used a label-based approach, in which emotions were defined and trained as categories such as 'joy' or 'sadness,' or required separate high-quality reference photos showing the desired emotion. In particular, when based on voice, there was a limitation in that the content of speech and emotion signals became mixed together, making it virtually impossible to express unstructured emotions not included in the training data, such as sarcasm.

C-MET, unveiled by the research team, is a technology that connects different forms of information (modalities), namely voice and video. It calculates the difference between neutral speech and emotionally charged speech as a 'vector' (numerical information containing the direction and magnitude of change), and is designed so that the AI learns how this vector of emotional change in the voice translates into changes in facial expression.

Through this, it can extract only the pure emotion signal, separated from information about the content of speech or mouth shape, and transfer it into the facial expression space. Even for the same sentence, if the tone changes, the technology controls the expression so that the movements of the corners of the mouth, eyebrows, and the area around the eyes change delicately.

In particular, because it tracks the amount of change between two emotions rather than selecting one from a fixed list of emotions, it can draw out subtle emotions not directly encountered during training, such as sarcasm, empathy, and charisma. It also stands out in that, even without high-quality reference photos, with only voice it can naturally change just the facial expression while maintaining the original speaker's facial features and head movements.

Outstanding versatility is another strength. C-MET was developed as a lightweight 'plug-in' style module that can be inserted like a component, rather than being dedicated to a specific generative model. Simply replacing the heavy expression encoder of an existing talking-face generation AI with C-MET can improve performance without separate retraining.

In an actual experiment applying C-MET to the latest model, 'EDTalk,' the accuracy of emotional expression jumped about 14 percentage points, from 41.99% to 55.91% compared to before. In another model, 'PD-FGC,' emotion accuracy rose from 33.36% to 36.82%, and inference speed was confirmed to be noticeably shortened in both models.

"This research practically solves the limitations of existing methods, in that it can change the emotion of a facial video using only voice, without a reference image," Professor Kim Tae-hwan explained. "It is a foundational technology that can be widely used in various fields such as virtual human production, post-production work for films and content, and emotion-recognition AI."

This research, in which Choi Chan-hyuk, a master's student at the UNIST Graduate School of Artificial Intelligence, participated as the first author, was accepted to 'CVPR (Computer Vision and Pattern Recognition) 2026,' the world's most prestigious international conference in the field of AI and computer vision. The related code and demo video can be viewed through the research team's project page.

Original reporting by Jang Ji-seung, Ulsan for Seoul Economic Daily.

AI-translated from Korean. Quotes from foreign sources are based on Korean-language reports and may not reflect exact original wording.

Watch · Seoul Economic Daily

More →
5:23

AI KEY

Preview
Korean Corporate Intelligence HubKOSPI · KOSDAQ · 12 sectors

A live, cap-weighted view of every KOSPI and KOSDAQ sector, with same-day Korean reporting distilled by company — built for foreign investors, correspondents and analysts who need to scan Korea before the next session.

Korea Chaebol Tree

Preview
Families Behind the GroupsKFTC May 2026 · DART filings

An English-first interactive map of Samsung, SK, Hyundai, LG and Lotte — built for foreign investors, correspondents and analysts. Korea translates companies into English. We translate the families behind them.

SIGNAL

Pre-register
English Edition · Capital MarketsM&A · IPO · PE · Fund Flows

Pre-register for SIGNAL English Edition — a premium subscription bringing Korean capital markets coverage (M&A, IPOs, private equity, fund flows) to global institutional investors. First access to the 50% introductory rate.