SEOUL, August 14 (AJP) - A research team at Sookmyung Women's University has built an artificial intelligence model that works out when a listener should react during a conversation, treating that as a separate question from what the reaction should look like.
The model, named CaReDiff, took second place in the REACT26 Grand Challenge organized by ACM Multimedia, and the work has been accepted for an oral presentation session at the conference, which opens in Rio de Janeiro in November.
CaReDiff generates a video of a listener whose face responds to what a speaker is saying and to how the speaker looks while saying it. The team is led by Kim Byung-gyu, a professor in the university's Department of Artificial Intelligence Engineering.
The difficulty the work addresses is timing. Earlier systems learned to react at the right moment and to use the right expression as a single task, which left them reacting at the wrong points in a live conversation. The same design meant that tuning a model to one person's habits could damage the reaction patterns it had already learned.
The team split the two problems apart. One technique, which the researchers call causal lead-lag cross attention, lines up the moment of the reaction explicitly rather than leaving it to be inferred. A second works in stages, planning the overall arc of the expression first and then refining the detail.
Personal style is handled by a separate lightweight module, the Personalised Residual Adapter, which sits alongside the trained model instead of altering it. It carries personality traits and a record of how the individual has reacted in the past. Less than one percent of the model's total settings are trained to produce a personal reaction style, and the rest are left as they are.
The team tested the system on the MARS dataset. On FRCorr, the measure the challenge uses to score how closely a generated reaction matches the reactions real people gave, the model reached 0.7671 in the offline setting. In the real-time setting, it reached 0.6462, up from 0.5680.
With the personalization module applied, the offline figure rose to 0.7868 from 0.5899 and the real-time figure to 0.6510 from 0.5475. The reaction timing the model had learned before personalization was applied stayed intact.
Those figures come from tests against recorded conversation data. The team has not reported results from human viewers judging whether the generated faces read as natural, and has not shown the system running inside a working chatbot or robot.
The researchers expect the technology to be used in chatbots, social robots and humanoid robots, in services where a machine has to hold an emotional exchange with a person.
Choi Ji-hoon was the first author. Choi Se-in, Lee Han-seul and Qurat Ul Ain Aisha took part in the research.
"What matters about this result is that we have secured the core underlying technology that lets AI understand the context of a conversation and the characteristics of an individual, and respond at the right moment," Kim said. "We will refine this human-centered empathetic AI so that it understands human emotion and intent and communicates naturally, and carry it through to real-world demonstration."
Copyright ⓒ Aju Press All rights reserved.


