SEOUL, July 31 (AJP) - South Korean researchers have developed two technologies that stop artificial intelligence from claiming to hear sounds that were never recorded or see objects that do not exist, a step toward AI systems reliable enough for self-driving cars, rescue robots, and hospitals.
The Korea Advanced Institute of Science and Technology (KAIST) said Friday that a team led by Professor Ro Yong-man from the institute's School of Electrical Engineering developed techniques to address two persistent weaknesses in multimodal AI, systems that process text, images, and sound together. One teaches AI to correctly read specialized sensors, such as thermal cameras and X-rays. The other prevents the system from confusing what it sees with what it hears. Neither requires retraining the AI from scratch, which the institute said makes both fast and cheap to deploy in industry.
The problem the team set out to solve is known as hallucination, the tendency of AI models to generate confident but false output. In multimodal systems the errors take a particular form. Shown silent security footage of a passing car, existing models often report hearing an engine, simply because a car is visible on screen.
The first technology, which the team calls Diverse Negative Attributes (DNA), addresses how AI misreads sensors that measure signals other than visible light. A person looking at a thermal image knows the bright areas are hot. Conventional AI, trained mostly on ordinary photographs, tends to interpret the brightness as light. The team developed a training method that repeatedly feeds the model examples it gets wrong, teaching it the physical meaning behind thermal, depth, and X-ray imagery. The result is a system that recognizes objects more accurately in darkness and smoke, using only a small amount of additional data.
The second technology, Modality-Adaptive Decoding (MAD), tackles the crossed wires between sight and sound. Before answering a question, the AI first asks itself whether the answer should come from the video or the audio, then concentrates on the relevant channel while suppressing interference from the other. Because the method works at the moment the AI generates its answer, it can be added to existing models like a software plug-in, with no retraining at all.
KAIST said the technologies could serve autonomous vehicles that must spot pedestrians at night or in fog, rescue robots searching for survivors in smoke-filled buildings, drones carrying thermal cameras, airport security systems that analyze X-ray scans, and medical AI that reads CT, MRI, and X-ray images together.
"For multimodal AI to be used in real environments, it is important that it accurately understands the characteristics of various sensors and does not confuse different sensory inputs," Ro said. "This research lays the groundwork for multimodal AI that can be trusted in daily life and industrial settings by reducing sensory bias and hallucination without large-scale retraining."
Doctoral student Chung Sang-yun served as first author on both studies, with Yoo Young-jun joining as co-first author on the DNA research. The MAD study was presented in June at the Conference on Computer Vision and Pattern Recognition (CVPR), the top international venue in the field, and the DNA study was published in the journal IEEE Transactions on Image Processing.
(Reference Information)
Journal/Source: IEEE Transactions on Image Processing and IEEE / CVF Conference on Computer Vision and Pattern Recognition
Title: Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking MAD: Modality-Adaptive Decoding for Mitigating Cross-Modal Hallucinations in Multimodal Large Language Models
Link/DOI: 10.48550/arXiv.2412.20750 / 10.48550/arXiv.2601.21181
Copyright ⓒ Aju Press All rights reserved.


