News & Events

Press Release

AI Learns to Focus Like Humans to Speed Up Video Analysis

The model identifies key video moments, reducing analysis time for 2-minute interview clips by 65%

As AI systems increasingly process video, audio, and language data, reducing unnecessary computation has become a requirement. Researchers from the Japan Advanced Institute of Science and Technology (JAIST) developed an AI model that identifies important video segments and processes only relevant information. The model reduced processing time for 2-minute clips by 65%, from 52 to 18 seconds, while maintaining high accuracy. This approach could enable efficient, real-time AI coaching and communication tools on everyday devices.

Artificial neural networks were originally inspired by the human brain, but they are still far less efficient at processing information. One reason the human brain is so efficient is its ability to focus only on the most relevant information and allocate cognitive effort based on the task. As artificial intelligence (AI) systems increasingly analyze multiple types of data, such as text, video, audio, and images, this ability to focus on the most important information is becoming increasingly important for reducing computational time and resource use. Most video frames contain little useful information, yet conventional AI systems still analyze them all, increasing computational cost and sometimes introducing noise that can reduce prediction accuracy.

A research team, led by Professor Shogo Okada from the Japan Advanced Institute of Science and Technology (JAIST), Japan, along with his Doctoral Student Hung Le from JAIST, has developed an AI model that identifies the key moments in a video, reducing the processing time for a 2-minute clip, from 52 seconds to 18 seconds. The findings of the study were made available online on July 11, 2026, and will be published in Volume 137 of the journal Information Fusion on January 01, 2027.

Doctoral Candidate Le, who is the first author of the paper, compares the model to the way humans naturally pay attention during a conversation: "Humans do not constantly keep their eyes on their conversation partner during a conversation. They first notice changes in sound and only direct their gaze at the specific moments they perceive as important."

The model, called EMF-dVAE (Efficient Multimodal Fusion with a discrete Variational Autoencoder), consists of two main components: a discrete variational autoencoder (dVAE) and a multimodal fusion (MF) network.

During training, the dVAE is given partially corrupted visual data, where the accompanying audio is first used to identify which visual segments should be masked. The model learns to reconstruct the missing visual information and, through this process, learns which visual regions are most informative for the task. Once training is complete, the dVAE identifies only the most important visual segments during inference. These selected visual features are then combined with the audio and language information by the MF network to produce the final prediction.

Since the model only analyzes a fraction of the video during inference, it requires much less computation while maintaining high accuracy. When tested on the ETS-Interview dataset, which contains 1,891 2-minute job interview videos from 260 participants, EMF-dVAE achieved state-of-the-art performance while using only 15.42% of the available visual features. It also reduced the processing time per video by approximately 65%, from 52 seconds to 18 seconds. Remarkably, by ignoring nearly 85% of the visual data, the model became both faster and more accurate, as redundant video frames can obscure informative signals while increasing computational cost.

"Much like the human brain, which focuses its attention on the most relevant moments, the AI automatically adjusts how much video data it analyzes for each clip, allowing it to allocate computational resources more efficiently," says Prof. Okada.

By making video analysis more efficient, the framework could enable AI-powered video interview coaches and communication-training tools that provide users with affordable, real-time feedback at near-expert levels. Because the framework processes only the most informative visual information, it also has the potential to reduce computational cost and energy consumption, supporting the development of more sustainable AI systems.

"Video is becoming the dominant form of data, and AI that must watch everything will not scale-- economically or environmentally. Within 5-10 years, AI that budgets its attention the way humans do could make multimodal assistants, interview coaches, tutoring systems, and communication-support robots affordable and responsive on everyday devices," says Prof. Okada.

pr20260805-1e.jpg

Title: Listen first―look only when it matters
Caption: The AI first analyzes the audio to identify a few-second intervals worth examining. Visual analysis is conducted only for the selected intervals, while the remaining frames are ignored. The extracted visual information is then combined with the audio and transcript data to generate the final evaluation.
Credit: Professor Shogo Okada from JAIST, Japan
License type: Original content
Usage restrictions: Cannot be reused without permission

Reference

Title of original paper: Audio-guided visual selection for efficient multimodal fusion via a discrete variational autoencoder
Authors: Hung Le*, Hung-Hsuan Huang, Candy Olivia Mawalim, Chee Wee Leong, and Shogo Okada*
Journal: Information Fusion
DOI: 10.1016/j.inffus.2026.104613

Additional information for EurekAlert

Latest Article Publication Date: 01 January 2027
Method of Research: Computational simulation/modeling
Subject of Research:  Not Applicable
Conflicts of Interest Statement: The authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper.

Funding information

This work was partially supported by JSPS KAKENHI (26K03016), and JST CREST, Japan (JPMJCR2563), and JST CRONOS (JPMJCS24K7).

August 5, 2026

PAGETOP