TFI2F-Net: Traffic-Facial Intra-Inter Frame Fusion Network for driver emotion recognition under naturalistic driving
Date published
Free to read from
Supervisor/s
Industry supervisor/s
Journal Title
Journal ISSN
Volume Title
Department
Course name
Type
ISSN
Format
Citation
Abstract
Driver emotion is a critical factor affecting both road traffic safety and human–machine interaction experience. Existing recognition methods mainly rely on individual affective cues, such as facial expressions, while rarely considering their collaborative modeling with traffic context. To address this limitation, this paper proposes a Traffic–Facial Intra–Inter Frame Fusion Network (TFI2F-Net), which explicitly models the stagewise interaction between traffic and facial frame sequence. The proposed framework consists of three key stages. First, an intra-frame cross-modal encoder based on dual-path attention mechanism is designed to adaptively model fine-grained interactions between traffic and facial features at the frame level. Second, an inter-frame reweighted semantic fusion module is developed to emphasize key frames and integrate temporal representations under disentangled semantic guidance. Third, auxiliary unimodal supervision and modality separation losses are incorporated to enhance single modality discriminability and promote effective disentanglement. In addition, we construct a naturalistic scenario driver emotion dataset (Scenario-Emo) comprising outside-view and inside-view video clips, where a participant–expert joint annotation protocol is employed to ensure the reliability. Extensive experiments on the Scenario Emo and AIDE datasets demonstrate that the proposed TFI2F Net consistently outperforms state-of-the-art methods. Overall, our model and dataset are expected to facilitate the development of affective interaction in intelligent cockpits.
