A Multimodal Privacy Filtering System Using Deep Learning for Visual-Audio Input Streams

  • Kim, Kunwoo
  • Park, Sang-Won
  • Lee, Hyeon-Ju
  • Choi, Hoyong
  • Buu, Seok-Jun
Citations

WEB OF SCIENCE

0
Citations

SCOPUS

0

초록

Protecting personally identifiable information (PII) in visual and audio data streams that are continuously being captured by sensing systems remains a critical challenge. Devices such as IoT cameras and smart microphones routinely capture sensitive elements such as faces, voices, and behavioral or contextual cues, often without the subject's awareness or consent. To address this issue, we propose a multimodal PII filtering system designed for privacy protection in environments where visual and auditory data are persistently recorded. The proposed system detects and irreversibly anonymizes PII at the point of capture, before the data is transmitted or stored in vulnerable infrastructure. It incorporates a visual object detection module based on YOLOv12 and a sentence-level classifier based on BERT, applied to speech transcriptions generated by a speech-to-text module. These modules operate within a unit-based processing framework that segments incoming data into short temporal units, enabling low-latency operation while maintaining semantic consistency across modalities. Although each modality is processed independently, the system maintains temporal synchronization to ensure coherent filtering decisions. We evaluate the system using both in-house and public datasets across diverse conditions including variable lighting and background noise. The system achieves a unified false negative rate of about 3%, showing reliable performance for real-world multimodal privacy protection. Furthermore, the system employs parallel unit-based processing to maximize computational efficiency, and its modular design supports flexible component combinations, confirming suitability for edge or cloud deployment. These findings demonstrate that the proposed system provides an efficient and scalable solution for real-world multimodal privacy protection.

키워드

Data privacyProtectionVideosCamerasImage edge detectionIdentification of personsVisualizationReal-time systemsPrivacyEncryptionPersonally identifiable information (PII)multimodal privacy filteringdeep learningvisual-audio data processingspeech processingINTERNET
제목
A Multimodal Privacy Filtering System Using Deep Learning for Visual-Audio Input Streams
저자
Kim, KunwooPark, Sang-WonLee, Hyeon-JuChoi, HoyongBuu, Seok-Jun
DOI
10.1109/ACCESS.2026.3670925
발행일
2026-03
유형
Article
저널명
IEEE Access
14
페이지
36491 ~ 36504