상세 보기
A Multimodal Privacy Filtering System Using Deep Learning for Visual-Audio Input Streams
- Kim, Kunwoo;
- Park, Sang-Won;
- Lee, Hyeon-Ju;
- Choi, Hoyong;
- Buu, Seok-Jun
WEB OF SCIENCE
0SCOPUS
0초록
Protecting personally identifiable information (PII) in visual and audio data streams that are continuously being captured by sensing systems remains a critical challenge. Devices such as IoT cameras and smart microphones routinely capture sensitive elements such as faces, voices, and behavioral or contextual cues, often without the subject's awareness or consent. To address this issue, we propose a multimodal PII filtering system designed for privacy protection in environments where visual and auditory data are persistently recorded. The proposed system detects and irreversibly anonymizes PII at the point of capture, before the data is transmitted or stored in vulnerable infrastructure. It incorporates a visual object detection module based on YOLOv12 and a sentence-level classifier based on BERT, applied to speech transcriptions generated by a speech-to-text module. These modules operate within a unit-based processing framework that segments incoming data into short temporal units, enabling low-latency operation while maintaining semantic consistency across modalities. Although each modality is processed independently, the system maintains temporal synchronization to ensure coherent filtering decisions. We evaluate the system using both in-house and public datasets across diverse conditions including variable lighting and background noise. The system achieves a unified false negative rate of about 3%, showing reliable performance for real-world multimodal privacy protection. Furthermore, the system employs parallel unit-based processing to maximize computational efficiency, and its modular design supports flexible component combinations, confirming suitability for edge or cloud deployment. These findings demonstrate that the proposed system provides an efficient and scalable solution for real-world multimodal privacy protection.
키워드
- 제목
- A Multimodal Privacy Filtering System Using Deep Learning for Visual-Audio Input Streams
- 저자
- Kim, Kunwoo; Park, Sang-Won; Lee, Hyeon-Ju; Choi, Hoyong; Buu, Seok-Jun
- 발행일
- 2026-03
- 유형
- Article
- 저널명
- IEEE Access
- 권
- 14
- 페이지
- 36491 ~ 36504