Poster: SAF: Semantic-Aware Flushing for Latency and Jitter Suppression in Continuous VLA Inference on Edge Devices

Citations

SCOPUS

0

초록

In memory-constrained edge devices, continuous Vision-Language-Action (VLA) inference commonly offloads KV caches via mmap to NVMe storage. However, OS background flushers such as pdflush/kworker are unaware of inference timing, triggering bursty writeback I/O during execution and causing high tail latency and jitter that undermine real-Time control. To address this, we propose Semantic-Aware Flushing (SAF), which detects the logical completion of each frame response as a semantic boundary and asynchronously flushes dirty pages during the inter-frame idle interval. We evaluate four flush policies using Moondream2 on Jetson Orin NX. SAF reduces mean latency by 24.2% (from 9,804 ms to 7,427 ms), P99 tail latency by 27.7% (from 16,174 ms to 11,696 ms), and jitter (IQR) by 77.8% (from 2,893 ms to 642 ms), while eliminating all 295 OS flusher interventions observed under the default policy. © 2026 IEEE.

키워드

Edge Computing; I/O Scheduling; Jitter; KV Cache Offloading; Tail Latency; VLA
제목
Poster: SAF: Semantic-Aware Flushing for Latency and Jitter Suppression in Continuous VLA Inference on Edge Devices
저자
Jeong, Junhyeok; Kim, Jaeho
DOI
10.1109/NVMSA71223.2026.11658877
발행일
2026-08
유형
Conference paper
저널명
Proceedings - 2026 IEEE 15th Non-Volatile Memory Systems and Applications Symposium, NVMSA 2026