SAGE-Prompt: Structured Attribution Guarded Explanation for Explainable Deepfake Question Answering

Citations

SCOPUS

0

초록

Deepfake detectors often fail to generalize when manipulation methods, compression settings, or capture pipelines change. Large vision-language models (LVLMs) can be used in zero-/few-shot mode and provide natural-language rationales, but ad-hoc prompt search leads to format violations, hallucinated evidence, and frequent "I am not sure"deferrals, making evaluation hard to reproduce. We propose SAGE-Prompt (Structured Attribution Guarded Explanation), a meta-prompting framework tuned via Optimization by PROmpting (OPRO) that replaces such ad-hoc prompts with a compact JSON schema comprising (1) a tri-state artifact label (yes/no/reject), (2) region selection from a facial codebook, and (3) short, concrete region-level clues. SAGE-Prompt adds schema validation and cross-region consistency checks as guardrails, standardizing both prompts and outputs while leaving backbone LVLMs and preprocessing unchanged. On cross-dataset deepfake benchmarks with local and API LVLMs, SAGE-Prompt substantially reduces format violations, stabilizes region-level explanations, and yields more controlled rejection behavior, although LVLM predictions alone remain insufficient as robust deepfake detectors under distribution shift. We view SAGE-Prompt as a practical baseline for hybrid pipelines where LVLMs supply structured attribution, while separate modules handle calibration and final decisions. © 2026 Owner/Author.

키워드

deepfake detection and qaexplanation consistencymeta-promptingselective predictionvision-language models
제목
SAGE-Prompt: Structured Attribution Guarded Explanation for Explainable Deepfake Question Answering
저자
Park, Jong-ChanKim, MyeongjunLim, So-HeeChoi, Sang-MinKim, Gun-Woo
DOI
10.1145/3774904.3792898
발행일
2026-04
유형
Conference paper
저널명
WWW 2026 - Proceedings of the ACM Web Conference 2026
페이지
8513 ~ 8516