Probability versus Prompting: Language Model Performance on Dependencies beyond English

초록

Recent work on Tranformer-based language models (LMs) has raised a methodological debate over how to best evaluate LMs’ linguistic knowledge: through direct probability-based measures of token likelihood or through prompt-based interactions that elicit explicit responses. While prompting has become increasingly popular, its reliability as a diagnostic tool for linguistic competence remains unclear. Moreover, most prior evaluations have focused on English, leaving open questions about cross-linguistic generalizability. This paper compares probability-based and prompt-based evaluation methods in two typologically distinct languages, Hindi and Korean, targeting politeness dependencies that involve long-distance coherence. Using controlled minimal pairs, we assess models of different scales. Our results show a clear advantage for probability-based evaluation. Notably, larger prompt-based models do not outperform smaller probability-accessible models, suggesting that model scale or training data does not necessarily compensate for methodological limitations in evaluation. Our findings suggest that probability-based methods provide a more accurate and efficient window into LMs’ linguistic representations than prompt-based approaches. This study further underscores the importance of evaluation methodology in cross-linguistic LM research and calls for the need to move beyond English-centered linguistic assessments.

키워드

language modelsprobabilitypromptinglong-distance dependencypoliteness
제목
Probability versus Prompting: Language Model Performance on Dependencies beyond English
저자
이수환Shaonan Wang
발행일
2026-04
유형
Y
저널명
언어학 연구
79
페이지
205 ~ 221