대규모 언어모델의 건축환경분야 성능평가를 위한 벤치마크 구축과 관련한 탐색적 연구

An Exploratory Study of Benchmark Construction for Performance Evaluation of Large Language Models in the Building Environmental Domain
Citations

SCOPUS

0

초록

This study proposes a benchmark construction method for systematically evaluating the performance of large language models (LLMs) in thebuilding environmental domain and presents an exploratory experiment applying the benchmark to on-device models. A dual-tier benchmarkframework was developed, categorizing items into core performance indicators and extended performance indicators based on expert consensus. A total of 120 question?answer items were created using educational materials in the building environmental field, and their importance wasassessed by experts. As a result, 32 items, or 26.7 percent, were classified as core performance indicators, 83 items, or 69.2 percent, asextended performance indicators, and 5 items, or 4.2 percent, were excluded from the benchmark. The proposed benchmark was then appliedto evaluate two on-device LLMs. The results showed that the models achieved accuracy rates of 43.8 to 59.4 percent on core performanceindicators and 39.1 to 52.2 percent on extended performance indicators. Both models demonstrated higher accuracy on the core performanceindicators, suggesting that concepts with stronger expert consensus were more likely to be reflected in training data for LLMs. Overall, thefindings indicate that a dual-tier benchmark based on expert consensus can serve as an effective tool for evaluating domain-specificknowledge in LLMs within the building environmental field.

키워드

Large language modelBuilt EnvironmentBenchmarkDomain Performance EvaluationExpert Consensus대형언어모델건축환경벤치마크도메인 성능평가전문가 합의
제목
대규모 언어모델의 건축환경분야 성능평가를 위한 벤치마크 구축과 관련한 탐색적 연구
제목 (타언어)
An Exploratory Study of Benchmark Construction for Performance Evaluation of Large Language Models in the Building Environmental Domain
저자
정창헌
DOI
10.5659/JAIK.2026.42.5.291
발행일
2026-05
유형
Y
저널명
대한건축학회논문집
42
5
페이지
291 ~ 299