상세 보기
대규모 언어모델의 건축환경분야 성능평가를 위한 벤치마크 구축과 관련한 탐색적 연구
SCOPUS
0초록
This study proposes a benchmark construction method for systematically evaluating the performance of large language models (LLMs) in thebuilding environmental domain and presents an exploratory experiment applying the benchmark to on-device models. A dual-tier benchmarkframework was developed, categorizing items into core performance indicators and extended performance indicators based on expert consensus. A total of 120 question?answer items were created using educational materials in the building environmental field, and their importance wasassessed by experts. As a result, 32 items, or 26.7 percent, were classified as core performance indicators, 83 items, or 69.2 percent, asextended performance indicators, and 5 items, or 4.2 percent, were excluded from the benchmark. The proposed benchmark was then appliedto evaluate two on-device LLMs. The results showed that the models achieved accuracy rates of 43.8 to 59.4 percent on core performanceindicators and 39.1 to 52.2 percent on extended performance indicators. Both models demonstrated higher accuracy on the core performanceindicators, suggesting that concepts with stronger expert consensus were more likely to be reflected in training data for LLMs. Overall, thefindings indicate that a dual-tier benchmark based on expert consensus can serve as an effective tool for evaluating domain-specificknowledge in LLMs within the building environmental field.
키워드
- 제목
- 대규모 언어모델의 건축환경분야 성능평가를 위한 벤치마크 구축과 관련한 탐색적 연구
- 제목 (타언어)
- An Exploratory Study of Benchmark Construction for Performance Evaluation of Large Language Models in the Building Environmental Domain
- 저자
- 정창헌
- 발행일
- 2026-05
- 유형
- Y
- 저널명
- 대한건축학회논문집
- 권
- 42
- 호
- 5
- 페이지
- 291 ~ 299