AI Judges in Design: Toward Expert-Equivalent Design Evaluations With Vision-Language Models and In-Context LearningSource: Journal of Mechanical Design:;2026:;volume( 148 ):;issue:007::page 41DOI: 10.1115/1.4071835Publisher: The American Society of Mechanical Engineers (ASME)
Abstract: Abstract. The subjective evaluation of early-stage engineering designs, such as concept sketches, traditionally relies on human experts. However, expert evaluations are time-consuming, expensive, and sometimes inconsistent. Recent advances in vision-language models (VLMs) offer the potential to automate design assessments, but it is crucial to ensure that these artificial intelligence (AI) “judges” perform on par with human experts. This work introduces in-context learning (ICL)-enhanced VLM judges and a comprehensive statistical framework (including agreement, error, correlation, statistical difference checks, equivalence testing, and top-set overlap) to rigorously assess AI–expert equivalence. Across two case studies, we show that reasoning-enabled VLMs are the strongest-performing AI judges. They consistently outperform two-third trained novices across all metrics, and for measures such as uniqueness, creativity, and drawing quality, they approach expert-equivalent performance. In specific cases, they even exceed expert–expert agreement, attaining lower mean absolute error and higher rank correlations than the expert baseline. These findings suggest that, on certain statistical tests, AI judges are not only approaching expert–expert equivalence but in some cases surpassing it.
|
Collections
Show full item record
| contributor author | Edwards, Kristen M. | |
| contributor author | Tehranchi, Farnaz | |
| contributor author | Miller, Scarlett | |
| contributor author | Ahmed, Faez | |
| date accessioned | 2026-08-23T07:19:50Z | |
| date available | 2026-08-23T07:19:50Z | |
| date copyright | 2026/07/01 | |
| date issued | 2026 | |
| identifier issn | 1050-0472 | |
| identifier other | md-25-1709.pdf | |
| identifier uri | http://yetl.yabesh.ir/yetl1/handle/yetl/4314951 | |
| description abstract | Abstract. The subjective evaluation of early-stage engineering designs, such as concept sketches, traditionally relies on human experts. However, expert evaluations are time-consuming, expensive, and sometimes inconsistent. Recent advances in vision-language models (VLMs) offer the potential to automate design assessments, but it is crucial to ensure that these artificial intelligence (AI) “judges” perform on par with human experts. This work introduces in-context learning (ICL)-enhanced VLM judges and a comprehensive statistical framework (including agreement, error, correlation, statistical difference checks, equivalence testing, and top-set overlap) to rigorously assess AI–expert equivalence. Across two case studies, we show that reasoning-enabled VLMs are the strongest-performing AI judges. They consistently outperform two-third trained novices across all metrics, and for measures such as uniqueness, creativity, and drawing quality, they approach expert-equivalent performance. In specific cases, they even exceed expert–expert agreement, attaining lower mean absolute error and higher rank correlations than the expert baseline. These findings suggest that, on certain statistical tests, AI judges are not only approaching expert–expert equivalence but in some cases surpassing it. | |
| publisher | The American Society of Mechanical Engineers (ASME) | |
| title | AI Judges in Design: Toward Expert-Equivalent Design Evaluations With Vision-Language Models and In-Context Learning | |
| type | Journal Paper | |
| journal volume | 148 | |
| journal issue | 7 | |
| journal title | Journal of Mechanical Design | |
| identifier doi | 10.1115/1.4071835 | |
| journal fristpage | 41 | |
| journal lastpage | 54 | |
| page | 14 | |
| tree | Journal of Mechanical Design:;2026:;volume( 148 ):;issue:007 | |
| contenttype | Fulltext |