YaBeSH Engineering and Technology Library

    • Journals
    • PaperQuest
    • YSE Standards
    • YaBeSH
    • Login
    View Item 
    •   YE&T Library
    • ASME
    • Journal of Mechanical Design
    • View Item
    •   YE&T Library
    • ASME
    • Journal of Mechanical Design
    • View Item
    • All Fields
    • Source Title
    • Year
    • Publisher
    • Title
    • Subject
    • Author
    • DOI
    • ISBN
    Advanced Search
    JavaScript is disabled for your browser. Some features of this site may not work without it.

    Archive

    AI Judges in Design: Toward Expert-Equivalent Design Evaluations With Vision-Language Models and In-Context Learning

    Source: Journal of Mechanical Design:;2026:;volume( 148 ):;issue:007::page 41
    Author:
    Edwards, Kristen M.
    ,
    Tehranchi, Farnaz
    ,
    Miller, Scarlett
    ,
    Ahmed, Faez
    DOI: 10.1115/1.4071835
    Publisher: The American Society of Mechanical Engineers (ASME)
    Abstract: Abstract. The subjective evaluation of early-stage engineering designs, such as concept sketches, traditionally relies on human experts. However, expert evaluations are time-consuming, expensive, and sometimes inconsistent. Recent advances in vision-language models (VLMs) offer the potential to automate design assessments, but it is crucial to ensure that these artificial intelligence (AI) “judges” perform on par with human experts. This work introduces in-context learning (ICL)-enhanced VLM judges and a comprehensive statistical framework (including agreement, error, correlation, statistical difference checks, equivalence testing, and top-set overlap) to rigorously assess AI–expert equivalence. Across two case studies, we show that reasoning-enabled VLMs are the strongest-performing AI judges. They consistently outperform two-third trained novices across all metrics, and for measures such as uniqueness, creativity, and drawing quality, they approach expert-equivalent performance. In specific cases, they even exceed expert–expert agreement, attaining lower mean absolute error and higher rank correlations than the expert baseline. These findings suggest that, on certain statistical tests, AI judges are not only approaching expert–expert equivalence but in some cases surpassing it.
    • Download: (1.369Mb)
    • Show Full MetaData Hide Full MetaData
    • Get RIS
    • Item Order
    • Go To Publisher
    • Statistics

      AI Judges in Design: Toward Expert-Equivalent Design Evaluations With Vision-Language Models and In-Context Learning

    URI
    https://yetl.yabesh.ir/yetl1/handle/yetl/4314951
    Collections
    • Journal of Mechanical Design

    Show full item record

    contributor authorEdwards, Kristen M.
    contributor authorTehranchi, Farnaz
    contributor authorMiller, Scarlett
    contributor authorAhmed, Faez
    date accessioned2026-08-23T07:19:50Z
    date available2026-08-23T07:19:50Z
    date copyright2026/07/01
    date issued2026
    identifier issn1050-0472
    identifier othermd-25-1709.pdf
    identifier urihttp://yetl.yabesh.ir/yetl1/handle/yetl/4314951
    description abstractAbstract. The subjective evaluation of early-stage engineering designs, such as concept sketches, traditionally relies on human experts. However, expert evaluations are time-consuming, expensive, and sometimes inconsistent. Recent advances in vision-language models (VLMs) offer the potential to automate design assessments, but it is crucial to ensure that these artificial intelligence (AI) “judges” perform on par with human experts. This work introduces in-context learning (ICL)-enhanced VLM judges and a comprehensive statistical framework (including agreement, error, correlation, statistical difference checks, equivalence testing, and top-set overlap) to rigorously assess AI–expert equivalence. Across two case studies, we show that reasoning-enabled VLMs are the strongest-performing AI judges. They consistently outperform two-third trained novices across all metrics, and for measures such as uniqueness, creativity, and drawing quality, they approach expert-equivalent performance. In specific cases, they even exceed expert–expert agreement, attaining lower mean absolute error and higher rank correlations than the expert baseline. These findings suggest that, on certain statistical tests, AI judges are not only approaching expert–expert equivalence but in some cases surpassing it.
    publisherThe American Society of Mechanical Engineers (ASME)
    titleAI Judges in Design: Toward Expert-Equivalent Design Evaluations With Vision-Language Models and In-Context Learning
    typeJournal Paper
    journal volume148
    journal issue7
    journal titleJournal of Mechanical Design
    identifier doi10.1115/1.4071835
    journal fristpage41
    journal lastpage54
    page14
    treeJournal of Mechanical Design:;2026:;volume( 148 ):;issue:007
    contenttypeFulltext
    DSpace software copyright © 2002-2015  DuraSpace
    نرم افزار کتابخانه دیجیتال "دی اسپیس" فارسی شده توسط یابش برای کتابخانه های ایرانی | تماس با یابش
    yabeshDSpacePersian
     
    DSpace software copyright © 2002-2015  DuraSpace
    نرم افزار کتابخانه دیجیتال "دی اسپیس" فارسی شده توسط یابش برای کتابخانه های ایرانی | تماس با یابش
    yabeshDSpacePersian