CAD-Coder: An Open-Source Vision-Language Model for Computer-Aided Design Code GenerationSource: Journal of Mechanical Design:;2026:;volume( 148 ):;issue:007::page 147DOI: 10.1115/1.4071305Publisher: The American Society of Mechanical Engineers (ASME)
Abstract: Abstract. Efficient creation of accurate and editable 3D CAD models is critical in engineering design, significantly impacting cost and time to market in product innovation. Current manual workflows remain highly time consuming and demand extensive user expertise. While recent developments in AI-driven CAD generation show promise, existing approaches are often constrained by limited CAD representations, inability to generalize to real-world images, or low output accuracy due to a lack of domain-specific knowledge. To address these challenges, this article introduces CAD-Coder, an open-source vision-language model (VLM) explicitly fine-tuned to generate editable CAD code (CadQuery Python) directly from visual input. Leveraging a large-scale dataset that we constructed—GenCAD-Code, consisting of over 163k CAD model image and code pairs—CAD-Coder outperforms VLM baselines such as GPT-4.5 and Qwen2.5-VL-72B on image-conditioned CAD code generation, achieving a 100% valid syntax rate and the highest accuracy in 3D solid similarity. Notably, our VLM exhibits initial signs of generalizability, generating CAD code from real-world images and utilizing a CAD operation not explicitly seen during fine-tuning. The performance and adaptability of CAD-Coder highlight the potential of VLMs fine-tuned on code to streamline CAD workflows for engineers and designers. CAD-Coder is publicly available.
|
Collections
Show full item record
| contributor author | Doris, Anna C. | |
| contributor author | Alam, Ferdous | |
| contributor author | Heyrani Nobari, Amin | |
| contributor author | Ahmed, Faez | |
| date accessioned | 2026-08-23T07:19:43Z | |
| date available | 2026-08-23T07:19:43Z | |
| date copyright | 2026/07/01 | |
| date issued | 2026 | |
| identifier issn | 1050-0472 | |
| identifier other | md-25-1707.pdf | |
| identifier uri | http://yetl.yabesh.ir/yetl1/handle/yetl/4314947 | |
| description abstract | Abstract. Efficient creation of accurate and editable 3D CAD models is critical in engineering design, significantly impacting cost and time to market in product innovation. Current manual workflows remain highly time consuming and demand extensive user expertise. While recent developments in AI-driven CAD generation show promise, existing approaches are often constrained by limited CAD representations, inability to generalize to real-world images, or low output accuracy due to a lack of domain-specific knowledge. To address these challenges, this article introduces CAD-Coder, an open-source vision-language model (VLM) explicitly fine-tuned to generate editable CAD code (CadQuery Python) directly from visual input. Leveraging a large-scale dataset that we constructed—GenCAD-Code, consisting of over 163k CAD model image and code pairs—CAD-Coder outperforms VLM baselines such as GPT-4.5 and Qwen2.5-VL-72B on image-conditioned CAD code generation, achieving a 100% valid syntax rate and the highest accuracy in 3D solid similarity. Notably, our VLM exhibits initial signs of generalizability, generating CAD code from real-world images and utilizing a CAD operation not explicitly seen during fine-tuning. The performance and adaptability of CAD-Coder highlight the potential of VLMs fine-tuned on code to streamline CAD workflows for engineers and designers. CAD-Coder is publicly available. | |
| publisher | The American Society of Mechanical Engineers (ASME) | |
| title | CAD-Coder: An Open-Source Vision-Language Model for Computer-Aided Design Code Generation | |
| type | Journal Paper | |
| journal volume | 148 | |
| journal issue | 7 | |
| journal title | Journal of Mechanical Design | |
| identifier doi | 10.1115/1.4071305 | |
| journal fristpage | 147 | |
| journal lastpage | 162 | |
| page | 16 | |
| tree | Journal of Mechanical Design:;2026:;volume( 148 ):;issue:007 | |
| contenttype | Fulltext |