MCERF: Advancing Multimodal Large Language Model Evaluation of Engineering Documentation With Enhanced RetrievalSource: Journal of Mechanical Design:;2026:;volume( 148 ):;issue:011::page 19528Author:Naghavi Khanghah, Kiarash
,
Anh Nguyen, Hoang
,
Doris, Anna C.
,
Mohammad Vahedi, Amir
,
Grandi, Daniele
,
Ahmed, Faez
,
Xu, Hongyi
DOI: 10.1115/1.4072033Publisher: The American Society of Mechanical Engineers (ASME)
Abstract: Abstract. Engineering rulebooks and technical standards contain multimodal information like dense text, tables, and illustrations that are challenging for retrieval augmented generation (RAG) systems. Building upon the DesignQA framework (Doris, A. C., Grandi, D., Tomich, R., Alam, M. F., Ataei, M., Cheong, H., and Ahmed, F., 2025, “Designqa: A Multimodal Benchmark for Evaluating Large Language Models’ Understanding of Engineering Documentation,” J. Comput. Inf. Sci. Eng., 25(2), p. 021009. 10.1115/1.4067333), which relied on full-text ingestion and text-based retrieval, this work establishes a multimodal ColPali-enhanced retrieval and reasoning framework (MCERF), a system that couples a multimodal retriever with large language model reasoning for accurate and efficient question answering from engineering documents. The system employs ColPali, which retrieves both textual and visual information, and multiple retrieval and reasoning strategies: (i) hybrid lookup mode for explicit rule mentions, (ii) vision to text fusion for figure- and table-guided queries, (iii) high-reasoning LLM mode for complex multi modal questions, and (iv) SelfConsistency decision to stabilize resp onses. The modular framework design provides a reusable template for future multimodal systems regardless of the underlying model architecture. Furthermore, this work establishes and compares two routing approaches: a single-case routing approach and an agent-based system, both of which dynamically allocate queries to optimal pipelines. Evaluation on the DesignQA benchmark illustrates that this system improves average accuracy across all tasks with a relative gain of +32.6% from baseline RAG best results, which is a significant improvement in multimodal and reasoning-intensive tasks without complete rulebook ingestion. This shows how vision-language retrieval, modular reasoning, and adaptive routing enable scalable document comprehension in engineering use cases. MCERF is publicly available online.
|
Collections
Show full item record
| contributor author | Naghavi Khanghah, Kiarash | |
| contributor author | Anh Nguyen, Hoang | |
| contributor author | Doris, Anna C. | |
| contributor author | Mohammad Vahedi, Amir | |
| contributor author | Grandi, Daniele | |
| contributor author | Ahmed, Faez | |
| contributor author | Xu, Hongyi | |
| date accessioned | 2026-08-23T07:30:57Z | |
| date available | 2026-08-23T07:30:57Z | |
| date copyright | 2026/11/01 | |
| date issued | 2026 | |
| identifier issn | 1050-0472 | |
| identifier other | md-26-1085.pdf | |
| identifier uri | http://yetl.yabesh.ir/yetl1/handle/yetl/4315206 | |
| description abstract | Abstract. Engineering rulebooks and technical standards contain multimodal information like dense text, tables, and illustrations that are challenging for retrieval augmented generation (RAG) systems. Building upon the DesignQA framework (Doris, A. C., Grandi, D., Tomich, R., Alam, M. F., Ataei, M., Cheong, H., and Ahmed, F., 2025, “Designqa: A Multimodal Benchmark for Evaluating Large Language Models’ Understanding of Engineering Documentation,” J. Comput. Inf. Sci. Eng., 25(2), p. 021009. 10.1115/1.4067333), which relied on full-text ingestion and text-based retrieval, this work establishes a multimodal ColPali-enhanced retrieval and reasoning framework (MCERF), a system that couples a multimodal retriever with large language model reasoning for accurate and efficient question answering from engineering documents. The system employs ColPali, which retrieves both textual and visual information, and multiple retrieval and reasoning strategies: (i) hybrid lookup mode for explicit rule mentions, (ii) vision to text fusion for figure- and table-guided queries, (iii) high-reasoning LLM mode for complex multi modal questions, and (iv) SelfConsistency decision to stabilize resp onses. The modular framework design provides a reusable template for future multimodal systems regardless of the underlying model architecture. Furthermore, this work establishes and compares two routing approaches: a single-case routing approach and an agent-based system, both of which dynamically allocate queries to optimal pipelines. Evaluation on the DesignQA benchmark illustrates that this system improves average accuracy across all tasks with a relative gain of +32.6% from baseline RAG best results, which is a significant improvement in multimodal and reasoning-intensive tasks without complete rulebook ingestion. This shows how vision-language retrieval, modular reasoning, and adaptive routing enable scalable document comprehension in engineering use cases. MCERF is publicly available online. | |
| publisher | The American Society of Mechanical Engineers (ASME) | |
| title | MCERF: Advancing Multimodal Large Language Model Evaluation of Engineering Documentation With Enhanced Retrieval | |
| type | Journal Paper | |
| journal volume | 148 | |
| journal issue | 11 | |
| journal title | Journal of Mechanical Design | |
| identifier doi | 10.1115/1.4072033 | |
| journal fristpage | 19528 | |
| journal lastpage | 19540 | |
| page | 13 | |
| tree | Journal of Mechanical Design:;2026:;volume( 148 ):;issue:011 | |
| contenttype | Fulltext |