-
LEGO Co-builder: Exploring Fine-Grained Vision-Language Modeling for Multimodal LEGO Assembly Assistants2025 arXiv (Cornell University) preprint Computer Science Multimodal Machine Learning Applications Open access
Hao‐Chen Huang, Y W Su, Xin Sun, Moonisa Ahsan, Mohammad Aliannejadi, Irene Viola, +8 more
1citations -
Case-Grounded Evidence Verification: A Framework for Constructing Evidence-Sensitive Supervision2026 arXiv (Cornell University) preprint Computer Science Multimodal Machine Learning Applications Open access
Soroosh Tayebi Arasteh, Mehdi Joodaki, Mahshad Lotfinia, Sven Nebelung, Daniel Truhn
0citations -
Vision-language models for chest radiography do not always need the image2026 arXiv (Cornell University) preprint Computer Science Multimodal Machine Learning Applications Open access
Mahshad Lotfinia, Sebastian Ziegelmayer, Lisa Adams, Tri-Thien Nguyen, Daniel Truhn, Andreas Maier, +1 more
0citations -
Optimizing VLP-aligned Multimodal Intent Representation with Correct Visual Instantiation for Zero-Shot Composed Image Retrieval2026 arXiv (Cornell University) preprint Computer Science Multimodal Machine Learning Applications Open access
Xuri Ge, Chunhao Wang, Junchen Fu, Haokun Wen, Zhiwei Xu, Ying Zhou, +4 more
0citations
15 results