Not finding it? Sign in to also search OpenAlex live.
-
From Superficial to Deep: Language Bias driven Curriculum Learning for Visual Question Answering2021 ACM International Conference on Multimedia (ACM MM) conference-paper Computer Science Multimodal Machine Learning Applications
Mingrui Lao, Yanming Guo, Yu Liu, Wei Chen, Nan Pu, Michael S. Lew
22citations -
COCA: COllaborative CAusal Regularization for Audio-Visual Question Answering2023 Proceedings of the AAAI Conference on Artificial Intelligence conference-paper Computer Science Multimodal Machine Learning Applications Open access
Mingrui Lao, Nan Pu, Yu Liu, Kai He, Erwin M. Bakker, Michael S. Lew
22citations -
Optimizing VLP-aligned Multimodal Intent Representation with Correct Visual Instantiation for Zero-Shot Composed Image Retrieval2026 arXiv (Cornell University) preprint Computer Science Multimodal Machine Learning Applications Open access
Xuri Ge, Chunhao Wang, Junchen Fu, Haokun Wen, Zhiwei Xu, Ying Zhou, +4 more
0citations -
Sparse Local Latents for Explainable Zero-Shot Reasoning in Medical Vision-Language Models2026 Lecture notes in computer science conference-paper Computer Science Multimodal Machine Learning Applications
Martin Goetze, Patrick Wienholt, Dennis Eschweiler, Marvin Gazibaric, Christiane Kuhl, Sven Nebelung, +1 more
0citations -
HotelMatch-LLM: Joint Multi-Task Training of Small and Large Language Models for Efficient Multimodal Hotel Retrieval2025 Annual Meeting of the Association for Computational Linguistics (ACL) conference-paper Computer Science Multimodal Machine Learning Applications Open access
Arian Askari, Emmanouil Stergiadis, Ilya Gusev, Moran Beladev
0citations -
Prompt Injection Attacks on Large Language Models in Oncology2024 arXiv (Cornell University) preprint Computer Science Adversarial Robustness in Machine Learning Open access
Jan Clusmann, Dyke Ferber, Isabella C. Wiest, Carolin Victoria Schneider, Titus Josef Brinker, Sebastian Foersch, +2 more
5citations -
LEGO Co-builder: Exploring Fine-Grained Vision-Language Modeling for Multimodal LEGO Assembly Assistants2025 arXiv (Cornell University) preprint Computer Science Multimodal Machine Learning Applications Open access
Hao‐Chen Huang, Y W Su, Xin Sun, Moonisa Ahsan, Mohammad Aliannejadi, Irene Viola, +8 more
1citations -
LEGO Co-builder: Exploring Fine-Grained Vision-Language Modeling for Multimodal Assembly Assistants2026 INTERNATIONAL CONFERENCE ON MULTIMODAL INTERACTION conference-paper Computer Science Multimodal Machine Learning Applications Open access
Haochen Huang, Yue Su, Xin Sun, Moonisa Ahsan, Mohammad Aliannejadi, Irene Viola, +8 more
0citations