TY - GEN
T1 - Cross-Modal Feature Alignment and MMD Improve Robustness of Prompt Tuning
AU - Sun, Jingchen
AU - Sharma, Rohan
AU - Lokhande, Vishnu Suresh
AU - Chen, Changyou
N1 - Publisher Copyright:
© 2025 IEEE.
PY - 2025
Y1 - 2025
N2 - Prompt Tuning has emerged as a prominent research paradigm for adapting vision-language models to various downstream tasks. However, recent research indicates that prompt tuning methods often lead to overfitting due to limited training samples. In this paper, we propose a Cross-modal Aligned Feature Tuning (CRAFT) method to address this issue. Cross-modal alignment is conducted by first selecting anchors from the alternative domain and deriving relative representations of the embeddings for the selected anchors. Optimizing for a feature alignment loss over anchor-aligned text and image modalities creates a more unified text-image common space. Overfitting in prompt tuning also deteriorates model performance on out-of-distribution samples. To further improve the prompt model's robustness, we propose minimizing Maximum Mean Discrepancy (MMD) over the anchor-aligned feature spaces to mitigate domain shift. The experiment on four different prompt tuning structures consistently shows the improvement of our method, with increases of up to 6.1% in the Base-to-Novel generalization task, 5.8% in the group robustness task, and 2.7% in the out-of-distribution tasks. The code is available at https://github.com/lingchensun/Craft.
AB - Prompt Tuning has emerged as a prominent research paradigm for adapting vision-language models to various downstream tasks. However, recent research indicates that prompt tuning methods often lead to overfitting due to limited training samples. In this paper, we propose a Cross-modal Aligned Feature Tuning (CRAFT) method to address this issue. Cross-modal alignment is conducted by first selecting anchors from the alternative domain and deriving relative representations of the embeddings for the selected anchors. Optimizing for a feature alignment loss over anchor-aligned text and image modalities creates a more unified text-image common space. Overfitting in prompt tuning also deteriorates model performance on out-of-distribution samples. To further improve the prompt model's robustness, we propose minimizing Maximum Mean Discrepancy (MMD) over the anchor-aligned feature spaces to mitigate domain shift. The experiment on four different prompt tuning structures consistently shows the improvement of our method, with increases of up to 6.1% in the Base-to-Novel generalization task, 5.8% in the group robustness task, and 2.7% in the out-of-distribution tasks. The code is available at https://github.com/lingchensun/Craft.
KW - cross-modal alignment
KW - out-of-distribution
KW - prompt tuning
KW - vision-language model
UR - https://www.scopus.com/pages/publications/105003643617
U2 - 10.1109/WACV61041.2025.00462
DO - 10.1109/WACV61041.2025.00462
M3 - Conference contribution
AN - SCOPUS:105003643617
T3 - Proceedings - 2025 IEEE Winter Conference on Applications of Computer Vision, WACV 2025
SP - 4714
EP - 4724
BT - Proceedings - 2025 IEEE Winter Conference on Applications of Computer Vision, WACV 2025
PB - Institute of Electrical and Electronics Engineers Inc.
T2 - 2025 IEEE/CVF Winter Conference on Applications of Computer Vision, WACV 2025
Y2 - 28 February 2025 through 4 March 2025
ER -