TP 3 : Multimodal Learning and its application
We will explore multimodal models, focusing on how to fuse information from different data types including text, images, and audio.
We will dive into advanced topics, including Joint Optimization of modalities using Contrastive Language–Image Pre-training and diffusion models.
[Read More]