Theoretical identifiability analysis of contrastive multimodal learning
ICLR 2023Multimodal models.
One research library.
Explore the architectures, papers, and open resources
shaping multimodal intelligence.
A survey of multimodal foundation models.
Explore the list and contribute.
Explore the library
Interactive browsing could not load. Browse the complete list on GitHub, or reload to try again.
All entries
Browse the complete collection.
360 entries
Introduces CutMix-style augmentation for unpaired VLP
ICML 2022Balances modality learning via dynamic gradient reweighting
CVPR 2022Unified architecture for vision-language understanding and generation
arXiv 2021Single transformer for multiple multimodal tasks
arXiv 2021Benchmark suite for multimodal learning evaluation
NeurIPS 2021General-purpose architecture for high-dimensional multimodal inputs
ICML 2021Contrastive vision-language pretraining at scale
arXiv 2021Improved visual features for VL tasks
arXiv 2021Early large-scale vision-language contrastive learning
arXiv 2020Unified multi-task learning across 12 VL tasks
CVPR 2020Contrastive transformer for video representation learning
arXiv 2019Counts reflect list entries; a model may appear in more than one category. Resource links do not imply a particular license or available weights.
Four perspectives on
multimodal intelligence.
Four complementary views.
Membership can overlap.
Traditional
Embedding alignment and cross-modal feature interaction are distinct patterns.
Explore collection ↗MLLMs
A pretrained language backbone receives perceptual evidence through a modality interface.
Explore collection ↗UMMs
Understanding and generation share a modeling framework; encoders and decoders may differ.
Explore collection ↗NMMs
Fusion structure and foundation-training history are assessed separately. UMM membership may overlap.
Explore collection ↗Availability does not imply an architectural category.
UMMs are defined by understanding–generation unification. NMMs distinguish architectural integration from training history; early fusion does not automatically mean training every component from scratch. Read the full definitions ↗
Built for shared discovery.
A community-maintained library from OpenEnvision.
Help make the map more complete.