Multimodal models.
One research library.

Explore the architectures, papers, and open resources
shaping multimodal intelligence.

Survey paperComing soon

A survey of multimodal foundation models.

GitHub

Explore the list and contribute.

Many modalities, shared intelligenceText, images, and audio flow into a shared model that supports understanding and generation. A conceptual illustration, not a taxonomy assignment. TIA TextImageAudioUnderstandingGeneration
TextImageAudio
UnderstandingGeneration
Many modalities. Shared intelligence.
A living research collection
360Entries
5Collections
2011–2026Research span

Explore the library

All entries

Browse the complete collection.

360 entries

Theoretical identifiability analysis of contrastive multimodal learning

ICLR 2023
Paper

Introduces CutMix-style augmentation for unpaired VLP

ICML 2022
Paper

Balances modality learning via dynamic gradient reweighting

CVPR 2022
Paper

Unified architecture for vision-language understanding and generation

arXiv 2021
Paper

Single transformer for multiple multimodal tasks

arXiv 2021
Paper

Benchmark suite for multimodal learning evaluation

NeurIPS 2021
Paper

General-purpose architecture for high-dimensional multimodal inputs

ICML 2021
Paper

Contrastive vision-language pretraining at scale

arXiv 2021
Paper

Improved visual features for VL tasks

arXiv 2021
Paper

Early large-scale vision-language contrastive learning

arXiv 2020
Paper

Unified multi-task learning across 12 VL tasks

CVPR 2020
Paper

Contrastive transformer for video representation learning

arXiv 2019
Paper

Counts reflect list entries; a model may appear in more than one category. Resource links do not imply a particular license or available weights.

01 / The taxonomy

Four perspectives on
multimodal intelligence.

Four complementary views.
Membership can overlap.

Defining component or relationRepresentative patterns · Image + text
01

Traditional

Two distinct patterns: CLIP-style image and text encoders align embeddings through comparison; fusion models combine modality features before task prediction.

Embedding alignment and cross-modal feature interaction are distinct patterns.

Explore collection ↗
02

MLLMs

Text tokens enter a pretrained LLM. Visual features condition it through a modality interface, either as input tokens or through layer-wise attention. The dashed route is an alternative.

A pretrained language backbone receives perceptual evidence through a modality interface.

Explore collection ↗
03

UMMs

Text and visual encodings enter a common understanding-and-generation modeling framework. Text decoding and a visual decoder or generator produce the outputs. Encoders may be shared or task-specific.

Understanding and generation share a modeling framework; encoders and decoders may differ.

Explore collection ↗
04

NMMs

Early-fusion example: modality interfaces feed packed multimodal states before the first shared backbone block. Joint state evolution is an architectural property; early multimodal optimization of core weights is a separate training property. Generation is optional.

Fusion structure and foundation-training history are assessed separately. UMM membership may overlap.

Explore collection ↗
Closed-source is a separate collection.

Availability does not imply an architectural category.

Browse closed-source ↗

UMMs are defined by understanding–generation unification. NMMs distinguish architectural integration from training history; early fusion does not automatically mean training every component from scratch. Read the full definitions ↗

02 / Open research

Built for shared discovery.

A community-maintained library from OpenEnvision.
Help make the map more complete.

Contribute on GitHub ↗
On Hugging FaceUnified Multimodal Model Zoo ↗On Hugging FaceNative Multimodal Model Zoo ↗Keep exploringSurveys, tools & related lists ↗
From the research library