Multimodal agents.
One research library.

Explore the systems, papers, and open resources
connecting perception, action, and feedback.

GoalComplete a task in the world
  1. 01

    Observe

    Language, vision
    and other sensors

  2. 02

    Decide

    Use context to choose
    the next action

  3. 03

    Act

    Use tools, interfaces
    or robots

New evidence changes the next action.
A living research collection
447Entries
8Collections
2018–2026Research span

Explore the library

All entries

Browse the complete collection.

447 entries

Control over generation conditions, execution, feedback, and reusable experience across agentic image, video, 3D, and multimodal generation

Preprint
PaperCollection

Observation compression, memory, action efficiency, and runtime optimization for GUI agents; connects capability evaluation with token, latency, and execution costs

Preprint
Paper

Evidence audit of 336 GUI-agent papers across architecture, interactive evaluation, recovery, lifecycle engineering, privacy, observability, human oversight, and deployment readiness

Preprint
Paper

Agentic internal intelligence, external tool invocation, environment interaction, training resources, evaluation, and applications for multimodal large language models

Preprint
PaperCollection

Architectures, tool use, collaboration, applications, and evaluation for LLM-driven multimodal agents

Preprint
Paper

Interactive agents that integrate environmental perception, multisensory inputs, external knowledge, embodied action, and human feedback

Preprint
Paper

Connects computer-use and robot-use through Perceive, Anticipate, Plan, Act, and Verify (PAPAV), with attention to physical constraints and gaps in evaluation

Preprint
PaperProjectCollection

Phillip Isola examines general-purpose agents operating robots through sensor and actuator interfaces, including latency, task feedback, and the limits of current deployments

Perspective
Article

General Robotics describes GRID workflows for hardware integration, simulation, skill construction, and deployment repair; laboratory examples are developer reports

Research blog
Article

Evidence-audited review of 42 intelligent-transportation agent families, separating model, system, and hybrid multimodality across capability, validation, deployment readiness, and bounded orchestration

Preprint
PaperCollection

Five functional roles for language, the embodied modules that consume it, and an evidence audit spanning actions, interventions, and causal grounding claims

Preprint
Paper

Contrasts end-to-end VLA weight policies with generated code skills, including closed-loop repair, skill memory and evolution, and interoperable skill ecosystems

Preprint
Paper

Research has one canonical placement. Engineering resources and skill packages are counted separately. Resource links do not imply released weights or a particular license.

01 / The research map

Two ways into
agent research.

Start with what a work contributes.
Then explore where an agent acts.

By contribution

What does the work add?

Research is grouped by its main evaluated contribution.

By action domain

Where does the agent act?

Systems and benchmarks share the same six domains.

Each work has one primary placement. Input modalities, memory, real-time behavior, and collaboration are described in its metadata. Robot-use belongs to the physical action domain. Read the taxonomy definitions ↗

02 / Open research

Built for shared discovery.

A community-maintained library from OpenEnvision.
Help make the map more complete.

Contribute on GitHub ↗
OpenEnvision LibraryMultimodal Modeling ↗Reusable proceduresMultimodal Agent Skills ↗Keep exploringRelated research libraries ↗
From the research library