Control over generation conditions, execution, feedback, and reusable experience across agentic image, video, 3D, and multimodal generation
PreprintMultimodal agents.
One research library.
Explore the systems, papers, and open resources
connecting perception, action, and feedback.
- 01
Observe
Language, vision
and other sensors - 02
Decide
Use context to choose
the next action - 03
Act
Use tools, interfaces
or robots - 04
Get feedback
Read results, changes
and responses
Explore the library
Interactive browsing could not load. Browse the complete list on GitHub, or reload to try again.
All entries
Browse the complete collection.
447 entries
Observation compression, memory, action efficiency, and runtime optimization for GUI agents; connects capability evaluation with token, latency, and execution costs
PreprintEvidence audit of 336 GUI-agent papers across architecture, interactive evaluation, recovery, lifecycle engineering, privacy, observability, human oversight, and deployment readiness
PreprintAgentic internal intelligence, external tool invocation, environment interaction, training resources, evaluation, and applications for multimodal large language models
PreprintArchitectures, tool use, collaboration, applications, and evaluation for LLM-driven multimodal agents
PreprintInteractive agents that integrate environmental perception, multisensory inputs, external knowledge, embodied action, and human feedback
PreprintConnects computer-use and robot-use through Perceive, Anticipate, Plan, Act, and Verify (PAPAV), with attention to physical constraints and gaps in evaluation
PreprintPhillip Isola examines general-purpose agents operating robots through sensor and actuator interfaces, including latency, task feedback, and the limits of current deployments
PerspectiveGeneral Robotics describes GRID workflows for hardware integration, simulation, skill construction, and deployment repair; laboratory examples are developer reports
Research blogEvidence-audited review of 42 intelligent-transportation agent families, separating model, system, and hybrid multimodality across capability, validation, deployment readiness, and bounded orchestration
PreprintFive functional roles for language, the embodied modules that consume it, and an evidence audit spanning actions, interventions, and causal grounding claims
PreprintContrasts end-to-end VLA weight policies with generated code skills, including closed-loop repair, skill memory and evolution, and interoperable skill ecosystems
PreprintResearch has one canonical placement. Engineering resources and skill packages are counted separately. Resource links do not imply released weights or a particular license.
Two ways into
agent research.
Start with what a work contributes.
Then explore where an agent acts.
What does the work add?
Research is grouped by its main evaluated contribution.
Agent systems
Complete runtimes with multimodal observations, grounded actions, and feedback.
Models & components
Reusable policies, perception, memory, and other learned capabilities.
Learning & planning
Algorithms for learning, tool use, world models, and coordination.
Benchmarks & data
Tasks, datasets, environments, and protocols that measure agent behavior.
Where does the agent act?
Systems and benchmarks share the same six domains.
Information & media
Search, inspect, and gather evidence
Creative workflows
Generate, edit, and verify artifacts
Conversation & interaction
Listen, respond, and coordinate
Software & interfaces
Operate browsers, desktops, and apps
Virtual & 3D worlds
Navigate and act in simulated spaces
Robotics & physical worlds
Sense, manipulate, and coordinate robots
Each work has one primary placement. Input modalities, memory, real-time behavior, and collaboration are described in its metadata. Robot-use belongs to the physical action domain. Read the taxonomy definitions ↗
Built for shared discovery.
A community-maintained library from OpenEnvision.
Help make the map more complete.