World models · image-to-video · camera-controlled-video · interactive-world-model · long-video-generation

AlayaWorld

LTX-2.3 World Models

The runtime reuses WorldFoundry's existing LTX-2.3 Transformer, VAE, Gemma connector, RoPE, patchifier, attention dispatcher, and distributed helpers.

IntegratedUnified environmentPinned sourcealayaworld

Build your run

Choose a recorded variant, then copy the exact setup or inspection command.

worldfoundry.pipeline
Taskimage-to-video
Environmentworldfoundry-unified-cu128
DeviceCUDA 12.8
worldfoundry-eval evaluate \
  --mode model \
  --model-id alayaworld \
  --model-runner worldfoundry:pipeline \
  --model-manifest-dir worldfoundry/data/models/catalog \
  --requests-path tmp/requests.jsonl \
  --output-dir tmp/model_eval/alayaworld \
  --metric artifact_count \
  --json
01

Compatibility & versions

Manifest-backed recipe

Integration
integrated
Runner evidence
checkpoint gpu validated
Environment
worldfoundry-unified-cu128
Python
3.11
CUDA
CUDA 12.8
PyTorch
torch>=2.6
Source revision
58678a00eb6fde0a77206b7466c65ab622acfeb9
Checkpoint revision
e3f1adebec5cda1c77c55c115cde47fba279f887
Runtime profile
alayaworld
Pipeline binding
alayaworld
Runner
worldfoundry.pipeline
Pipeline target
worldfoundry.pipelines.alayaworld.pipeline_alayaworld:AlayaWorldPipeline
Backend stage
independent_in_tree_runtime
Runtime status
checkpoint_gpu_validated_official_case1_single_chunk
Driver status
compatible
Environment kind
Unified environment
02

Install environment

The environment resolver reads the recorded profile and chooses the unified or dedicated environment shown here.

bash scripts/setup/model_env_install.sh --model alayaworld
Package constraints 20
  • torch>=2.6
  • torchvision
  • transformers>=4.49
  • tokenizers>=0.20.3
  • safetensors
  • huggingface-hub
  • accelerate
  • sentencepiece
  • einops
  • numpy
  • scipy
  • pillow
  • opencv-python
  • imageio
  • imageio-ffmpeg
  • av
  • omegaconf
  • easydict
  • tqdm
  • torchao
Conda packages 4
  • python
  • pip
  • ffmpeg
  • ninja
03

Checkpoints & assets

Run the local check before allocating compute. Gated, private, and license fields below come directly from the checkpoint manifest.

worldfoundry-eval zoo model-download --model-id alayaworld --check-local --json
AlayaLab/AlayaWorld
Revision
e3f1adebec5cda1c77c55c115cde47fba279f887
License
Gated
Private
google/gemma-3-12b-it-qat-q4_0-unquantized
Revision
68f7ee4fbd59087436ada77ed2d62f373fdd4482
License
Gated
true
Private
depth-anything/DA3NESTED-GIANT-LARGE-1.1
Revision
b2359bdf726fb44ef62acca04d629dcf158053e7
License
Gated
Private
04

Launch & outputs

The generated command uses the shared evaluation boundary and writes durable result manifests and artifacts.

worldfoundry-eval evaluate \
  --mode model \
  --model-id alayaworld \
  --model-runner worldfoundry:pipeline \
  --model-manifest-dir worldfoundry/data/models/catalog \
  --requests-path tmp/requests.jsonl \
  --output-dir tmp/model_eval/alayaworld \
  --metric artifact_count \
  --json

Input contract

FieldRecorded contract
promptRequired
imageRequired
videoOptional
actionscamera_trajectory, camera_path

Artifact contract

Artifact kindFilename / path
generated_world_video
generated_worldalayaworld.mp4
05

Evidence & sources

Catalog integration, native-demo parity, and runner parity are independent records.

Integration
integrated
Runner evidence
checkpoint gpu validated
Native demo evidence
Not recorded
Validation imports
torch, torchvision, transformers, safetensors, einops, cv2, worldfoundry.base_models.diffusion_model.video.ltx2.alayaworld, worldfoundry.base_models.three_dimensions.depth.depth_anything.depth_anything_v3.api, worldfoundry.synthesis.visual_generation.alayaworld