Video · video-to-audio · audio-generation

ThinkSound

Worldfoundry Native

Hidden from Studio until the infer-only runtime and ThinkSound diffusion checkpoint are both available in-tree/Hugging Face cache.

Metadata OnlyUnified environmentthinksound

Build your run

Choose a recorded variant, then copy the exact setup or inspection command.

worldfoundry.pipeline
Taskvideo-to-audio
Environmentworldfoundry-unified-cu128
DeviceCUDA 12.8
worldfoundry-eval evaluate \
  --mode model \
  --model-id thinksound \
  --model-runner worldfoundry:pipeline \
  --model-manifest-dir worldfoundry/data/models/catalog \
  --requests-path tmp/requests.jsonl \
  --output-dir tmp/model_eval/thinksound \
  --metric artifact_count \
  --json
01

Compatibility & versions

Manifest-backed recipe

Integration
metadata only
Runner evidence
pending
Environment
worldfoundry-unified-cu128
Python
3.11
CUDA
CUDA 12.8
PyTorch
torch
Source revision
Checkpoint revision
Runtime profile
thinksound
Pipeline binding
thinksound
Runner
worldfoundry.pipeline
Pipeline target
worldfoundry.pipelines.video_official.pipeline_official_video:ThinkSoundPipeline
Backend stage
official_runtime_bridge
Runtime status
pending_checkpoint_and_in_tree_runtime
Driver status
compatible_unified_default
Environment kind
Unified environment
02

Install environment

The environment resolver reads the recorded profile and chooses the unified or dedicated environment shown here.

bash scripts/setup/model_env_install.sh --model thinksound
Package constraints 4
  • torch
  • torchvision
  • torchaudio
  • xfuser
Conda packages 5
  • python=3.11
  • pip
  • ffmpeg
  • setuptools
  • wheel
03

Checkpoints & assets

Run the local check before allocating compute. Gated, private, and license fields below come directly from the checkpoint manifest.

worldfoundry-eval zoo model-download --model-id thinksound --check-local --json
FunAudioLLM/ThinkSound
Revision
License
Gated
Private
04

Launch & outputs

The generated command uses the shared evaluation boundary and writes durable result manifests and artifacts.

worldfoundry-eval evaluate \
  --mode model \
  --model-id thinksound \
  --model-runner worldfoundry:pipeline \
  --model-manifest-dir worldfoundry/data/models/catalog \
  --requests-path tmp/requests.jsonl \
  --output-dir tmp/model_eval/thinksound \
  --metric artifact_count \
  --json

Input contract

FieldRecorded contract
promptRequired
videooptional

Artifact contract

Artifact kindFilename / path
generated_audiothinksound.wav
05

Evidence & sources

Catalog integration, native-demo parity, and runner parity are independent records.

Integration
metadata only
Runner evidence
pending
Native demo evidence
Not recorded
Validation imports
torch, torchvision, numpy, transformers, diffusers, xfuser, worldfoundry
Recipe provenance
worldfoundry/data/models/catalog/video/thinksound.yamlGitHubhttps://huggingface.co/FunAudioLLM/ThinkSound