Video · video-to-audio · audio-generation
ThinkSound
Worldfoundry Native
Hidden from Studio until the infer-only runtime and ThinkSound diffusion checkpoint are both available in-tree/Hugging Face cache.
Build your run
Choose a recorded variant, then copy the exact setup or inspection command.
Taskvideo-to-audio
Environmentworldfoundry-unified-cu128
DeviceCUDA 12.8
worldfoundry-eval evaluate \
--mode model \
--model-id thinksound \
--model-runner worldfoundry:pipeline \
--model-manifest-dir worldfoundry/data/models/catalog \
--requests-path tmp/requests.jsonl \
--output-dir tmp/model_eval/thinksound \
--metric artifact_count \
--jsonCompatibility & versions
Manifest-backed recipe
- Integration
- metadata only
- Runner evidence
- pending
- Environment
- worldfoundry-unified-cu128
- Python
- 3.11
- CUDA
- CUDA 12.8
- PyTorch
- torch
- Source revision
- —
- Checkpoint revision
- —
- Runtime profile
- thinksound
- Pipeline binding
- thinksound
- Runner
- worldfoundry.pipeline
- Pipeline target
- worldfoundry.pipelines.video_official.pipeline_official_video:ThinkSoundPipeline
- Backend stage
- official_runtime_bridge
- Runtime status
- pending_checkpoint_and_in_tree_runtime
- Driver status
- compatible_unified_default
- Environment kind
- Unified environment
Install environment
The environment resolver reads the recorded profile and chooses the unified or dedicated environment shown here.
bash scripts/setup/model_env_install.sh --model thinksoundPackage constraints 4
torchtorchvisiontorchaudioxfuser
Conda packages 5
python=3.11pipffmpegsetuptoolswheel
Checkpoints & assets
Run the local check before allocating compute. Gated, private, and license fields below come directly from the checkpoint manifest.
worldfoundry-eval zoo model-download --model-id thinksound --check-local --jsonFunAudioLLM/ThinkSound
- Revision
- —
- License
- —
- Gated
- —
- Private
- —
Launch & outputs
The generated command uses the shared evaluation boundary and writes durable result manifests and artifacts.
worldfoundry-eval evaluate \
--mode model \
--model-id thinksound \
--model-runner worldfoundry:pipeline \
--model-manifest-dir worldfoundry/data/models/catalog \
--requests-path tmp/requests.jsonl \
--output-dir tmp/model_eval/thinksound \
--metric artifact_count \
--jsonInput contract
| Field | Recorded contract |
|---|---|
prompt | Required |
video | optional |
Artifact contract
| Artifact kind | Filename / path |
|---|---|
generated_audio | thinksound.wav |
Evidence & sources
Catalog integration, native-demo parity, and runner parity are independent records.
- Integration
- metadata only
- Runner evidence
- pending
- Native demo evidence
- Not recorded
- Validation imports
- torch, torchvision, numpy, transformers, diffusers, xfuser, worldfoundry
Recipe provenance
worldfoundry/data/models/catalog/video/thinksound.yamlGitHubhttps://huggingface.co/FunAudioLLM/ThinkSound


