Video · video-to-audio · audio-generation

ThinkSound

Worldfoundry Native

Hidden from Studio until the infer-only runtime and ThinkSound diffusion checkpoint are both available in-tree/Hugging Face cache.

Metadata Only统一环境thinksound

构建运行命令

选择仓库中已记录的 variant,然后复制对应的准备、检查或运行命令。

worldfoundry.pipeline
任务video-to-audio
环境worldfoundry-unified-cu128
设备CUDA 12.8
worldfoundry-eval evaluate \
  --mode model \
  --model-id thinksound \
  --model-runner worldfoundry:pipeline \
  --model-manifest-dir worldfoundry/data/models/catalog \
  --requests-path tmp/requests.jsonl \
  --output-dir tmp/model_eval/thinksound \
  --metric artifact_count \
  --json
01

兼容性与版本

由 Manifest 生成

集成状态
metadata only
Runner 证据
pending
环境
worldfoundry-unified-cu128
Python
3.11
CUDA
CUDA 12.8
PyTorch
torch
源码 revision
Checkpoint revision
Runtime profile
thinksound
Pipeline binding
thinksound
Runner
worldfoundry.pipeline
Pipeline target
worldfoundry.pipelines.video_official.pipeline_official_video:ThinkSoundPipeline
Backend stage
official_runtime_bridge
Runtime 状态
pending_checkpoint_and_in_tree_runtime
Driver 状态
compatible_unified_default
环境类型
统一环境
02

安装环境

环境解析器会读取已记录 profile,并选择此处显示的统一或独立环境。

bash scripts/setup/model_env_install.sh --model thinksound
依赖版本约束 4
  • torch
  • torchvision
  • torchaudio
  • xfuser
Conda 依赖 5
  • python=3.11
  • pip
  • ffmpeg
  • setuptools
  • wheel
03

Checkpoint 与资产

分配算力前先做本地检查;gated、private 与 license 字段直接来自 checkpoint manifest。

worldfoundry-eval zoo model-download --model-id thinksound --check-local --json
FunAudioLLM/ThinkSound
Revision
License
Gated
Private
04

运行与输出

生成的命令通过共享 evaluation 边界运行,并持久保存结果 manifest 与 artifact。

worldfoundry-eval evaluate \
  --mode model \
  --model-id thinksound \
  --model-runner worldfoundry:pipeline \
  --model-manifest-dir worldfoundry/data/models/catalog \
  --requests-path tmp/requests.jsonl \
  --output-dir tmp/model_eval/thinksound \
  --metric artifact_count \
  --json

输入契约

字段已记录契约
promptRequired
videooptional

Artifact 契约

Artifact 类型文件名 / 路径
generated_audiothinksound.wav
05

证据与来源

Catalog 集成、原生 demo parity 与 runner parity 是三条独立记录。

集成状态
metadata only
Runner 证据
pending
原生 Demo 证据
未记录
验证 Imports
torch, torchvision, numpy, transformers, diffusers, xfuser, worldfoundry
配方溯源
worldfoundry/data/models/catalog/video/thinksound.yamlGitHubhttps://huggingface.co/FunAudioLLM/ThinkSound