Benchmark Hub
Find a benchmark, inspect its execution boundary, and open its manifest-backed run recipe.
chronomagic_scoreVerified evidenceVideo GenerationOfficial runtimeVerifiedHosted APIEvalCrafterevalcrafterIn-tree official runtime scores generated videos per promptevalcrafter_totalVerified evidenceVideo GenerationOfficial runtimeVerifiedLocal / no APIFETVfetvIn-tree official runtime scores generated videos per promptfetv_averageVerified evidenceVideo GenerationOfficial runtimeBounded Official Clipscore VerifiedLocal / no APIT2V-CompBencht2v-compbenchIn-tree official runtime scores generated videos per prompt; some dimensions may call a hosted APIt2v_compbench_averageVerified evidenceVideo GenerationOfficial runtimeBounded Official Generative Numeracy VerifiedHosted APIVBenchvbenchIn-tree official runtime scores generated videos per promptoverall_qualityVerified evidenceVideo GenerationOfficial runtimeVerifiedLocal / no APIVBench-2.0vbench-2.0In-tree official runtime scores generated videos per prompt; some dimensions may call a hosted APIvbench2_totalVerified evidenceVideo GenerationOfficial runtimeVerifiedHosted APIVBench++vbench-plus-plusIn-tree official runtime scores generated videos per prompt; some dimensions may call a hosted APIvbench_plus_plus_averageVerified evidenceVideo GenerationOfficial runtimeVerifiedHosted APIVideo-Benchvideo-benchIn-tree official runtime scores generated videos per prompt; some dimensions may call a hosted APIvideobench_averageVerified evidenceVideo GenerationOfficial runtimeVerifiedHosted APIVideoScorevideoscoreIn-tree official runtime scores generated videos per promptvideoscore_averageVerified evidenceVideo GenerationOfficial runtimeVerifiedLocal / no APIWorldBenchworldbenchIn-tree evaluator scores artifact bundles per the task protocolforeground_miouVerified evidenceWorld ModelsIn Tree Model Backed EvaluatorBounded Fixture VerifiedLocal / no APIWorldModelBenchworldmodelbenchIn-tree official runtime scores generated videos per promptworld_model_averageVerified evidenceWorld ModelsOfficial runtimeVerifiedLocal / no APIWorldScoreworldscoreIn-tree official runtime scores generated videos per promptworldscore_averageVerified evidenceWorld ModelsOfficial runtimeVerifiedLocal / no API4DWorldBench4dworldbenchIn-tree official runtime scores generated videos per promptfour_d_worldbench_averageIntegratedWorld ModelsOfficial runtimeNot recordedLocal / no APIAI2-THORai2thorClosed-loop simulator rollouts, or normalization of official rollout resultssuccess_rateIntegratedEmbodied AIClosed-loop simIntegration ReadyLocal / no APIAIGCBenchaigcbenchImports precomputed metrics and normalizes them into a scorecardaigcbench_averageIntegratedVideo GenerationArtifact importIn Tree Artifact Import ReadyLocal / no APICameraBenchcamerabenchIn-tree official runtime scores generated videos per prompt; some dimensions may call a hosted APIcamerabench_averageIntegratedVideo GenerationNot recordedPendingHosted APIDEVIL Dynamicsdevil-dynamicsIn-tree official runtime scores generated videos per prompt; some dimensions may call a hosted APIdevil_dynamics_averageIntegratedVideo GenerationOfficial runtimePendingHosted APIEWMBenchewmbenchIn-tree official runtime scores generated videos per promptewmbench_averageIntegratedWorld ModelsOfficial runtimePendingLocal / no APIGenAI-Benchgenai-benchIn-tree official runtime scores generated videos per promptgenai_bench_averageIntegratedVideo GenerationOfficial runtimePendingLocal / no APIIPV-Benchipv-benchIn-tree official runtime scores generated videos per promptipv_bench_averageIntegratedVideo GenerationOfficial runtimePendingLocal / no APIiWorld-Benchiworld-benchIn-tree official runtime scores generated videos per promptiworldbench_averageIntegratedWorld ModelsOfficial runtimeOfficial Runtime Ready Real Data Validation PendingLocal / no APIMemoBenchmemobenchIn-tree official runtime scores generated videos per promptmemobench_averageIntegratedVideo GenerationOfficial runtimePendingLocal / no APIMiraBenchmirabenchIn-tree official runtime scores generated videos per promptmirabench_averageIntegratedVideo GenerationOfficial runtimePendingLocal / no APIPhyEduVideophyeduvideoImports precomputed metrics and normalizes them into a scorecardphyeduvideo_averageIntegratedVideo GenerationArtifact importResult Importer ReadyLocal / no APIPhyFPS-Bench-Genphyfps-bench-genIn-tree official runtime scores generated videos per promptphyfps_bench_gen_averageIntegratedVideo GenerationOfficial runtimeOfficial Runtime ReadyLocal / no APIPhyGenBenchphygenbenchImports precomputed metrics and normalizes them into a scorecardphygenbench_averageIntegratedVideo GenerationArtifact importBounded Component ReadyLocal / no APIPhyGroundphygroundImports precomputed metrics and normalizes them into a scorecardphyground_overallIntegratedVideo GenerationArtifact importResult Importer ReadyLocal / no APIPhysics-IQ Originalphysics-iqIn-tree official runtime scores generated videos per promptphysics_iq_scoreIntegratedVideo GenerationOfficial runtimeIn Tree Raw Video EvaluatorLocal / no APIPhysics-IQ Verifiedphysics-iq-verifiedIn-tree official runtime scores generated videos per promptphysics_iq_verified_scoreIntegratedVideo GenerationOfficial runtimeIn Tree Raw Video EvaluatorLocal / no APIPhysVidBenchphysvidbenchImports precomputed metrics and normalizes them into a scorecardphysvidbench_averageIntegratedVideo GenerationArtifact importBounded Caption Qa ReadyLocal / no APIT2VSafetyBencht2v-safety-benchIn-tree official runtime scores generated videos per promptnsfw_averageIntegratedVideo GenerationOfficial runtimePendingLocal / no APIT2VWorldBencht2vworldbenchIn-tree official runtime scores generated videos per promptworld_knowledge_averageIntegratedWorld ModelsOfficial runtimePendingLocal / no APIVideoPhyvideophyIn-tree judge/VLM scores generated videos per promptvideophy_averageIntegratedVideo GenerationJudge runtimeIn Tree Official Runtime ReadyLocal / no APIVideoPhy2videophy2In-tree judge/VLM scores generated videos per promptjoint_scoreIntegratedVideo GenerationJudge runtimeIn Tree Official Runtime ReadyLocal / no APIVideoScience-Benchvideoscience-benchIn-tree official runtime scores generated videos per promptvideoscience_averageIntegratedVideo GenerationOfficial runtimePendingLocal / no APIVideoVersevideoverseIn-tree judge/VLM scores generated videos per prompt; some dimensions may call a hosted APIvideoverse_averageIntegratedVideo GenerationJudge runtimeOfficial Runtime ReadyHosted APIVisual Chronometervisual-chronometerIn-tree official runtime scores generated videos per promptvisual_chronometer_averageIntegratedVideo GenerationOfficial runtimeOfficial Runtime ReadyLocal / no APIVMBenchvmbenchIn-tree official runtime scores generated videos per promptvmbench_averageIntegratedVideo GenerationOfficial runtimePendingLocal / no APIWBenchwbenchIn-tree official runtime scores generated videos per promptwbench_averageIntegratedWorld ModelsOfficial runtimeIn Tree RuntimeLocal / no APIWorld-in-Worldworld-in-worldImports precomputed metrics and normalizes them into a scorecardworld_in_world_averageIntegratedWorld ModelsArtifact importResult Importer ReadyLocal / no APIWorldArenaworldarenaClosed-loop simulator rollouts, or normalization of official rollout resultsewm_scoreIntegratedWorld ModelsOfficial runtimeIn Tree Video Quality RuntimeLocal / no APIWRBenchwrbenchIn-tree official runtime scores generated videos per promptwrbench_averageIntegratedWorld ModelsOfficial runtimePendingLocal / no APIBEHAVIOR-1Kbehavior1kImports official or upstream results into a normalized scorecardsuccess_rateNormalizerEmbodied AIOfficial resultsNormalizer onlyLocal / no APIBridgeData V2bridgedata-v2Imports official or upstream results into a normalized scorecardsuccess_rateNormalizerEmbodied AIOfficial resultsNormalizer onlyLocal / no APICALVINcalvinImports official or upstream results into a normalized scorecardsuccess_rateNormalizerEmbodied AIOfficial resultsNormalizer onlyLocal / no APIKinetixkinetixImports official or upstream results into a normalized scorecardsuccess_rateNormalizerEmbodied AIOfficial resultsNormalizer onlyLocal / no APILIBEROliberoImports official or upstream results into a normalized scorecardsuccess_rateNormalizerEmbodied AIClosed-loop simNormalizer onlyLocal / no APILIBERO-Memlibero-memImports official or upstream results into a normalized scorecardsuccess_rateNormalizerEmbodied AIOfficial resultsNormalizer onlyLocal / no APIWhat a benchmark recipe records
Each detail page is generated from benchmark catalogs and task manifests rather than a hand-maintained link list. A recipe keeps these claims separate:
| Record | What it tells you |
|---|---|
| Catalog | Stable benchmark ID, domains, aliases, official sources, and integration status. |
| Evaluation surface | Primary metrics, artifact layout, judge requirements, and supported execution path. |
| Runtime | In-tree evaluator, external result importer, hosted judge, or simulator boundary. |
| Verification | The strongest checked runner or normalizer evidence currently recorded in the manifest. |
| Eligibility | Whether available evidence is sufficient for leaderboard-facing claims. |
This hub does not publish leaderboard rankings. Missing or pending verification remains visible instead of being promoted to a stronger support claim.
Before installing benchmark-specific packages, read the runtime environment matrix. It separates the unified environment, managed dedicated environments, hosted judge services, and dimension-specific native/checkpoint stacks.
How to run one benchmark
Open the benchmark recipe first. Prompt files, generated-artifact layouts, checkpoint names, judge models, and official result formats are benchmark-specific.
Import official result files
worldfoundry-eval zoo benchmark-run \
--benchmark-id <benchmark-id> \
--mode official-validation \
--official-results-path <official-result-file-or-dir> \
--generated-artifact-dir <generated-artifact-dir> \
--output-dir tmp/<benchmark-id>/official-validation \
--jsonScore generated artifacts
Use official-run only when the benchmark recipe says the in-tree runtime can recompute scores from generated artifacts and the required checkpoints are staged.
worldfoundry-eval zoo benchmark-run \
--benchmark-id <benchmark-id> \
--mode official-run \
--generated-artifact-dir <generated-artifact-dir> \
--output-dir tmp/<benchmark-id>/official-run \
--jsonTo generate artifacts with a registered model before scoring:
worldfoundry-eval run \
--benchmark <benchmark-id> \
--model <model-id> \
--mode official-run \
--output-dir tmp/<benchmark-id>/<model-id> \
--json





































