Benchmark Hub

Find a benchmark, inspect its execution boundary, and open its manifest-backed run recipe.

On this page
64 benchmarkscatalog data from repository manifests

64 benchmarks

BenchmarkPrimary metricsStatusExecutionVerificationJudge
ChronoMagic-Benchchronomagic-benchIn-tree official runtime scores generated videos per prompt; some dimensions may call a hosted APIchronomagic_scoreVerified evidenceVideo GenerationOfficial runtimeVerifiedHosted APIEvalCrafterevalcrafterIn-tree official runtime scores generated videos per promptevalcrafter_totalVerified evidenceVideo GenerationOfficial runtimeVerifiedLocal / no APIFETVfetvIn-tree official runtime scores generated videos per promptfetv_averageVerified evidenceVideo GenerationOfficial runtimeBounded Official Clipscore VerifiedLocal / no APIT2V-CompBencht2v-compbenchIn-tree official runtime scores generated videos per prompt; some dimensions may call a hosted APIt2v_compbench_averageVerified evidenceVideo GenerationOfficial runtimeBounded Official Generative Numeracy VerifiedHosted APIVBenchvbenchIn-tree official runtime scores generated videos per promptoverall_qualityVerified evidenceVideo GenerationOfficial runtimeVerifiedLocal / no APIVBench-2.0vbench-2.0In-tree official runtime scores generated videos per prompt; some dimensions may call a hosted APIvbench2_totalVerified evidenceVideo GenerationOfficial runtimeVerifiedHosted APIVBench++vbench-plus-plusIn-tree official runtime scores generated videos per prompt; some dimensions may call a hosted APIvbench_plus_plus_averageVerified evidenceVideo GenerationOfficial runtimeVerifiedHosted APIVideo-Benchvideo-benchIn-tree official runtime scores generated videos per prompt; some dimensions may call a hosted APIvideobench_averageVerified evidenceVideo GenerationOfficial runtimeVerifiedHosted APIVideoScorevideoscoreIn-tree official runtime scores generated videos per promptvideoscore_averageVerified evidenceVideo GenerationOfficial runtimeVerifiedLocal / no APIWorldBenchworldbenchIn-tree evaluator scores artifact bundles per the task protocolforeground_miouVerified evidenceWorld ModelsIn Tree Model Backed EvaluatorBounded Fixture VerifiedLocal / no APIWorldModelBenchworldmodelbenchIn-tree official runtime scores generated videos per promptworld_model_averageVerified evidenceWorld ModelsOfficial runtimeVerifiedLocal / no APIWorldScoreworldscoreIn-tree official runtime scores generated videos per promptworldscore_averageVerified evidenceWorld ModelsOfficial runtimeVerifiedLocal / no API4DWorldBench4dworldbenchIn-tree official runtime scores generated videos per promptfour_d_worldbench_averageIntegratedWorld ModelsOfficial runtimeNot recordedLocal / no APIAI2-THORai2thorClosed-loop simulator rollouts, or normalization of official rollout resultssuccess_rateIntegratedEmbodied AIClosed-loop simIntegration ReadyLocal / no APIAIGCBenchaigcbenchImports precomputed metrics and normalizes them into a scorecardaigcbench_averageIntegratedVideo GenerationArtifact importIn Tree Artifact Import ReadyLocal / no APICameraBenchcamerabenchIn-tree official runtime scores generated videos per prompt; some dimensions may call a hosted APIcamerabench_averageIntegratedVideo GenerationNot recordedPendingHosted APIDEVIL Dynamicsdevil-dynamicsIn-tree official runtime scores generated videos per prompt; some dimensions may call a hosted APIdevil_dynamics_averageIntegratedVideo GenerationOfficial runtimePendingHosted APIEWMBenchewmbenchIn-tree official runtime scores generated videos per promptewmbench_averageIntegratedWorld ModelsOfficial runtimePendingLocal / no APIGenAI-Benchgenai-benchIn-tree official runtime scores generated videos per promptgenai_bench_averageIntegratedVideo GenerationOfficial runtimePendingLocal / no APIIPV-Benchipv-benchIn-tree official runtime scores generated videos per promptipv_bench_averageIntegratedVideo GenerationOfficial runtimePendingLocal / no APIiWorld-Benchiworld-benchIn-tree official runtime scores generated videos per promptiworldbench_averageIntegratedWorld ModelsOfficial runtimeOfficial Runtime Ready Real Data Validation PendingLocal / no APIMemoBenchmemobenchIn-tree official runtime scores generated videos per promptmemobench_averageIntegratedVideo GenerationOfficial runtimePendingLocal / no APIMiraBenchmirabenchIn-tree official runtime scores generated videos per promptmirabench_averageIntegratedVideo GenerationOfficial runtimePendingLocal / no APIPhyEduVideophyeduvideoImports precomputed metrics and normalizes them into a scorecardphyeduvideo_averageIntegratedVideo GenerationArtifact importResult Importer ReadyLocal / no APIPhyFPS-Bench-Genphyfps-bench-genIn-tree official runtime scores generated videos per promptphyfps_bench_gen_averageIntegratedVideo GenerationOfficial runtimeOfficial Runtime ReadyLocal / no APIPhyGenBenchphygenbenchImports precomputed metrics and normalizes them into a scorecardphygenbench_averageIntegratedVideo GenerationArtifact importBounded Component ReadyLocal / no APIPhyGroundphygroundImports precomputed metrics and normalizes them into a scorecardphyground_overallIntegratedVideo GenerationArtifact importResult Importer ReadyLocal / no APIPhysics-IQ Originalphysics-iqIn-tree official runtime scores generated videos per promptphysics_iq_scoreIntegratedVideo GenerationOfficial runtimeIn Tree Raw Video EvaluatorLocal / no APIPhysics-IQ Verifiedphysics-iq-verifiedIn-tree official runtime scores generated videos per promptphysics_iq_verified_scoreIntegratedVideo GenerationOfficial runtimeIn Tree Raw Video EvaluatorLocal / no APIPhysVidBenchphysvidbenchImports precomputed metrics and normalizes them into a scorecardphysvidbench_averageIntegratedVideo GenerationArtifact importBounded Caption Qa ReadyLocal / no APIT2VSafetyBencht2v-safety-benchIn-tree official runtime scores generated videos per promptnsfw_averageIntegratedVideo GenerationOfficial runtimePendingLocal / no APIT2VWorldBencht2vworldbenchIn-tree official runtime scores generated videos per promptworld_knowledge_averageIntegratedWorld ModelsOfficial runtimePendingLocal / no APIVideoPhyvideophyIn-tree judge/VLM scores generated videos per promptvideophy_averageIntegratedVideo GenerationJudge runtimeIn Tree Official Runtime ReadyLocal / no APIVideoPhy2videophy2In-tree judge/VLM scores generated videos per promptjoint_scoreIntegratedVideo GenerationJudge runtimeIn Tree Official Runtime ReadyLocal / no APIVideoScience-Benchvideoscience-benchIn-tree official runtime scores generated videos per promptvideoscience_averageIntegratedVideo GenerationOfficial runtimePendingLocal / no APIVideoVersevideoverseIn-tree judge/VLM scores generated videos per prompt; some dimensions may call a hosted APIvideoverse_averageIntegratedVideo GenerationJudge runtimeOfficial Runtime ReadyHosted APIVisual Chronometervisual-chronometerIn-tree official runtime scores generated videos per promptvisual_chronometer_averageIntegratedVideo GenerationOfficial runtimeOfficial Runtime ReadyLocal / no APIVMBenchvmbenchIn-tree official runtime scores generated videos per promptvmbench_averageIntegratedVideo GenerationOfficial runtimePendingLocal / no APIWBenchwbenchIn-tree official runtime scores generated videos per promptwbench_averageIntegratedWorld ModelsOfficial runtimeIn Tree RuntimeLocal / no APIWorld-in-Worldworld-in-worldImports precomputed metrics and normalizes them into a scorecardworld_in_world_averageIntegratedWorld ModelsArtifact importResult Importer ReadyLocal / no APIWorldArenaworldarenaClosed-loop simulator rollouts, or normalization of official rollout resultsewm_scoreIntegratedWorld ModelsOfficial runtimeIn Tree Video Quality RuntimeLocal / no APIWRBenchwrbenchIn-tree official runtime scores generated videos per promptwrbench_averageIntegratedWorld ModelsOfficial runtimePendingLocal / no APIBEHAVIOR-1Kbehavior1kImports official or upstream results into a normalized scorecardsuccess_rateNormalizerEmbodied AIOfficial resultsNormalizer onlyLocal / no APIBridgeData V2bridgedata-v2Imports official or upstream results into a normalized scorecardsuccess_rateNormalizerEmbodied AIOfficial resultsNormalizer onlyLocal / no APICALVINcalvinImports official or upstream results into a normalized scorecardsuccess_rateNormalizerEmbodied AIOfficial resultsNormalizer onlyLocal / no APIKinetixkinetixImports official or upstream results into a normalized scorecardsuccess_rateNormalizerEmbodied AIOfficial resultsNormalizer onlyLocal / no APILIBEROliberoImports official or upstream results into a normalized scorecardsuccess_rateNormalizerEmbodied AIClosed-loop simNormalizer onlyLocal / no APILIBERO-Memlibero-memImports official or upstream results into a normalized scorecardsuccess_rateNormalizerEmbodied AIOfficial resultsNormalizer onlyLocal / no API

What a benchmark recipe records

Each detail page is generated from benchmark catalogs and task manifests rather than a hand-maintained link list. A recipe keeps these claims separate:

RecordWhat it tells you
CatalogStable benchmark ID, domains, aliases, official sources, and integration status.
Evaluation surfacePrimary metrics, artifact layout, judge requirements, and supported execution path.
RuntimeIn-tree evaluator, external result importer, hosted judge, or simulator boundary.
VerificationThe strongest checked runner or normalizer evidence currently recorded in the manifest.
EligibilityWhether available evidence is sufficient for leaderboard-facing claims.

This hub does not publish leaderboard rankings. Missing or pending verification remains visible instead of being promoted to a stronger support claim.

Before installing benchmark-specific packages, read the runtime environment matrix. It separates the unified environment, managed dedicated environments, hosted judge services, and dimension-specific native/checkpoint stacks.

How to run one benchmark

Open the benchmark recipe first. Prompt files, generated-artifact layouts, checkpoint names, judge models, and official result formats are benchmark-specific.

Import official result files

worldfoundry-eval zoo benchmark-run \
  --benchmark-id <benchmark-id> \
  --mode official-validation \
  --official-results-path <official-result-file-or-dir> \
  --generated-artifact-dir <generated-artifact-dir> \
  --output-dir tmp/<benchmark-id>/official-validation \
  --json

Score generated artifacts

Use official-run only when the benchmark recipe says the in-tree runtime can recompute scores from generated artifacts and the required checkpoints are staged.

worldfoundry-eval zoo benchmark-run \
  --benchmark-id <benchmark-id> \
  --mode official-run \
  --generated-artifact-dir <generated-artifact-dir> \
  --output-dir tmp/<benchmark-id>/official-run \
  --json

To generate artifacts with a registered model before scoring:

worldfoundry-eval run \
  --benchmark <benchmark-id> \
  --model <model-id> \
  --mode official-run \
  --output-dir tmp/<benchmark-id>/<model-id> \
  --json