Code for the Goku video generation foundation models, a CVPR 2025 highlight paper.
Worth knowing
Price: Open source repo (license unknown)
Status: sourced. We have not run this ourselves yet (no "Tested" badge).
Quality bar (computed by code): 4/5 (primary source ✓, reproducible ✗, new ✓, example output ✓, safe ✓)
Get Atlas Dispatch
The best new AI video workflows and prompts, weekly.
Weekly: 5 new prompts and workflows. Unsubscribe anytime.
From the source
Setup. This is a research repo for the paper 'Goku: Flow Based Video Generative Foundation Models' (HKU and ByteDance, CVPR 2025 highlight, arXiv 2502.04896). The README is a paper overview and doesn't document a public inference setup or weights download.
Inputs. Per the paper overview: text-to-video, image-to-video and text-to-image, using rectified-flow Transformers trained jointly on images and video.
What the source says about results. The authors report 84.85 on VBench text-to-video (No. 2 on VBench as of 2024-10-07, per the README), 0.76 on GenEval and 83.65 on DPG-Bench for text-to-image. These are author-reported benchmark numbers.
Connect Claude, ChatGPT, Cursor, Gemini or Grok Bot to Atlas, and your agent can plan shots, pick models and price them from the data on this site. Free. No sign-up or API key.