Text-to-video model from scratch (2 brothers, 2 years, 2B params)
What they did
Two brothers trained a 2B-parameter text-to-video model from scratch. Their write-up shows both good and bad samples.
How it went
The authors say it took 2 years. The HN post drew 158 points.
Worth knowing
- Price: unknown
- Status: sourced. We have not run this ourselves yet (no "Tested" badge).
- Quality bar (computed by code): 4/5 (primary source ✓, reproducible ✗, new ✓, example output ✓, safe ✓)
The best new AI video workflows and prompts, weekly.
Related
▶ YouTube demoVeo (Google DeepMind)
Google DeepMind's page for the Veo video model family, with capability notes and sample clips straight from the lab.
▶ demo at sourceVeo 3 on fal
Veo 3 served through fal's API, with text-to-video and optional generated audio. Price depends on whether audio is on.
▶ demo at sourceVeo 3 Fast on fal
A faster, cheaper Veo 3 variant on fal. It's a sensible default for drafts before you spend on the full model.
▶ YouTube demoVeo on Gemini API docs
Google's developer docs for generating video through the Gemini API, including when to pick Veo 3.1 versus faster editing-oriented models.
Google Flow
Google's filmmaking app built on Veo and its image models. You assemble shots and scenes in one place instead of generating clips one by one.
▶ demo at sourceReplicate video collection
Replicate's curated collection of text-to-video models you can run by API, an alternative host to fal.
