daVinci-MagiHuman (GAIR & Sand.ai)
What they did
A 15B single-stream transformer that generates human-centric video and speech audio together in six languages. The authors report a 5-second 1080p clip in 38 seconds on one H100, and pairwise human-eval wins over Ovi 1.1 and LTX 2.3. The base, distilled and super-resolution models are all released.
Worth knowing
- Price: Open weights on Hugging Face (apache-2.0)
- License: apache-2.0
- Status: sourced. We have not run this ourselves yet (no "Tested" badge).
- Quality bar (computed by code): 4/5 (primary source ✓, reproducible ✗, new ✓, example output ✓, safe ✓)
Get Atlas Dispatch
The best new AI video workflows and prompts, weekly.
Weekly: 5 new prompts and workflows. Unsubscribe anytime.
