Prism: native 2K joint video-audio model
What they did
A research model from Fudan, Tencent Hunyuan and Zhejiang University that natively trains joint video and audio generation at 2K. Its dynamic sparse attention splits each clip into zones and sizes the attention blocks by how much the image and the audio change there.
Worth knowing
- Price: Open weights on Hugging Face (mit)
- License: mit
- Status: sourced. We have not run this ourselves yet (no "Tested" badge).
- Quality bar (computed by code): 4/5 (primary source ✓, reproducible ✗, new ✓, example output ✓, safe ✓)
Get Atlas Dispatch
The best new AI video workflows and prompts, weekly.
Weekly: 5 new prompts and workflows. Unsubscribe anytime.
