Guides / Prism: native 2K joint video-audio model

model

Prism: native 2K joint video-audio model

AI-generated illustration, not actual output from this listing

What they did

A research model from Fudan, Tencent Hunyuan and Zhejiang University that natively trains joint video and audio generation at 2K. Its dynamic sparse attention splits each clip into zones and sizes the attention blocks by how much the image and the audio change there.

Worth knowing

  • Price: Open weights on Hugging Face (mit)
  • License: mit
  • Status: sourced. We have not run this ourselves yet (no "Tested" badge).
  • Quality bar (computed by code): 4/5 (primary source ✓, reproducible ✗, new ✓, example output ✓, safe ✓)
Get Atlas Dispatch

The best new AI video workflows and prompts, weekly.

Weekly: 5 new prompts and workflows. Unsubscribe anytime.
For your agent

Give your AI agent a video director

Connect Claude, ChatGPT, Cursor, Gemini or Grok Bot to Atlas, and your agent can plan shots, pick models and price them from the data on this site. Free. No sign-up or API key.

Step 1 of 4

Which tool do you use?