Meta Introduces Muse Image and Muse Video: Advancing Media Generation with Agentic Tools and Compute Scaling
Meta 發表 Muse Image 與 Muse Video:首款導入 Agentic 機制與推理期算力擴展的媒體生成模型

Meta Superintelligence Labs introduced Muse Image and previewed Muse Video. Unlike traditional models, Muse Image acts as an agent, using search and coding tools (to render plots/QR codes) and self-refining its drafts via Chain of Thought (CoT). By scaling test-time compute across text reasoning and visual generation tokens, it achieves superior quality over simple Best-of-N sampling. Currently, Muse Image ranks No. 2 on Arena for text-to-image and editing, while Muse Video ranks No. 3.
Key points
Agentic Tool Integration
Integrates coding (for precise plots and QR codes) and web search to ground generated images in factual, real-world information.
Emergent Self-Refinement
The model reflects on its drafts within its chain of thought, choosing to make local edits or regenerate from scratch to achieve better results.
Test-Time Compute Scaling
Quality scales log-linearly with increased inference-time compute, combining text reasoning and visual generation tokens effectively.
Multi-Reference Composition
Seamlessly composes elements like people, clothing, style, and environments from multiple inline reference images.
How it works
Why it matters
This marks a shift from passive prompt-to-image mapping to active agentic generation. By enabling models to search, code, and self-correct before outputting the final image, Meta addresses major industry pain points: lack of factual grounding, difficulty in precise control, and weak iterative editing capabilities. It proves that LLM-style reasoning and compute scaling can successfully generalize to visual generation.
Who it affects
- AI Developer
- Content Creator
- Product Manager
- Enterprise Leader
How to use it
- 1Creating animated GIFs, web pages, or interactive visual games containing precise generated code.
- 2Composing multiple product and environment reference images to generate marketing assets for small businesses.
- 3Performing coherent, multi-turn, and local image edits and brainstorming on social platforms like Instagram.
Limitations & caveats
- Muse Video currently has performance gaps in audio-video synchronization and physically accurate fast motion.
- Although detector tools are in preview, the absolute robustness of Content Seal invisible watermarks under extreme transformations remains to be broadly proven.