Qualitative samples over training iterations. The latent space model consistently demonstrates faster image structure learning and superior final image quality in pre-training progress.
Evaluation curves over training steps. The model trained in the latent space consistently achieve superior results (measured by GenEval~geneval and DPG~dpg) and faster convergence compared to those trained in the pixel space.
SQuad: Sub-Quadratic Attention Distillation for Efficient Video Generation | Qualcomm AI Research | arXiv:2608.16585
前序问题:视频 DiT 的算力几乎全被自注意力吃掉,而它的成本随 latent token 数 n 呈 O(n²) 增长。视频的 token 数天生就大——Wan 2.2 5B 在 81 帧 704×1280 下 n 就到 18480——于是这一项直接主导了运行时与显存,成为分辨率和时长的硬天花板。学界的常规解法是把满 softmax 的 QK^T 换成更便宜的核:线性注意力 O(n) 或低秩近似 O(nk)。这些替身确实便宜,但几乎从没真正恢复原始注意力的表达力,留下一道顽固的质量鸿沟。问题的本质是效率与表达力之间的取舍被推到了极端——要么二次成本,要么掉点。
本文贡献:SQuad 选了取舍曲线上一个被忽略的中间点:O(n√n)。做法是把注意力拆成两段——先在局部窗口内做精细注意力,再以窗口为单位做全局注意力,两段复合出 n√n 的复杂度,天然平衡效率与表达力。更务实的是训练路径:从零训一个自己的视频 DiT 成本高到不现实,所以他们直接把预训练的满 softmax DiT 蒸馏成 SQuad-Attention 版本,分两阶段进行——先做 Flow-Matching 有监督微调对齐行为,再用改进的分布匹配蒸馏 DMD2,后者额外让采样本身也变少步。也就是说这套方法同时压了单步成本和总步数两个乘数。
% attention. Left: full softmax Self-Attention at $\mathcal{O}(n^2)$ complexity. Right: factorizes it into a local pass within $\mathcal{O}(\sqrt{n})$ windows followed by a global pass across the windows, giving a full receptive field at $\mathcal{O}(n\sqrt{n})$ complexity with a true softmax throughout.%
Overview. Video latent frames are embedded as tiles of a virtual image grid and edited by the image editing transformer; only the two projections (warm-started from the editor's own patchify/unpatchify layers) and the transformer weights (LoRA or full) are trained.
实验效果:从 Qwen-Image-Edit 出发,无需任何视频生成预训练即可完成指令式视频编辑。流程图显示除 In-Proj 与 Out-Proj 两个轻量投影和 DiT 本体外,Wan 2.1 VAE 编解码与 Wan 2.2 增强器全程冻结,可训练部分被压到极小。定性结果覆盖全局风格迁移与局部主体替换等任务,人物换成机器人这类大幅编辑仍保持了动作与场景的连贯。
批判点评:这篇最漂亮的地方不是结果而是论证方式:那条零训练观察链把「为什么这样可行」讲成了一个可验证的递进过程,而不是先做完再编故事。把视频摊成联络表这个操作听起来近乎取巧,但作者恰恰用它揭示了一个有意思的事实——强图像编辑模型内部已经隐含了相当程度的跨帧一致性能力,只是从没被这样调用过。投影层从 patchify 热启动的设计也很讲究,保证了初始化即合理。但取巧的代价必须点明。第一,把帧铺成 tile 意味着帧间关系只能靠二维位置编码隐式表达,模型没有任何显式的时间轴概念,短片段可能侥幸过关,帧数一多、运动一快,这种隐式关联大概率崩塌——而 tile 数量还直接受图像编辑器分辨率上限约束,这是个硬天花板,摘要没说能处理多少帧。第二,需要 Wan 2.2 做「可选时序增强」这件事本身就是自证:如果时序一致性真的够好,为什么还要额外几步去噪来补?这个「可选」组件在多少比例的案例里其实是必需的,是关键信息。第三,训练数据是公开的 Ditto-1M,编辑类型分布决定了能力边界,超出分布的指令表现如何未知。第四,全文只报了定性结果和「结果表明」,没有与专门视频编辑模型的定量对比——考虑到这是一份 report 而非完整论文,可以理解,但也意味着现在还无法判断它离 SOTA 有多远。
Overview of PixRestore. A frozen vision encoder extracts multi-layer features from the LQ image, and an adaptive layer router predicts per-layer weights to fuse them into a single conditioning feature. The LQ image and the noisy state $x_t$ are patchified, then processed by $N$ DiT blocks, and finally decoded into the HQ output.
PixRestore achieves the best overall performance in terms of restoration quality, model size, and inference speed. Top-left: radar charts comparing PSNR, LPIPS, and degradation removal performance across eight degradation types. Top-right: GFLOPs versus inference latency, where the bubble size denotes the number of parameters. Bottom: visual comparisons on eight restoration tasks. With only about 50M parameters and single-step inference, PixRestore achieves the best overall restoration quality while being the most efficient among diffusion-based methods.
Equilibrium Forcing Overview. is trained with a simple unmodulated equilibrium denoising objective (left), which can be shaped at inference time through the $\eta$ warp (middle). The sampling landscape here is shaped to be an attractor, with the magnitude decreasing as the sample descends and converges (right).
Overview of the improved synthesis framework. Stage 1: Library Construction leveraging LLM world knowledge. Stage 2: Semantic Matching and Instruction Generation including VQA checklists. Stage 3: Image Synthesis using various editing models. Stage 4: Instance Specific Verification using VQA.
(a) Edit Concept Scaling. Left: The previous paradigm is restricted by coarse categories and limited diversity. Right: Our approach scales up to 1,000+ fine-grained concepts to ensure a balanced and rich distribution. (b) Dense Supervision. Left: Conventional training relies on single edit pairs with sparse supervision signals. Right: Our composite edit strategy provides dense supervision, enhancing training efficiency.
Overview of the proposed synchronization-aware acceleration framework. Cross-modal interactions between audio and video branches are used to guide protected sparse attention and cache reuse.
Qualitative comparison of dense inference, unprotected acceleration, and our protected acceleration. The protected computation path better preserves salient visual structures and reduces local artifacts under acceleration.
批判点评:这篇的问题意识比方法本身更值得称道:它指出了一个正在发生但少有人明说的隐患——加速方法的迁移往往是模态盲视的,在纯视频上验证过的稀疏化和缓存策略被直接用到音视频模型上,而评测指标里如果没有同步性这一项,退化会完全隐形。「同步性集中在少数 token 上」这个观察也很自然地导出了保护式稀疏化,逻辑闭环。Δ_sync 取四路信号最大值的设计偏保守但合理——宁可多算几步也不冒同步崩掉的风险。不过整份摘要在效果层面几乎是零信息量:只说「提升效率」「保持质量」,没有任何加速倍率、没有基线名称、没有具体指标数值,这在一篇以效率为核心卖点的论文里是硬伤——读者无法判断它是快了 1.5 倍还是 5 倍,而这个区间决定了方法有没有实用价值。其次,Top-k 保护的 k 如何选择、保护比例与加速比之间的权衡曲线长什么样,是这类方法最核心的超参问题,摘要完全没碰。第三,方法建立在「能拿到双向交叉注意力图」的前提上,这对某些架构(如把音频当额外 token 拼进同一序列做自注意力的设计)并不成立,适用范围有限制。第四,摘要文本里出现了多处 this http URL 的替换残留,说明投稿时的自动化处理出了问题,虽然不影响技术内容,但反映出打磨程度。第五,与 SQuad 那种把加速做成蒸馏目标的思路相比,这里仍是训练无关的推理期启发式,上限受限于原模型。
Comparison with baseline. Evaluated across a diverse set of metrics for robust assessment, performs on par with standard flow matching and outperforms it in some cases.
Overview of CineDub under the ICHC paradigm. The holistic visual condition $\mathbf{c}_v$ (left), extracted by SynchFormer (Sec.~sec:visual_condition), encodes both event-level audio-visual correspondences and fine-grained lip-sync alignment. To address the speaker assignment ambiguity in $\mathbf{c}_v$, the semantic-bundled transcription $\mathbf{c}_t$ (right) provides per-segment speaker-utterance grounding cues, implicitly coupling with $\mathbf{c}_v$ via multi-conditional training (Sec.~sec:text_condition). CineDub adopts a unified DiT backbone that supports both joint video-to-speech-and-audio generation and each subtask. $\mathbf{c}_t$ and $\mathbf{c}_a$ are routed through decoupled textual branches to prevent cross-prompt interference, with learnable meta-tokens replacing inactive branches during single-task inference (Sec.~sec:decoupled_attn).
实验效果:实现了从未裁剪视频直接进行多说话人多轮对话配音,并扩展到语音与音效的联合生成。同时释出 CineDub-Multi(多说话人对话配音)等两个真实场景基准。论文已被 ACM MM 2026 接收。
Visualization of SynchFormer self-attention maps (magenta: active speaker; white: inactive speaker). (a)~Single-speaker setting: attention concentrates on the lip region. (b)~Two-speaker setting: attention dynamically shifts to the active speaker across turns. (c)~Failure cases: attention drifts to a non-speaking face during overlapping speech or becomes ambiguous at shot boundaries.
批判点评:「不需要人脸裁剪和说话人日志」是这篇最实在的卖点。做过配音管线的人都知道,那些预处理步骤是主要的故障源和数据规模瓶颈,能去掉它们,工程价值立刻显现。用带 / 标记的语义捆绑转写格式来消解说话人归属歧义,是个巧妙的取舍——把「谁在说」这个视觉难题部分转移到了文本侧的结构化表示上,绕开了视觉说话人定位的老大难。三支冻结编码器分工也清晰:Synchformer 管密集时序、CLIP 管全局语义、Gemma-T5 管文本,各司其职且都不训练,成本可控。ACM MM 2026 接收提供了一定的同行评议背书。但几点需要审视。第一,转写格式承担了消歧的主要负担,意味着推理时需要一份带说话人切分标记的高质量转写——这份输入从哪来?如果需要人工标注或依赖 ASR 加日志,那么「不需要说话人日志」只是把它从视觉管线挪到了文本管线,端到端的依赖并没有真正消除,摘要在这一点上表述得偏乐观。第二,同时生成语音与音效带来的「跨提示词干扰」用解耦文本分支来解决,但音效与语音在时间上天然重叠(说话时门在响),这种物理层面的耦合能否靠表示层解耦干净,值得怀疑。第三,两个基准都由作者团队自建,与训练数据的分布关系未说明,同分布优势难以排除。第四,摘要没有任何定量结果——没有同步性指标、没有音质指标、没有与层级式方法的对比数字,一篇声称解决多说话人歧义的论文却不给说话人归属准确率,是明显的缺口。第五,环境音到语言的课程学习是为缓解子任务退化而引入的,说明联合生成确实存在能力互斥,那么相比两个独立模型分别生成再混音,联合方案的净收益有多大并不明确。
评论 (0)