与前序工作的本质区别: 混合策略:学生继承教师轨迹点作锚点,再向自身策略局部精修 K 步,最后在自生成 roll-out 上做速度级监督
1.2 方法原理
Introduction to HPSD. (a) Off-policy SFT supervises on fixed teacher endpoints, which drift from the evolving student policy. (b) On-policy distillation queries the teacher at student-visited states, yet suffers from condition-state mismatch in TI2V models. (c) HPSD anchors the student on the teacher trajectory and supervises roll-outs evolved under its own policy. (d) HPSD substantially boosts the base generation quality of TI2V models (WAN-2.2-TI2V-5B in this figure), especially in cinematic lighting, fine details, and balanced composition. Prompts in the Appendix.
Overview of HPSD. (a) Offline Stage: Privileged conditions (enhanced prompt and first frame) are synthesized using auxiliary models. (b) Online Stage: The student evolves hybrid-policy sub-trajectories starting from the teacher's off-policy anchor states. Distillation on these states provides the student with precise policy correction anchored by the teacher's privileged content priors. Varying gray shading denotes the noise level over time, while solid white indicates clean frames.
Context-Matched Distillation. A causal student first generates an autoregressive rollout. Self-Forcing scores its complete noised rollout with a bidirectional teacher, allowing the score of a target to depend on future blocks. CMD instead uses a causal teacher: Base CMD conditions on preceding noised DMD blocks, Prefix Scoring uses the cached clean student prefix that produced the target, and Prefix Corruption perturbs this prefix to stabilize supervision when early rollouts are unreliable.
Efficient Prefix Scoring. The causal student first generates an on-policy rollout and caches the clean prefix used to produce each block. Corrupted prefixes and their corresponding noised DMD targets are then packed under a block-causal attention mask, allowing all targets to be scored in parallel. Finally, the real and fake scores are evaluated under the same prefix--target context to form the context-matched DMD update.
在短视频和长视频基准上,CMD 在自回归方法中取得 SOTA 综合性能,同时对时变相机控制的遵循度有大幅改善(substantially improved adherence to time-varying camera controls)。论文包含 chunk1 vs chunk4 的用户研究偏好对比。相机控制这一项的提升尤其能佐证方法动机——正是因为原来的双向教师会泄漏未来的控制信号,学生才学不好「只根据当前控制做出反应」这件事。
Conceptual comparison in the multi-task setting. Left: Standard OPD regresses a shared student toward task-specific teacher velocities, encouraging interpolation among the teachers. Right: DreOPD uses a shared degraded reference to construct targets that extrapolate through each task-specific teacher. Joint regression toward these targets moves the student beyond each teacher along the corresponding teacher-reference directions.
稀疏视图 3D 高斯泼溅(3DGS)修复的常规思路是不断堆更复杂的修复架构,但这篇论文指出了一个更根本的共性缺陷:基于扩散先验的方法,其监督发生在独立加噪的状态上,而这些状态并不覆盖推理时真正到达的状态。在稀疏视图场景下这个问题被放大——欠约束的几何从一开始就让去噪产生偏差,而这些偏差沿 rollout 逐步累积(compound along the rollout)。这正是本期专题的理论根基:off-policy 监督位置错了,误差就永远得不到纠正。
Qualitative comparison on DL3DV. GSFixer blurs reflective and under-observed regions (zoomed), whereas TRACE-GS restores sharper detail closer to the GT.
Overview of FlowErase-OPD. (a) illustrates the framework of our approach. Each mini-epoch we input a prompt contain a specific target concept to the student, corresponding teachers, and the anchor teacher. Then we compute the KL divergence and employ ARC to calculate the final loss and update the weight of each teacher models for the current mini-epoch. (b) details the Adaptive Retention Control. RAC dynamically adjusts the mini-epoch occupancy ratios between each concept erasure teachers and the anchor teacher based on the erasure efficacy for each individual concept.
从框架图可以看清 FlowErase-OPD 的结构:左侧是共享文本编码器,学生是一个带 LoRA 适配器的流匹配主干(Double Stream Blocks + Single Stream Blocks);中间是 n 个概念擦除教师(各自也是 LoRA 适配器,分别对应「梵高星空」「裸体」等概念);下方是锚点教师(以「a child playing soccer」这类正常提示词为条件),负责界定不该被破坏的生成行为。学生输出 μ_θ 与各教师输出 μ_φ 逐一计算 KL 散度,得到 ℓ_1..ℓ_n 与 ℓ_anchor。右半部分是 Adaptive Retention Control 的计算逻辑:每个擦除损失乘以权重 w_i、乘以采样概率 si、再乘以 (1-ρ),锚点损失乘以 w{n+1} 与 ρ,最后取期望合成总目标。ρ 即擦除与保留的动态配比,w 与 s 则让训练过程能自动把算力分配给当前还没擦干净的概念,避免某个概念主导优化。
扩散模型在视觉生成上性能占优但推理开销巨大。基于缓存的加速是一条有希望的路线,但现有策略都依赖局部相似度启发式——看当前步与上一步的特征差异是否够小来决定是否复用缓存。论文指出这类局部指标与最终生成质量存在显著错位(significantly misaligned with final generation quality),根源在于误差沿去噪轨迹的传播和累积是非均匀的:同样大小的局部误差,发生在不同时间步对最终画面的影响可能差好几倍。
% Comparison between local mismatch and global impact (lower is better). % We conduct a series of independent experiments in which a cached residual is reused at exactly one specific timestep. % For each timestep, the blue marker (Rel-L1) measures the local discrepancy between the ground-truth residual and the cached residual from the preceding step, while the red marker (LPIPS) reflects the resulting impact on final generation quality. % As shown, local mismatch is not a reliable proxy for global impact, as large Rel-L1 values around step~19 and in the final denoising stages correspond to only minor increases in LPIPS, i.e., a relatively small degradation in final output quality. % Note that the prominent Rel-L1 spike around step~19 is a known characteristic of Flux-dev~1.0 when using the Euler ODE solver. % Similar spike behaviors in different diffusion models have also been reported in prior work~DBLP:conf/cvpr25:TeaCache. %
几何条件下的多视角扩散能生成高质量 3D 纹理,但要对每个视角反复评估去噪器,计算开销很大。现有免训练加速器主要利用时间冗余——跨去噪步复用计算。可是在多视角纹理生成里,跳过一步同时也跳过了跨视角交互,而正是这种交互在持续对齐同一表面的不同观测;结果就是一致性与保真度快速退化。换句话说,通用的时间缓存方法在这个任务上水土不服。
The attachment point of , and one cached step. The denoising loop is 14.7% of the paint stage and the only component a neural cache can act on. Views occupy the batch axis, so restricting the forward to the $a$ anchor rows is the natural unit of saving; the anchors' per-step change in $\xz$ is transported through correspondence and added to each other view's own previous $\xz$.
Speed--fidelity ladders on the three backbones leads. Shading marks ${\geq}2\times$; the bottom right is best. Table~tab:main_results carries the headline rows.
评论 (0)