Overview of . Three strategies intervene at different levels of the self-forcing training pipeline: (a) Hybrid Forcing at the input level, (b) dynamics-aware reward regularization at the loss level, and (c) reference perturbation at the conditioning level.
Dynamic collapse in self-forcing distillation. (a)~Text-to-video (Self-Forcing~huang2025self): VBench~huang2023vbench dynamic degree drops from 0.80 to 0.00 while visual quality improves. (b)~Audio-driven avatar (LiveAvatar~huang2025liveavatar): Sync-C and ExpVar degrade while IQA and ASE improve. (c)~Comparing self-forcing with CausVid~yin2025slow (GT-anchored): CausVid's dynamic degree fluctuates around 0.43 without collapsing, confirming that unanchored self-conditioning, not DMD distillation itself, drives the collapse.
Overview of the VA-Judger training pipeline. Stage 1 cold-starts the model on easy preference pairs. Stage 2 aligns the model on harder pairs through human rejection sampling. Stage 3 performs GRPO using answer-level and dimension-level rewards grounded in human reasons.
Human preference rates over 200 three-way comparisons. For each prompt, participants select the best video-audio output among LTX-2, LTX-2 post-trained with OmniNFT, and LTX-2 post-trained with VA-Judger.
Partition the Support, Reconstruct the Residual: Training-Free Sparse Attention for Video Generation and World Models | Texas A&M University | arXiv:2608.18484
Overview of SparsePR. Response-Coupled Partitioning builds executable K/V and query groups, and Probe-Fitted Residual Reconstruction uses exact probe rows to correct the sparse output.
Cross-model density sensitivity. Mean (top) and p99 (bottom) normalized attention-output error versus total executed-pair density for semantic and response-coupled partitioning under the same residual-repair configuration. The dashed line marks the (22%) reference density.
An overview of FireRedTTS3, including (a) the RedAE Tokenizer with semantic supervision, (b) FireRedTTS3-Base for multilingual and multi-dialect voice cloning, and (c) FireRedTTS3-Instruct for voice cloning, instruction-controlled voice design, and speech editing.
批判点评:这篇的取向很对我的胃口——不加模块、不加阶段,只在表征约束上动手,用一个现成的冻结理解模型当语义锚。这个思路的可迁移性其实超出 TTS:任何连续自回归生成任务,只要该模态存在成熟的理解侧编码器,都可以照这个套路给特征空间加语义正则。小红书把 Base 和 Instruct 两个变体分开发布也很实用,前者专注克隆质量,后者覆盖指令与编辑。但要注意的地方不少。第一,「简单」是相对的——冻结音频编码器本身就是一个在多任务上训过的大模型,把它算进系统开销后,与「额外挂语义模块」的复杂度差距没有摘要说得那么大,只是训练流程更短。第二,效果论证全部依赖自动指标(可懂度、说话人相似度),而语音生成里最关键的自然度、韵律合理性与情感恰当性都必须靠 MOS 类主观评测,摘要层面完全没提。第三,语音编辑的评测只在 Ming-Freeform-Audio-Edit 上做,这个基准的覆盖广度和难度分布尚未被社区充分检验,单一基准上的领先说服力有限。第四,论文只给出一张架构图,缺少可视化的实验对照,无法判断误差累积到底被压住了多少——「更稳定」这个说法需要长句合成的失败率曲线来支撑。
6. MegaParts:300部件3D生成靠离散token压出来
MegaParts: Scaling Part-Aware 3D Object Generation to 300 Parts via Token-Efficient Autoregressive Modeling | 香港大学、上海人工智能实验室、复旦大学、同济大学、浙江大学 | arXiv:2608.14783
Overview of MegaParts. Our framework consists of two stages. Top: we learn a token-efficient vector-quantized representation for part geometry. Bottom: we train a part-aware autoregressive generator that represents a 3D object as a structured sequence conditioned on a text prompt and optional object- and part-level bounding boxes. If bounding boxes are absent, the model first predicts the object box and part boxes, followed by autoregressive generation of local shape tokens for each part. The decoded part geometries are then assembled into the final part-aware 3D asset.
. It consists of 1) MASD modulates base queries $Q_{base}$ with modality embeddings $E_m$ to distill visual tokens $T_{vis}$; 2) DCI applies a semantic shift to calibrate instruction extraction, yielding dual-context tokens. These unified embeddings are injected into a Video DiT to steer the generation process. All examples presented in the figure are collected from the open-source OpenVE~he2025openve.
实验效果:在 VicEditBench 上,VicEdit 在基础指令编辑和视觉上下文编辑两类任务上都取得最优。定性对比很直观:三个案例分别是局部风格(把花瓶涂成深番茄红)、全局风格(转成雪天)、局部移除(去掉时钟及其阴影)。基线们的失败模式各异——Lucy-Edit 和 NovaEdit 大量出现 No Change(干脆没动),VideoCoF 出现 Flat Coloring(上色平板无质感),Kiwi-Edit 则是 Severe Artifacts 和 Inpainting Failure;VicEdit 在三个案例上分别达到自然编辑、和谐风格化与完全移除。
Comparison with baselines on visual in-context editing tasks. All examples presented in the figure are collected from the open-source OpenVE~he2025openve and Senorita~zisenorita.
Overview of our method. Top: Given an input multi-shot video and corresponding editing masks, our method first trains a Supervisory Adapter to inject multi-shot information into the diffusion backbone. We also introduce two key strategies: Cross-Shot Packing and Sparse Cross-Attention to keep the consistency during training. Bottom: During inference, editing is performed on the first frame and produces coherent and identity-preserving multi-shot video output.
Qualitative comparison results with the state-of-the-art methods. We compare our multi-shot approach against leading single-shot video editing baselines on a four-shot editing task. The results show that our method successfully executes the complex subject replacement, demonstrating superior cross-shot temporal consistency.
批判点评:「拿多视角视频数据集当多镜头编辑的监督来源」是这篇最聪明的一步——它把一个数据不存在的问题,转化成一个已有数据的重新解释问题,成本几乎为零而监督信号的语义正好对上。跨镜头打包按语义相关性而非时间邻近去聚合上下文,也比简单扩窗口更贴合多镜头的结构特点。ECCV 2026 接收和港科大+百度的组合也说明这套方案经过了同行审视。但几点需要留意。第一,多视角与多镜头并不完全等价:多视角是同时刻不同机位,主体姿态与光照基本一致;真实多镜头往往跨时间、跨场景、主体状态本身在变化。用前者监督后者,在「状态确实应该变化」的镜头切换上可能反而过度约束。第二,评测基准由作者自行整理,规模与构成细节不明,「显著优于」缺少可核对的数值。第三,推理链路依赖首帧编辑质量,首帧一旦有瑕疵会被跨镜头传播机制放大到全序列,这个失效模式论文未讨论。第四,定性对比里 TokenFlow 出现「基本没改动」的结果,与 VicEdit 那篇里基线的 No Change 现象类似,都提示当前多镜头编辑的基线调用方式尚未标准化,横向对比结果需要谨慎解读。
9. AnyTalk:零动画数据给任意角色配口型
AnyTalk: Speech Animation for Arbitrary Characters Leveraging a Video Generation Model | KAIST(TVCG 接收) | arXiv:2608.16143
关键词:3D语音动画,视频扩散先验,零动画数据,blendshape优化,蒸馏实时化,KAIST
⚠️ 前序问题:给 3D 角色做音频驱动的语音动画,现有方法要么依赖该角色专属的动画训练数据,要么需要费力的绑定与重拓扑。这两条路都把成本卡在了「每来一个新角色就得重做一遍」上,风格化角色尤其吃亏——它们的面部拓扑和 blendshape 配置千差万别,几乎不可能有现成动画数据。作者想问的是:既然视频扩散模型已经在海量真人说话视频上学到了强大的运动先验,能不能把这份先验直接借给任意 3D 角色,从而完全绕开动画数据这一环。
Finetuning process of $D_{CsF}$. The rendered images of the target 3D character and the zero condition $C_{zero}$ are used to form a pair for fine-tuning the model. To preserve the motion prior, only the spatial residual network of denoising UNet is trained (flame icon), while the other layers remain frozen.
Comparison with baselines on Morphy. Each phoneme and its corresponding frame are presented for comparison. Ours best follows the given audio with wider mouth opening and precise lip closure, while ScanTalk, DiffSpeaker + NFR, and CodeTalker + NFR exhibit only subtle movement.
Overview of PersonaShot evaluation framework. Given generated multi-shot videos and prompts, PersonaShot extends conventional quality assessment with three specialist dimensions for narrative-level evaluation: causal physical continuity, affective dynamics, and cinematic grammar.
实验效果:评测揭示不同 SOTA 模型呈现出各异的能力剖面,并暴露出感知质量与跨镜头叙事连续性之间存在明显落差——视觉上很有说服力的视频,仍频繁出现物理状态重置、情绪突变和电影关系断裂。雷达图对比 Seedance、HoloCine、EchoShot 三个模型:Seedance 在物体状态、几何尺度、表情自然度、180 度规则上明显领先,HoloCine 在物理连续性维度接近但在情绪弧和视线匹配上落后,EchoShot 整体最弱且在电影语法一侧塌陷严重。失败案例图给出的典型问题包括空间布局漂移(人物与电话亭的相对位置在三个镜头间跳变)。人类研究表明作者的评估器与专家判断高度一致。
Overview of PersonaShot's person-centric evaluation. Top: Qualitative examples illustrating our three core dimensions: physical continuity, affective dynamics, and cinematic grammar. Bottom: Quantitative comparison across fine-grained metrics for specific state-of-the-arts, revealing their distinct capability profiles.
评论 (0)