Core intuition of ThermoDPO. Flow Matching (gray) transports noise to the pretrained data manifold. FlowDPO (green) may reach preferred regions through an off-manifold displacement, whereas ThermoDPO (red) adds a winner-side anchor intended to keep the redirected mass closer to the pretrained manifold. The formal statements and their assumptions are summarized in tab:notation_theory_map; the behavior is evaluated in the toy study (Fig.) and real-image study (Fig.).
Pairwise human evaluation of ThermoDPO-weighted against different baselines on text accuracy and visual quality over 30 prompts. Each stacked bar reports the percentage of prompts for which ThermoDPO-weighted is preferred, tied, or dispreferred relative to the corresponding baseline. ThermoDPO-weighted shows consistently stronger performance on visual quality while remaining competitive on text accuracy.
Overview of AViTS and spatiotemporal importance estimation. (a) Three-stage dynamic-resolution sampling: Stage~1 low-res denoising with signal collection over the last $N_T$ steps, Stage~2 selective upsampling of top-$K$ tokens, and Stage~3 full-res refinement. (b) Temporal importance from token-wise step variance across the collected snapshots. (c) Spatial importance from aggregated latent--text cross-attention (averaged over heads/blocks and collection steps); fused scores guide prioritization.
Qualitative comparison on GEdit-Bench with Qwen-Image-Edit. AViTS achieves superior semantic preservation and visual quality at a higher acceleration ratio (5.24$\times$) compared to state-of-the-art baselines.
Training architecture of SingDance. The first video frame is separately encoded as a clean reference latent, while subsequent frames are temporally compressed and noised in latent space. Wav2Vec~2.0 and MuQ features, together with the vocal-role condition, are selected by hard-compact routing and supplied to frame-wise joint audio injection. The repeated block is schematic: joint audio injection is applied immediately after each of 10 selected blocks in the 30-block DiT backbone. The denoised latents are decoded into output frames. Snowflakes and flames denote frozen and trainable modules, respectively.
Qualitative paired vocal-role switching on two SingDance examples. Each Lip-on/Lip-off pair shares all role-independent inputs and inference settings; only the vocal role changes.
AvatarDynamizer: From Static to Dynamic Human Avatars via Generative Dynamic Textures | 马普信息学研究所、VIA Research Center、EPFL | arXiv:2608.19900
关键词:4D数字人,动态纹理,3D高斯,视频扩散先验,马普所
前序问题:全身数字人要跨过恐怖谷,表面动态是关键——衣服的褶皱、布料的甩动,这些细节缺失时人眼立刻觉得假。但现有两条路各有硬伤:通用方法能从单张图、单目视频或文本提示重建静态 3D 数字人,可它们的动画是骨骼驱动的,衣服跟着骨头刚性变形,没有真实的表面动态;专用方法(为单个人训一套)渲染质量和动态都好,代价是每个人都要一套昂贵的多视角采集。近期的通用动态方法则难以把表面动态嵌进表示里,结果要么多视角不一致,要么动态表现力不足。
本文贡献:核心想法是把「数字人的表面动态」转成「纹理生成问题」,从而可以直接借用预训练视频扩散模型的先验。具体做法:设计一个纹理空间的表面动态嵌入,用编码器-解码器把姿态相关的动态编码进动态纹理图(dynamic texture map)——纹理图是二维的,天然与视频扩散模型的输入格式兼容;解码时再把纹理解成 3D 高斯,由高斯渲染器出图,多视角一致性因此得到保证。整条流水线的输入端很宽:多视角图、单目图、单目视频、文本提示、随机采样、3D 扫描都能作为静态数字人来源,经身份拟合得到身份纹理,再与用户指定动作产生的运动纹理一起送进动态纹理生成器,最后由通用高斯解码器渲染出自由视角结果。由于现有数据集在规模、序列长度或动作多样性上都不够,作者还自采了一个大规模多视角长序列数据集。
Overview. Given a static human avatar reconstructed from multimodal inputs, e.g. multi-view images, we first perform identity fitting to obtain its coarse shape and identity texture. Our Dynamic Texture Generator predicts dynamic texture maps conditioned on identity and novel pose. Then, our Generalized Gaussian Decoder recovers Gaussian splats for faithful and view-consistent renderings.
Qualitative Comparison on DynaHuman. Given a static avatar reconstructed from a monocular image using LHM or from multi-view images (MV), we compare our method with state-of-the-art generalizable surface dynamics methods. Our AvatarDynamizer produces faithful wrinkles while preserving both 3D geometry and identity consistency.
Overview of GenRec. Given posed input views and target poses, monocular depth and forward warping produce per-target observation masks and warped renders. These condition a (I) multi-view flow matching backbone, which jointly denoises RGB and scene-coordinate latents through cross-view and cross-modal attention. A (II) reconstruction branch then refines only the observed pixels via sparse 3D attention guided by the predicted scene coordinates, while the observation mask gates gradients so the generative prior is preserved in unobserved regions.
Overview of KeyID. (a) Reference-Aware Video Generation. A global prompt, optional temporal prompts, and a face reference are fused by GPT-5 into an enhanced prompt, from which a first frame is generated and optionally edited with non-human objects. The edited frame and the global prompt, along with optional temporal prompts, drive I2V to produce a video draft. (b) Identity Preserved Keyframe Editing. Keyframe windows are uniformly sampled in groups along the video draft, optimal keyframes are selected in keyframe identity refinement, and motion interpolation produces the final temporally coherent video with preserved identity and prompt adherence.
Qualitative comparison. Four evenly sampled frames are shown for each video. KeyID better preserves identity consistency and prompt adherence across complex actions.
Mise-en-Sc`ene. The method runs in two stages. In Stage~1 (simplified layout generation), a text brief, a background canvas, and the visual elements are mapped by frozen encoders into a joint token sequence and processed by a flow-matching Diffusion Transformer adapted only through a knockout-selected LoRA, which produces a layout-aware draft design. Because the draft comes from stochastic sampling, it still contains rendering artifacts such as distorted text and warped logos. In Stage~2 (deterministic match-and-place), a VLM grounder locates each original element in the draft independently, and the original high-resolution layers are alpha-composited at the grounded boxes to yield the final design, which keeps the emergent layout while restoring exact, pixel-faithful, editable assets.
Qualitative comparison on the PrismLayersPlus test set. Each row shows, from left to right, the input design elements provided without spatial context, followed by the composites from FlexDM, LaDeCo, our Mise-en-Sc`ene, and the ground truth. Every method is rendered from its own predicted layout; our column is the match-and-place output. Our designs follow the ground-truth arrangement most closely; see text for a per-row discussion.
批判点评:第二阶段能保证素材保真,但它同时暴露了第一阶段的局限:草稿里文字和 logo 是失真的(框架图里明确标注了 text distortion、brand/logo distortion、detail inconsistency 三类问题),也就是说 DiT 学到的其实是「大致该放哪」而非真正理解了排版规则,match-and-place 是必要的补丁而非可选增强——一旦草稿里某个元素的位置判断错了,第二阶段只会把原始素材精确地搬到一个错误的位置上。评测的主指标是与真值的感知接近度,这个口径隐含假设了 PrismLayersPlus 的真值就是好设计,但平面设计本身允许多解,「接近真值」和「设计得好」并不等价,而论文没有人工设计师的偏好评测。另外基准单一,元素数量、语种、画布比例的覆盖范围未交代。
10. BeyondMasks:删掉物体也要删掉它的影子
BeyondMasks: Evaluating Causal and Physical Consistency in Video Object Removal | Bilkent 大学、Koç 大学、Hacettepe 大学(ECCV 2026) | arXiv:2608.20107
Overview of the BeyondMasks construction pipeline. Top: Paired intervention-based generation, where an object-present video is derived from a clean background video, ensuring temporal alignment. Bottom: Representative categories of object-induced causal after-effects, including illumination, reflection, volumetric, and dynamic interaction effects.
Pixel metrics fail to capture causal correctness. Although LPIPS scores are similar across methods, their ability to remove object-induced effects (e.g. shadows or reflections) differs substantially. CORE separates these failure modes by evaluating object removal (CORE-OS) and elimination of causal after-effects (CORE-AES). $\downarrow$ better for LPIPS; $\uparrow$ better for CORE scores.
评论 (0)