Overview of EchoWM. Media context, structured text, and user controls share a unified camera-intent interface. Discrete controls or metric poses are converted into globally calibrated relative 6-DoF trajectories and serialized as an event stream. A relative UCPE branch injects this trajectory into video self-attention before audio--visual cross-attention, while the audio stream receives no direct trajectory condition. Four progressive training stages yield synchronized audio--visual generation and support multi-turn inference through synchronized tail-window conditioning.
Qualitative scale-control comparison. LingBot-World-v2 and EchoWM receive matched reference observations and navigation conditions. When the camera approaches a foreground character or scene structure, LingBot-World-v2 exhibits larger changes in apparent subject scale and less predictable displacement response, whereas EchoWM preserves a more gradual scale change across the trajectory. This qualitative comparison complements the metric-scale and speed-response ablations: a fixed global translation scale keeps requested displacement amplitudes comparable across clips and makes approach behavior more controllable.
Overview of VMM-forcing unrolling. During training, we fully unroll streaming inference: the causal student generates each temporal block from noise with the target few-step schedule, commits the denoised block to the KV cache as history, and proceeds to the next block. At a sampled local transition for each selected block/rollout, we query the bidirectional teacher and auxiliary model at the student-induced states, and supervise the student by matching the teacher--auxiliary residual velocity moments under the same autoregressive context. The student-generated first-block condition is also fed into the teacher and auxiliary model.
Qualitative Text-to-Video (T2V) results on frequent user prompts. We compare the Self-forcing conversion and the proposed VMM-forcing pipeline. VMM-forcing produces richer motion dynamics and scene evolution while maintaining texture realism and temporal coherence, whereas Self-forcing often yields over-sharpened, plastic-like appearances and reduced temporal diversity.
实验效果:跨四种不同的 3D 表示,FixAnything 都用轻量微调稳定提升了渲染质量,证明单个通用视频先验可以替代多条专用精修管线。最直观的是点云与稀疏输入的修复结果:游乐场攀爬架序列中,输入的中间三帧几乎只剩黑底上的散点(点云稀疏到看不出结构),输出恢复出完整的金属管架、地面绿漆与远景树木,与 GT 高度接近;亭子序列里输入帧被大面积黑洞和撕裂纹理占据,输出补出了完整的木构屋檐、栏杆、石阶与背景树林,逆光帧的光晕也被合理重建。mask 感知条件的消融显示,不加 mask 时干净的训练视角也会被模型重画而丢失细节。框架的简洁性还带来一个附加好处:将来出现更强的视频模型可以直接换上,不需要重新设计架构。
FixAnything generalizes across input representations. The same model cleans up sparse COLMAP point clouds (top) and meshes (bottom), producing photorealistic output despite minimal visual input except for the clean first and last views.
Loopy: Seamless Video Loop Generation via Anchored Looping Shift of Positional Embedding | 天津大学、香港理工大学、腾讯 | arXiv:2608.23090
关键词:循环视频生成,RoPE位置编码,分层时序控制,锚定层,免训练,RGBA视频,ACM TOG
前序问题:循环视频在网页动效、游戏素材和社交媒体里需求很实在,但现有做法质量普遍不行,根因在于大家都绕过了一个基础问题:视频生成模型究竟如何感知时间顺序,这种感知又如何决定首尾能否接得上。常见的两条路各有硬伤——直接在提示词里写「scene looping seamlessly」,模型基本不理解,生成出来的首尾差异巨大(鲨鱼从侧身游到正面怼镜头,根本接不上);用首尾帧到视频(FLF2V)把首帧复制给尾帧作为参考,虽然强制首尾一致了,但模型为了同时满足两端约束会把中间的运动压得几乎静止,得到的是「能循环但不动」的伪循环。两条路都没有触及 DiT 内部时序表示这一层。
本文贡献:论文首次揭示:DiT 内不同注意力层的位置编码对时序顺序的控制强度差异很大,其中最强的那一层可以视为「锚点」。他们把这个锚定层设为循环视频的参考点,为其余层提供强上下文先验。基于此提出锚定位置编码偏移策略——按每层实际的时序控制效应给出层特定的偏移长度,把 DiT 对时间的感知从「一条直线」变成「一个圆」。实现上极简:在每个 DiT block 的自注意力里,Q/K 除了原始 RoPE R_{t,h,w},再并行算一份偏移后的 RoPE,偏移量为 (t − δ_l) mod T',其中 δ_l 是第 l 层的时间偏移,两者按层组合后再进 attention 与交叉注意力+FFN。整个过程不改权重、不需重训,是纯推理期干预,因此可以直接挂在 Wan2.1/2.2、HunyuanVideo 1.5 等现成基座上,并同时支持 RGB 与 RGBA 视频,还能叠加身份控制与风格迁移这类 AIGC 能力。
Comparison of Wan2.2 14B and HunyuanVideo 1.5 outputs when shifting RoPE at the layers with the first- and second-largest $\alpha_l$, relative to the original generated videos.
Comparison with Wan2.2 in text-to-video (T2V) and first-last-frame-to-video (FLF2V) modes. FLF2V and our method use the same prompt. Although the prompt explicitly requests a looping scene, T2V fails to produce a seamless loop, while FLF2V degenerates to a nearly static video. Neither mode generates a looping video with vivid motion.
Overview of Object-Uni. Our model unifies object-centric spatial understanding and orientation-controllable generation. The dotted modules denote object-centric novel view synthesis.
Evaluation of pose-controllable generation in single- and multi-object scenarios. The reference image in the first row serves as the source for the input text and 9D pose annotations. Object-Uni shows better text alignment and pose consistency under diverse pose conditions.
本文贡献:作者不走扩散,而是把下一尺度自回归(next-scale autoregressive)范式适配到以人为中心的视角合成上,使其支持更高分辨率、多视角输出与更强的跨视角一致性,且全部在单次前向里完成。管线是:N 张 512×512 输入图经多尺度 VQVAE 编码器得到逐尺度 token(i_1 为 N×1×32、i_2 为 N×4×32、i_3 为 N×9×32,依次加密),与 CLIP 特征、目标相机外参 (R_1,T_1)…(R_M,T_M) 组成的全局条件一同经 word embedding 进入 Causal Next-Scale Transformer,自回归产出 M 组目标视角的逐尺度 token r_1/r_2/r_3,再由多尺度 VQVAE 解码器一次性解出 M 张 512×512 视角;训练时目标视角走 teacher forcing。这个范式的关键红利是训练数据经济性:不需要 2D 预训练,且低分辨率、通用的预训练可以直接复用,只在最后几个训练阶段才用到全尺寸的任务专用图像,因此用一个更小但更真实的数据集就能收敛。他们在覆盖多样身份与服饰的合成人脸数据集上训练,并把输出接到已有的 GS-LRM 上做像素对齐的 3D 高斯提升,最终得到人脸的 3D 模型。
Our architecture: input images are encoded into multi-scale features by the VQ-VAE encoder and used as inputs for the transformer. They are also embedded into a global conditioning along with the camera-pose RT matrix, used both in the AdaLN layers and the SOS tokens. The transformer outputs one scale at a time for every output view simultaneously, up to the final scale. The output features are decoded by the VQ-VAE decoder to produce the final RGB images, which can be lifted to 3D gaussians by an off-the-shelf GS-LRM model. During training, previous scales are teacher-forced, while during inference the model is auto-regressive. The architecture is the same for all our models, with the only difference being the number of input and output views.
批判点评:最大的问题是训练数据全部为合成人脸,论文自己也强调「a smaller but more realistic training dataset」,但合成到真实的域差距未见量化评估,模型在真实拍摄人脸上的表现是空白,而这才是实际用途所在。量化结果同样只以「we observe gains」表述,没有 PSNR/LPIPS/身份相似度的具体数字,也没和 FaceLift、Splatter-Image 这些明确出现在其定性对比目录里的基线给出表格化比较。分辨率停在 512×512,对影视级人脸资产而言偏低,而下一尺度自回归的 token 数随尺度平方增长,往 1024 走的代价论文没有交代。第一作者的贡献声明为「在 Meta Reality Labs 实习期间完成」,且 LaTeX 源里仍留有多处 TODO FINAL: Replace with your institution list 未清理的模板注释,说明这是一版尚未定稿的投稿,结论宜谨慎引用。此外整套流程依赖外部 GS-LRM 做 3D 提升,端到端的几何误差如何在两个模型之间分摊也没有分析。
7. Lever-Edit:不用编辑奖励也能做编辑在线RL
Can We Perform Online RL for Image Editing without Editing Rewards? | 北京大学 | arXiv:2608.22780
Overall two-stage T2I reward transfer pipeline in Lever-Edit.
实验效果:实验显示 Lever-Edit 在编辑对齐度与源内容保持上都能与基于编辑奖励微调的方法竞争,同时明显优于几种直观的迁移基线。定性对比按提示遵循、美学、一致性、自然度四组维度组织,八个编辑任务里可以看到:不做对齐(w/o Alignment)时「把狗变成大理石雕像」只是把狗漂白、「换成极简白背景」直接把人像退化成线稿;Lever Pure Edit 与 Lever Pure VLM 各有偏差(前者美学不足,后者出现主体改变,如冲浪板换成桨板时人物姿态被重画);Lever Offline SOTA 在自然度上仍有塑料感;完整的 Lever-Edit 在四组维度上综合最好——大理石雕像有正确的石材质感与基座、红色跑车放到建筑前方且透视正确、热气球加入天空且尺度合理、白色背景替换保留了人物细节。
Qualitative analysis of different methods. Lever-Edit successfully edits the images with strong prompt following and high aesthetics, while maintaining both semantic and source consistency with naturalness. Zoom in for better visualizations.
批判点评:「competitive against editing-reward-based fine-tuning」这个表述值得留意——它说明本方法并未超越用真正编辑奖励训练的上界,价值在于省掉三元组标注成本,而不是效果更好;如果编辑奖励模型未来变得容易获得,这条路径的必要性会下降。摘要没有给出任何数字(编辑对齐、源保持的具体指标与基线差距均缺),也没说明用了哪几个 T2I 奖励模型及其聚合权重,而这直接决定结果可复现性。方法引入的 captioner 是新的失效点:反事实目标描述如果写错(把「移除」理解成「替换」),策略就会被系统性地引向错误方向,论文未讨论 caption 错误率及其对训练的影响。「参考一致性靠粗粒度语义转换」是全流程最薄弱的一环——T2I 奖励天生不看源图,无法真正惩罚身份改变,从定性图里也能看到 Lever Pure VLM 出现主体被重画的情况;这类结构性缺陷不是靠 captioner 能补齐的。另外两阶段交替冻结的训练配方对超参可能较敏感,稳定性未见报告。
8. SACHA:2.23MB存下动态高斯头像
SACHA: Semantic-Aware Compression for 3D Gaussian Head Avatars | 香港城市大学、阿里巴巴达摩院 | arXiv:2608.23133
关键词:3D高斯头像,语义感知密度控制,外观-运动解耦压缩,率失真优化,头部先验参数,阿里达摩院
前序问题:可驱动的 3D 高斯头像渲染质量高、控制灵活,但高斯基元数量庞大,存储和传输成本都很重,这是它进入视频会议、VR 社交、移动端这些实际场景的主要障碍。现有的高斯头像方法有两处共同盲区。第一,忽视了头部不同语义区域的视觉显著性差异——眼睛、嘴唇这些高频且被注意力集中的区域和后脑、脖颈这类低频区域,如果按同样的密度分配高斯基元,就是把预算浪费在看不出差别的地方。第二,几乎没人做训练完成后的头像序列压缩:一段动态头像在时间维度上高度冗余(外观基本不变、只有表情和姿态在动),但主流做法仍然逐帧存储完整的高斯参数。
本文贡献:DefaultShift 是一套成对审计流程。上半部分 Default Shift-Audit 分三段:成对采样阶段,对同一提示(如「a photo of a helicopter」)分别用参考模型 T(多步扩散)和加速替换模型 S 各生成一批图;语义默认阶段,用 VLM 评估器按封闭语义词表(颜色、背景、光照、视角等,每类若干离散取值)给每张图打标,得到两个直方图;漂移审计阶段,用 TV 距离等指标算概率质量的移动,并做教师自举校正(TV_adj = TV(T,S) − η_T)以剔除参考模型自身的采样噪声,再经 q=32 筛查、q=64 确认两级预算与 cell-level bootstrap,输出按配方排序的结果和方向性签名(偏暖还是偏灰)。下半部分 Default Shift-Select 是离线校准:给定参考直方图模板与 CLIP 质量下限,从 M=128 的候选池里做分布匹配(min TV),挑出 N=32 张既满足质量门槛、又让语义分布贴近参考模型的图。方法把「解释性排序」与「确证性的 cross-fit 推断」分开,避免用同一批数据既选又证。
Overview of DefaultShift. The audit generates matched samples from a declared reference and accelerated replacement, maps each image to a closed semantic vocabulary, and compares their conditional distributions. Cell-level discrepancies are aggregated for recipe ranking and directional analysis. DefaultShift-Select uses reference statistics and a quality-filtered replacement pool to construct a distribution-matched offline subset.
Acceleration recipes differ in magnitude, attribute, and direction. a The unified $q=64$ audit covers all 14 replacement pairs. Lines are 95% object-cluster intervals. b Color dominates the four-attribute profile, while viewpoint shifts are small. c Replacement-minus-reference changes in gray and warm mass reveal opposing recipe fingerprints.
批判点评:整套审计建立在 VLM 评估器加封闭语义词表之上,两处都会引入偏差:VLM 自身对颜色、光照的判断有系统性倾向,论文虽给出了人类与 VLM 的混淆矩阵,但那只覆盖「clear case」,模糊样本上的一致性未知;封闭词表则把连续属性强行离散化,词表粒度直接决定测出多大的 TV 距离,换一套词表数值不可比。「颜色差异 0.054 到 0.303」是 TV 距离而非人眼可感差异,缺少感知尺度上的锚定,读者很难判断 0.2 究竟意味着轻微偏暖还是明显变色。DefaultShift-Select 只是事后筛选,本质是丢弃不合分布的样本来贴近参考分布,代价是有效产出下降(128 选 32,即弃掉 75%),这个吞吐损失论文没有和「干脆用慢速模型」做成本对比。审计限定在提示未指定的属性上,但「未指定」的边界本身依赖提示写法,作者用的是「a photo of a {object}」这类极简模板,真实业务里的长提示会指定更多属性,结论的外推性有限。最后,下游 4.3/7.5 点的增益只在其自建的均衡评测设定下成立,任务范围较窄。
10. FIRM-Video:先核查再打分的8B视频奖励模型
FIRM-Video: Check Before You Score for Reliable Text-to-Video Reward Modeling | 上海交通大学、腾讯优图实验室、同济大学 | arXiv:2608.21839
Overview of the FIRM-Video data construction framework and reward model training. Given a prompt and its generated video, FIRM-Video applies a unified check-before-score pipeline to construct fine-grained supervision for instruction following, world coherence, and perceptual quality. The resulting dimension-specific analyses and scores form FIRM-Video-90K, which is then used to train the FIRM-Video-8B reward model.
评论 (0)