本文贡献:Ring Forcing 的核心是一个环形训练策略:把原视频与反转后的视频首尾相接成环,随机裁掉开头一段,再在中间「挖掉」一部分作为隐藏的真值(Hidden GT),同时施加 Ring Context Drop。这样一来,要预测的目标帧其内容只能从远距离历史里取回,模型没有走捷径的余地,从而在「严格遵循历史」与「保持生成多样性」之间取得平衡。为扩容记忆,作者用压缩加时间步组合的策略,在固定序列长度下把有效历史跨度推到分钟级,并让感受野覆盖整段历史。第三个部件是稀疏 RoPE:让位置编码可以灵活、可扩展地适配记忆条目,同时尽量复用预训练先验,实验显示它比标准按索引排布的 RoPE 收敛更快、重建质量更好。作者还设计了双流结构,在不增加序列长度的前提下让组合策略稳定优于两个单独基线。
Overview of the Ring-Structured Training Strategy. We construct a sequence ring by concatenating a linear raw video with its reversed counterpart. By unrolling the ring into a conditioning sequence, the target frames are naturally embedded into the distant history as the Hidden GT, which essentially forces the autoregressive generator to learn explicit long-range retrieval. To prevent trivial shortcuts, a Random Head Crop is applied to truncate boundary information leakage, while a Random Context Drop is used to balance history adherence and generation diversity.
Qualitative results of 60-second long video generation. The visualization demonstrates the model's capability to maintain overall consistency and visual quality throughout a 1-minute duration in general scenarios.
批判点评:这篇的价值在于把长视频一致性从「算力/上下文长度问题」重新定义成「检索问题」,而环形训练是一个相当巧妙的构造——它用数据组织方式强制模型放弃预测邻近帧的捷径,比单纯加损失项更根本。稀疏 RoPE 也和昨天 JoyAI-Echo-1.5 的「给记忆单独编址」思路撞在一起,可以认为这是当前的一个共识方向。需要保留的疑问有三点:一是自建 A-D-R Benchmark 虽然针对性强,但也意味着「恒存性」这个最关键指标缺少第三方基准交叉验证,作者说的 SOTA 主要在自家尺子上量出;二是环形训练依赖把视频反转,这对有明确因果方向的内容(倒水、物体破碎)是否引入了反物理先验,论文没有讨论;三是分钟级的说法建立在「压缩+时间步组合」之上,压缩必然有损,论文没有量化到底丢了什么、在多长时间尺度上开始失效。
2. EditaLive:两步采样做到14.47FPS直播换装
EditaLive! Unified Character Video Editing for Live Streaming | 澳门大学、vivo BlueImage Lab、大湾区大学 GVC Lab | arXiv:2608.27123
Overview of the three-stage training pipeline of EditaLive. (a) Appearance--Motion Decoupled Editing reformulates character video editing as appearance editing with explicit body and facial motion preservation. (b) Causal Streaming Adaptation converts bidirectional modeling into causal generation using clean historical context and causal attention. (c) Aligned Self-Rollout Distillation compresses the model into a two-step sampler, where Align Forcing and Fixed RoPE align training with streaming inference, and FPSA improves long-term appearance stability.
Tether the Subject, Release the Scene: Query-Aware Memory Routing for Long-Horizon Autoregressive Video Generation | 华中科技大学、小米 MiLM Plus | arXiv:2608.26902
Overview of TetherMem. (a) End-to-end routing preserves identity memory for subject queries and favors recent backgrounds for scene queries. (b) Prior methods select memory frames, whereas TetherMem routes query roles within the retrieved history.
The evaluation landscape exposes two distinct gaps. (a) Identity EP versus scene-progression EP shows that strong identity can coexist with weak progression. (b) nT versus human progression shows that coherent translation is informative but incomplete: Causal Forcing has high nT and low human progression. Values are the method-level aggregates from Table.
What should a long-term video memory store? The same rollout, indexed two ways. Top: the man is recorded at his first appearance ($t_1$); each later frame depicts something new: a passing car ($t_2$), a pedestrian ($t_3$), another passer-by ($t_4$). Bottom Left: frame-indexed memory re-stores old and new content alike at every step, then compresses or evicts it as time passes, so memory pressure grows with video length. Bottom Right: appearance-indexed memory (RECAP-Forcing) stores content when it newly appears, at canonical temporal position $0$, memory scales with the amount of newly appearing content, not with its duration.
实验效果:作为推理期免训练方法,RECAP-Forcing 在多个强基线上一致提升了视觉质量与语义保真度,并优于既有的记忆方法。一分钟长视频的定性对比很直观:基线在「第二个主体入画」的场景里逐渐丢失内容身份并开始漂移退化——尾随的 SUV 化成尘土并出现复制、岩石人裂开并偏离路径;而 RECAP-Forcing 在保住主体身份与视觉一致性的同时没有牺牲运动幅度。记忆可视化图显示,被选中留存的 patch 确实集中在新出现的角色与新露出的区域上,说明新颖度判据按预期工作。
Qualitative results on four one-minute prompts. Across all scenes, the baseline loses subject identity, style, scale, or background consistency. In contrast, RECAP-Forcing preserves the subject(s) and the full scene throughout the minute.
Magpie system overview. Designers define layouts, interactions, rules, and numerical parameters in the game engine. The User Client forwards player actions to the Game Engine, which executes gameplay and emits white-box observations. The independent Render Server converts those observations into generated video and streams the result back to the client. Gameplay state and rule execution remain inside the engine.
Chunk-level interaction and rendering timeline. At 24 FPS, the Game Engine records 20 white-box frames in approximately 830 ms. After transfer, the Render Server generates the corresponding chunk in approximately 620 ms; transfer and buffering account for the short intervals between stages. The first action-aligned visual response therefore appears after roughly 1550 ms, while subsequent output chunks are displayed over approximately 830 ms.
本文贡献:SpatialCrafter 引入一个全局 3D 代理,把生成过程拆成「代理生成」与「外观精修」两阶段。代理阶段用 Point-anchored Sparse Structure(PaSS)Flow 模块,从输入图与点云出发预测空间对齐、几何一致的稀疏体素结构,再经 SLAT 扩散得到粗糙 3D 场景。精修阶段把视频扩散模型重新定位为「生成式延迟渲染器」,只负责在代理给定的几何之上合成高频写实细节。为了让代理与预训练视频模型更好结合,作者提出 Parallel Geometry Injection(把粗糙 RGBD 视频经 VAE 编码后与噪声潜变量做宽度/通道拼接注入)和 Proxy-Aware Corruption 训练策略,提升对代理瑕疵的鲁棒性,同时避免破坏预训练的生成流形;视频 DiT 主体冻结、仅用 LoRA 微调。此外作者构建了 115K 场景的大规模数据集,称是首个面向图生场景任务的混合数据集。
Overview of SpatialCrafter. Given a single image and a camera trajectory, we first construct a global 3D proxy using a native 3D generator with Point-Anchored Sparse Structure (PaSS) Flow Matching. This proxy then serves as a reliable coarse 3D prior for the Generative Deferred Refiner, which transforms it into photorealistic RGB-D video sequences via Parallel Geometry Injection and Proxy-Aware Corruption.
Qualitative comparison with SOTA methods. SpatialCrafter significantly surpasses the baseline methods under extreme challenging camera motions, yielding photorealistic and spatially consistent novel view synthesis. Please refer to the supplementary material for more comparison results.
Overview of LiveVVT. (a) Rolling streaming try-on emits one clean chunk per update from a staggered-noise window with persistent global appearance memory $\mathcal{A}=(\mathcal{A}_g,\mathcal{A}_f)$ and temporal memory $\mathcal{H}_k$. (b) Bidirectional VVT learning, teacher-trajectory regression, and CoMD progressively transfer offline bidirectional priors to few-step streaming generation.
Self-OPD pipeline. Top: The student branches its prediction into $K$ SDE paths (blue), generates images via ODE rollouts (purple), and scores them ($r^+, r^{\mathrm{ref}}, r^-$). Bottom-left: Rewards define a self-referenced advantage $A_k$ that pulls toward high-reward velocities $v_+$ and pushes away from low-reward velocities $v_-$. Bottom-right: A direction-aware coefficient $d_k$ gates the push to prevent it from counteracting the pull, resulting in the multi-branch loss $\mathcal{L}_{\text{Self-OPD}}$.
Inference and training pipeline of RubricRM. (A) RubricRM supports both text-to-image and image-editing preference selection. At inference time, the model first dynamically generates a rubric conditioned on the input, including scoring criteria and corresponding weights, and then scores each image along each dimension based on the rubric, all within a single inference pass. (B) Training is conducted in two stages. The model is first supervised fine-tuned on synthesized rubric data to learn the paradigm of rubric generation followed by scoring, and is then further optimized with GRPO using dimension-level rewards.-0.6cm
We present StreamAV-Bench, a comprehensive benchmark for streaming audio-video generation. (a) The benchmark includes a progressive track to assess instruction adherence and long-horizon stability, and an interactive track to measure interactive response alongside state retention and reuse. (b) The benchmark covers content complexity in both tracks and update complexity in the interactive track, with 320 expert-verified scenarios that span 8 scene domains, 5 audio domains, 5 subject categories, and 4 visual styles. (c) The evaluation is driven by a unified framework of 32 fine-grained dimensions, computed with expert models, MLLMs, and corresponding checklists.
生成模型开始被当作「渲染器」嵌进既有工业管线 — Magpie 的做法很克制也很务实:不让生成模型接管游戏逻辑,而是把玩法执行留在游戏引擎里,引擎只输出白盒画面,再由独立 Render Server 生成最终视觉。这保住了游戏最不能丢的东西——规则可复现、状态可控。SpatialCrafter 也是类似思路,先出一个几何正确的 3D proxy,再让视频扩散模型只负责补高频细节。两者都在退一步:不追求端到端全能,而是把生成模型放在「渲染」这一层,让可控性由传统方法兜底。
评论 (0)