Overview of ours. Given a reference image and streaming pose signals, our system generates video blocks autoregressively. Each block undergoes 3-step denoising followed by a clean KV update. PR-Sink augments a three-block rolling window with the first generated block and a pose-matched historical block selected from a compact memory bank. Ulysses sequence parallelism distributes attention computation across GPUs.
Quality and identity over a three-minute rollout. Metrics are computed on a 0--10,s prefix, the 0--30,s initial window, and three subsequent temporal segments. The full model is highlighted in red. Flat trajectories indicate resistance to accumulated degradation. Higher is better for ASE, IQA, and DINO-S; lower is better for FID and V-MAE.
Overview of Avatar-Forever. Top: Existing streaming avatar systems employ sequential distillation pipelines, leading to stage-wise dependence and objective interference. Bottom: Avatar-Forever learns few-step efficiency and long-horizon robustness in parallel. A lightweight adapter is trained under accumulated rollout errors, and it is merged with the distilled generator in inference. In addition, ForeverCache is proposed to reuse historical features for efficient streaming inference.
Visual results of extended-duration generation. We uniformly sample frames from a continuous video generated for more than 11 minutes. Avatar-Forever maintains stable identity, facial structure, and scene content while producing natural expressions, coherent head and body motion, and accurate audio--visual synchronization throughout the extended autoregressive rollout. The bottom row shows the corresponding driving-audio waveform.
Training loss and STFT L2 distance curves for different DiT width/depth configurations. Topline denotes DashengTokenizer encode-to-decode reconstruction.
Overview of . Left: video diffusion attention contains a low-mass tail, allowing approximately $40\%$ of the block area to be removed while retaining approximately $99\%$ of the attention mass. Right: after a dense warm-up, measures exact block masses at $t_0$ and constructs the mask $\Omega$ by retaining, for each layer, head, and query block, the smallest prefix whose cumulative mass reaches $\theta$. The same frozen mask is reused at all remaining steps, while Q/K/V and attention weights are recomputed from the current hidden states. Bottom: with optional feature caching, the cache selects which steps are computed, while accelerates self-attention within every computed step.
Speed--quality trade-off on Wan2.1-1.3B. Each point reports VBench Overall versus end-to-end speedup. and its cached variants form the entire Pareto frontier: every baseline configuration is matched or dominated by a configuration, showing that near-lossless sparse attention preserves quality better than aggressive sparsity or cache-only acceleration.
Overview of UniSwap. The three-stage pipeline comprises (a) In-Context Pretraining for joint audio-video replacement, (b) Conditional Streaming Adaptation with block-causal masking and KV-cached inference, and (c) Efficient Self-Forcing DMD, which reduces denoising to 3 steps per block. Feature-RoPE Decomposition bounds cached positions while preserving cross-modal physical-time alignment.
Qualitative comparison on the long-video benchmark. Frames are sampled every 10 seconds from 1-minute generations. The baselines exhibit identity drift and visual artifacts as generation progresses (highlighted in red), while UniSwap preserves the reference identity throughout the full duration. Zoom in for details.
Schematic of the proposed GeoFlow framework. % Instead of initializing from an uninformative noise, % Our method leverages multi-view geometry to bridge the gap between the source and target manifolds. % The pipeline projects visual features into a latent point cloud, which is rendered to the target view to form a geometric prior. To bridge the gap between the source and target manifolds, we leverage multi-view geometry to construct a Geometry-Aligned Prior Distribution, where an Spatially-Adaptive Noise Injection strategy is introduced to reduce error accumulation due to depth and rendering artifacts. Bottom Panel: As visualized in the sampling trajectory comparison, our initialization significantly shortens and straighten the flow, % yielding a nearly straight generation path (green) compared to the highly curved trajectory of the standard Gaussian baseline (red), thereby enabling high-fidelity synthesis within several inference steps.
本文贡献:提出 StateFlow,以状态为中心的生成式预演框架。不做一次性视频生成,而是用一个可编辑的 3D 世界来组织场景结构、演化与相机,需要更高保真时再让现成视频模型来增强画质。这个世界被维护为场景元素与相机配置的持久结构化 3D 状态,作为预演的核心工作表示。三个阶段:State construction 通过先验引导、冲突感知的双视图初始化,把生成的 2D 内容抬升为连贯 3D 世界;State evolution 把用户意图翻译成结构化的状态转移并保留世界记忆,避免每次编辑都全场景重生成;State access 用渲染反馈反思机制把相机方案细化为视觉上可行的轨迹,而不是只依赖 VLM 的语义判断。
Qualitative comparison on Scene Generation
实验效果:实验显示 StateFlow 能为视频创作与类游戏原型制作产出高质量 3D 世界。
Qualitative comparison on previsualization applications. Case~1 and Case~3 showcase video creation, while Case~2 shows 3D game prototyping. Video-generation baselines consistently suffer from spatial--temporal inconsistency, identity drifting, and restricted camera motion, whereas StateFlow operates on a persistent and interactive 3D world, yielding precise, highly controllable, and geometry--appearance-consistent results.
批判点评:这篇提的问题比它给的方案更重要。当前视频生成的交互范式基本是「抽奖」——改一个词,整段重出,上一版的所有满意之处一并丢失,这在需要迭代收敛的专业创作流程里是致命的。StateFlow 把矛头指向缺少持久状态,这个诊断很准:把「视频」当作输出、把「3D 世界状态」当作可编辑的真源,编辑就从重新生成变成状态转移,世界记忆得以保留,这才是专业工具该有的形态。三阶段拆分也各有对症之处,尤其 State access 用渲染反馈来验证相机轨迹的物理可行性、不轻信 VLM 的语义判断,是个务实的清醒设计——VLM 说「镜头从这里推过去很好」和这条轨迹真的不穿墙是两件事。但这篇的实验部分是明显短板:摘要只给了「能产出高质量 3D 世界」这种定性结论,没有任何量化指标、没有与一次性视频生成基线的可控性对比、也没有用户研究,而「可控性」和「迭代效率」恰恰是它的核心主张,最该被量化。方案本身也把难题转移而非消除了:把 2D 内容抬升为连贯 3D 世界(单视图/双视图 3D 重建)本身就是未解问题,「冲突感知的双视图初始化」听起来是在缓解不一致而非解决它,一旦初始状态几何有误,后续所有编辑都建在歪地基上。它还依赖现成视频模型做画质增强,那么最终输出的一致性又部分回到了那个不可控的黑箱里。作为范式提案很有启发,作为可用工具还需要更硬的证据。
9. PEAK:擦概念把攻击成功率从96.5%压到5.6%
PEAK: Precise and Persistent Concept Erasure via k-Sparse Autoencoders | 合肥工大·中科大·澳门大学 | arXiv:2608.10985
The main pipeline of the proposed PEAK. (a) A frozen kSAE encodes diffusion-model activations induced by matched target and non-target prompts, and target-specific features are selected by contrasting their activation scores. (b) The selected target features are suppressed during fine-tuning, and the other features are aligned with those of the original model.
Qualitative evaluation of PEAK across different diffusion architectures. PEAK consistently erases target concepts while preserving non-target semantics in both SDXL and FLUX models.
Illustration of traditional UMM evaluation and our Self-Generative-Understanding (SGU) framework. (a) Existing evaluations often assess generation ($\mathcal{M}_G$) and understanding ($\mathcal{M}_U$) separately, producing capability-specific scores such as $s_g$ and $s_u$. These metrics remain valuable for component-wise diagnosis but do not directly provide a system-level view of how a UMM behaves when understanding and generation are jointly involved. (b) SGU provides a complementary closed-loop evaluation framework tailored to UMMs. It requires the model to understand an input image, generate a textual representation, reconstruct a visual context from that representation, and reason over the reconstructed image, yielding an outcome-based system-level score $s_{\text{umm}}$.
评论 (0)