Unified visual backbone for T2I generation and image editing. Semantic conditions from Qwen3-VL and image latents from FLUX.2 AE are jointly processed by parallel single-stream Transformer blocks.
Comparison for parameter and overall edit score with open-source models. The overall edit score is computed as the average across three public editing benchmarks and our benchmarks. 本方法 achieves leading performance with only 6B and 3B parameters.
Static-3D rewards freeze the scene; a 4D reward keeps it moving. Five uniformly-spaced frames from a $10.3$,s LongLive rollout (“a cat running away with a fish while people chase behind”). The distilled base contains large motion but the cat and fish drift across frames; the static-3DGS rewards World-R1 and VideoGPA reduce the scene to a more rigid, low-motion configuration. Stream4D keeps the cat running with clear forward motion while preserving a coherent subject and scene.
实验效果:定性结果是这篇最有力的部分:三场景 × 三时刻的对比图里,两个基于 3DGS 奖励的基线(VideoGPA、World-R1)在全部三个场景都被判定为 FROZEN,而 Stream4D 在全部场景保持 KEEPS MOTION;十秒时长的长程对比进一步显示基线在 t=5s 之后主体位移基本停滞。原始 Base 模型虽然 MOVES,但可以看到画面结构在后续帧出现明显漂移与内容突变。也就是说 Base 有运动但没有几何一致性,3DGS 奖励方法有一致性但杀掉了运动,Stream4D 试图同时拿住两者。
本文贡献:4DAnyone 的框架从源视频出发,一路 VAE 编码得到 Src Tokens,另一路过 HMR 模型估计人体,把目标视角的骨架渲染图经可训练的 Skel Enc 编码为 Skeleton Tokens。参考上下文侧维护一组 Ref Tokens 并配合 Pack Ref Context 机制压缩,与 Noise Tokens、Skeleton Tokens 一起拼接后送入 Diffusion Transformer。主干里每个 block 包含三种注意力,其中 Video Attention 与 Cross Attention 保持冻结、只有 Multiview Attention 可训练,作用维度上前者按 v(f h w) d 组织、后者按 f(v h w) d 组织——也就是说把「同一视角内跨帧」和「同一帧内跨视角」两种一致性显式分离到不同的注意力上,训练成本集中在真正需要新增的多视角一致性上。生成的多组目标视角视频最后统一用来训练 4DGS,得到可自由视角渲染的动态人体。
Overview of 本方法. Given a source video, we extract a 3D skeleton sequence via GVHMR and render skeleton videos and skeleton depths for $v$ target viewpoints. The 3D-aware skeleton encoder injects skeleton tokens into the DiT via residual addition. The source video is encoded by the pretrained VAE and concatenated with noisy target latents along the view dimension. The DiT denoises the target latents via Target Context Routing and the decoded target videos are packed into the Reference Context Packing module as references for subsequent generation rounds. The final 4DGS model is reconstructed from the generated multi-view videos using FreeTimeGS.
Qualitative comparison with baselines. We show target-view generated videos (Gen.) and their corresponding 4DGS renderings (Rend.) across diverse human-centric videos. 本方法 produces geometrically accurate and visually detailed results across viewpoints, while the baselines suffer from inaccurate camera control (our fine-tuned ReCamMaster$^\dagger$) or geometric distortions (MV-Performer). See the project page for dynamic results.
WithEveryone: Unified Planning and Identity Grounding for Group Image Generation | 复旦大学、腾讯混元、香港大学 | arXiv:2608.20336
关键词:身份保持生成,群体图像,布局规划,身份接地,ID Loss,统一骨干
前序问题:身份保持的图像生成在单人场景已经比较成熟,但一旦场景里要出现多个指定的人,可靠性就迅速下降。难点不止是「每张脸都要像」,还要把每个参考身份绑定到画面里一个特定的人和特定的位置——模型很容易把 A 的脸特征渗到 B 身上,或者干脆漏掉某个参考。更棘手的是训练侧:常用的身份损失需要在若干张噪声预测出来的人脸之间建立对应关系,而在多人场景下这个对应本身就是模糊的,监督信号会被错配稀释,等于用一个不可靠的目标去训一个本就困难的任务。
Overview of 本方法. The model loads selected reference identities as ID tokens, predicts structured Layout CoT, renders a visual layout condition, and predicts identity representations before generating the target group image. Modality-switch tokens that mark the boundaries between autoregressive reasoning and flow-based image generation are used in the sequence but omitted from the figure.
实验效果:方法支持最多十个参考身份的群体图像生成,配套的画廊素材(gallery1-3、quality)展示了多身份合影场景下的生成结果与质量对比,另有单步预测可视化与数据构造流程说明。消融部分覆盖较全:人脸尺寸的绝对/相对影响、高分辨率下的身份与规划表现、ID Loss 的布局项与权重设置、Rep Forcing 的余弦相似度与损失曲线,都各有独立分析图,说明作者对每个组件的作用边界做了系统检验。
Qualitative comparison with proprietary models. mark faces that reproduce a person already generated elsewhere in the image, that are not recognizable as their bound reference, and that appear transplanted without adapting to the pose, lighting, or viewpoint of the scene. These author annotations are illustrative; their quantitative counterparts are the duplicate rate, Sim(Ref), and Copy-Paste.
批判点评:这篇抓住的痛点很实在:多人身份保持的真正瓶颈不在人脸编码器强不强,而在「参考-人物-位置」这个三元绑定关系没有被显式监督。把布局作为中间产物先规划出来、再用人脸区域标注把 ID Loss 接地到具体位置,是对症的解法,比在损失里加更多身份约束更有希望。用离散坐标 token 让统一骨干直接输出布局规划,也避免了外挂一个布局模型带来的接口割裂。但几点需要保留判断。第一,整条链路依赖人脸区域标注来提供 Layout-Grounded ID Loss 的监督,这类标注在训练数据规模化时成本不低,且标注质量直接决定接地的准确性,论文的数据构造流程需要仔细看清楚这一步是人工还是自动、错误率多少。第二,「最多十个身份」是个上限声明而非能力曲线,实际效果在 2 人、5 人、10 人时大概率显著不同,摘要没有给出随身份数变化的量化衰减,而这恰恰是读者最关心的。第三,Rep Forcing 引入了额外的表征对齐损失,其权重与 ID Loss、NTP 损失、Flow Matching 损失之间的平衡是四项损失的调参问题,消融给了曲线但整体的调参难度可能被低估。第四,项目页与代码标注为「will be released」,当前无法验证,多阶段序列的实现细节(尤其坐标 token 的词表设计)对复现影响很大。
5. ear-VAE2:波形自编码器缺一根频率轴
Fourier is Frontier: Frequency-Aware Autoencoding for High-Fidelity Music Reconstruction | 阿里巴巴通义千问团队、Monash University | arXiv:2608.19843
Architecture of 本方法. Stereo audio is transformed via STFT into a complex spectrogram; a Spec Encoder compresses it to a 25,Hz continuous latent and a Spec Decoder reconstructs a coarse complex spectrogram; the highest-frequency (Nyquist) STFT bin is removed to facilitate frequency downsampling. The subsequently applies band-specific complex corrections and restores the Nyquist bin to produce the full-bin spectrogram, which the iSTFT resynthesizes into a waveform.
Latent temporal-frequency probe for one track (Track 1004034; temporal split at 0.5). Rows (top to bottom): ground-truth audio; full reconstruction; low-temporal-frequency latent half decoded; high-temporal-frequency latent half decoded. Columns (left to right): 本方法, LeVo~2, SAME-L, and Stable Audio Open. For 本方法, the rapidly varying (high) latent half decodes to audio with a lower spectral centroid, matching Table.
TextRefine: Improving Textual Fidelity, Spatial Placement, and Glyph Rendering for Text Editing in Product Posters | 快手科技、中国科学院大学人工智能学院 | arXiv:2608.19637
Comparison of binary OCR correctness and the proposed graded glyph signal. Whereas a binary reward retains only the final correctness decision, the target-glyph CTC posterior preserves continuous recognition evidence and provides a more informative signal for optimizing glyph fidelity.
Block3D overview. Text and optional bounding-box conditions guide left-to-right block generation. Completed blocks form a frozen causal prefix, while M2T and T2T edit only the active block before the frozen decoder maps the completed codes to a mesh.
Additional Block3D text-to-shape results. Each group contains multiple viewpoints of a single generated mesh, with its complete text prompt shown directly below.
批判点评:块状生成这个思路本身很自然——它就是在「逐 token 串行」和「全表示并行」这两个极端之间插入一个可调的中间点,用块大小当旋钮同时控制串行深度和单步成本。从工程角度这是个稳妥的改进,也容易与现有的形状 token 化方案组合。定性结果里对复杂组合描述(多部件、对称约束、计数约束如「两个头两条尾」)的服从度看起来不错,这类计数与对称约束恰恰是纯自回归容易翻车的地方,块内联合去噪的纠错能力在这里应该确实有帮助。但要保留几点。第一,块大小是核心超参,它直接决定效率-质量权衡的位置,论文若不给出块大小扫描曲线,读者无法判断报告的结果是在一个精心挑选的甜点上还是普遍成立。第二,效率对比表只给了端到端生成时间,缺少显存占用与质量指标的同表对照——单看时间快了但质量掉了多少,需要交叉验证。第三,块间自回归意味着误差仍会沿块序列累积,只是粒度变粗了,长序列(高分辨率形状)下这个累积是否可控,论文未专门检验。第四,作者机构跨六家单位,实验在单张 A100 上完成,规模适中;与工业级 3D 生成系统(如商业 3D 资产工具)的差距如何,缺少对照。
前序问题:音频驱动的唇形同步任务是把说话视频的嘴部区域改成匹配驱动音频,同时保持头部姿态、身份和背景不变。这在定义上是一个局部编辑任务,但主流做法却是用重型的 GAN 或扩散解码器把整个下半脸重建出来。代价有两层:一是延迟高,难以实时;二是更要紧的,重建式解码器会「幻觉」出口内细节——牙齿、唇纹这些高频结构被生成器凭先验编出来,而不是保留原视频里真实存在的纹理,结果是身份感和真实感都受损。作者的判断是,身份保持的瓶颈并不在参考帧数量不够,而在缺一个能把参考帧里已有的真实纹理忠实搬运过来的机制。
The framework or our EfficientSync. (a) STAR Sampling first selects, from the source video, a reference pool that is both sharp and topologically diverse, supplying high-quality material for deformation at no inference cost. (b) The Dynamic Texture Mixer then aggregates the spatially aligned references into a high-fidelity mouth feature through a context-aware, channel-wise selection mechanism. (c) The Spatio-Temporal Shifted Adaptive Masking strategy finally composites the synthesized mouth back into the source frame, blending it seamlessly with the preserved background.
实验效果:效率对比在单张 A100 上进行,给出 FPS 等指标的完整表格,与主流唇形同步方法横向对照,主张达到实时。质量侧配有与 SOTA 的定性对比(vssota)、路由策略对比(comparison_routing)、动态纹理混合器的动机说明(dtm_motivation)与用户研究结果。核心主张是:在保持实时性的同时,口内细节来自真实参考纹理而非生成幻觉,因此在身份保真与真实感上优于重建式方法。这篇为期刊格式(含作者简介),实验组织相对完整。
Qualitative comparisons with state-of-the-art works under the cross-audio setting. EfficientSync preserves sharp, identity-consistent intra-oral textures, whereas competing methods either blur the mouth region or hallucinate teeth that deviate from the source identity.
Overview. We augment the audio in our dataset resulting in motion sequences aligned with multiple audio variations to ensure generalization. Raw audio is then processed into explainable features to condition a diffusion-based Transformer Decoder. We employ a hybrid pose representation, using joint rotations for the body and Cartesian positions for drumsticks. The model is trained via a dual-objective loss to balance natural body dynamics with high-precision stick impacts. Finally, performance is evaluated using Impact Point Deviations and Percussive Alignment Scores to assess spatial and temporal fidelity.
Our model achieves temporal alignment comparable to ground-truth data. PAS distributions as violin plots for GT (leftmost, black), noisy GT (second and fourth, red), rotations-only model (middle, purple), and our model (rightmost, orange), when applied to our test dataset. Each sample is one-second long. Noisy GT is obtained by adding Gaussian noise with standard deviation $\sigma$=50 ms and $\sigma=25$ ms to every motion onset time. When noise is added to the data, PAS decreases, indicating that the metric successfully discriminates audio-motion alignment quality.
评论 (0)