f-loss(LIX 巴黎综合理工学院(CNRS, IP Paris)、AMIAD、LIGM 巴黎国立桥路学校、加州大学伯克利分校):频域配平让收敛快40%
MeRoPE(香港科技大学、卓驭科技):大基线相机控制不再爆范数
GlyphAnchor(复旦大学、小红书、上海创智学院):字形块锚定长文本渲染
今日论文速览
1. SolarWM:5秒训练撑起一小时世界推演
SolarWM: Open Data and Scalable Training for Long-Horizon Video World Models | 香港中文大学(深圳)、SLAI、新加坡国立大学、香港中文大学、香港科技大学、香港科技大学(广州)、NVIDIA、UCLA、微软亚洲研究院 | arXiv:2609.02886
Hour-scale world rollout generated by the SolarWM-wan2.2-5B-fast causal student. Sparse frames are sampled throughout a single uninterrupted autoregressive session initialized with a real first frame from the held-out validation pool, spanning the initial observation to the 60-minute endpoint. The scene prompt remains fixed and the model follows a predetermined camera trajectory.
Overview of SelfLift. Given the predicted low-resolution clean endpoint, SelfLift-zero constructs a trajectory-preserving latent candidate and a VAE-reachable pixel anchor. Their consistency residual yields an artifact-risk map and adaptive weights for selectively correcting high-risk regions before re-noising. SelfLift-rich distills this correction into a lightweight latent lifter and performs on-policy recovery on student-visited states, using an EMA self-teacher for privileged high-resolution guidance and a dynamics objective to preserve the pretrained few-step trajectory.
(Left) Images generated with our XL model at 256 resolution. Beware, 2 impostors (real images from the ImageNet training set) are hiding in these images. Place your bet on which ones they are and check the solution on page 10. (Right) FID vs. training epoch for an XL model. Our achieves faster convergence than JiT without requiring architectural modifications.
批判点评:这篇的诊断比方法更值得记住。把「细节学得慢」归因到损失函数对频率的隐含加权,而不是归因到容量或架构,是一次位置正确的归因——1/f² 是自然图像的统计事实,像素损失对空间误差的均匀对待则是实现上的默认选择,两者相乘就必然产出低频偏置。论文用两层证据把这条推理钉住:真实数据上功率谱的系统性偏差,以及只含两个频率的玩具实验里高频模式的完全缺席。后者尤其干净,因为它排除了容量不足这个替代解释。方法侧的克制也是优点:不改架构、可直接替换损失,这类改动的采用成本极低,而先频域后像素的两段式安排还有收敛曲线作为依据,不是拍出来的。带通 FID 分析是全篇最诚实的一段——它主动说明低频和中频最终会收敛到同一水平,优势集中在高频和更早的收敛时点,这比只报一个总 FID 要可信得多。几处需要留意。「最多 40%」是上界,跨规模的期望收益应当打折看;而高频段上基线中途曾追上、之后又被反超,说明这条优势并不单调,实际收益取决于训练预算落在曲线的哪一段。更实质的是适用范围:诊断成立的前提是像素空间训练,而当前主流做法大量在 latent 空间做 flow matching,latent 的谱统计与自然图像并不相同,这套配平能否平移过去,论文没有回答。还有 f-loss 引入的权重超参需要调(论文有相应消融),「drop-in」的说法在工程上仍有调参成本。最后必须提醒:源文件中带通 FID 与定性收敛这两张图的图注仍留着作者的「Placeholder / to be updated」字样,这是预印本未定稿的迹象,引用这两处时应当谨慎。
4. MeRoPE:大基线相机控制不再爆范数
MeRoPE: Metric Rotary Position Embedding for Camera-Controlled Video Generation | 香港科技大学、卓驭科技 | arXiv:2609.01252
Qualitative video generation under camera pose control. Visual results on two representative nuScenes scenes under three commanded trajectories (original, left turn, right turn). In each row, the leftmost plot displays the bird's-eye view of the commanded path (solid) and the path recovered from the generated video by VGGT-$\Omega$ (dashed), followed by generated video frames sampled at $t \in \{0, 16, 32, 48\}$.
Overview of GlyphAnchor. Prompt, optional source images, and rendered glyph patch images are encoded into tokens and concatenated before a text-to-image or image-editing diffusion-transformer backbone. The resulting glyph patch tokens are placed in an additional condition frame, forming a virtual glyph plane whose coordinates are anchored to the target layout.
Overview of CameraEditor. (a) Generation Pipeline: A video diffusion model generates the target Chain of Frames (CoF) sequence conditioned on the source image and reference visual sequence. (b) Dynamic Routing: A routing mechanism selects the optimal reference prior via GeoCalib during inference for precise viewpoint alignment.
实验效果:论文报告 CameraEditor 在相机控制精度与源身份保持两方面达到当前最优,优于既有方法。证据形态上,定性结果与几何误差图并排给出:每个方法的右侧面板分别显示 PF-I(上)与 PF-T(下)误差,暖色表示偏差更大,也就是说几何准确性是用透视场(Perspective Field)热力图来验证的,而非仅凭肉眼。扩展对比进一步显示,该框架在与目标相机几何对齐的同时,避免了标准扩散编辑中常见的身份漂移与结构幻觉。两组消融值得注意:一是相机感知模块的选择——随机路由或基于视觉语言模型的理解(Puffin 变体)给出的错误初始估计会导致灾难性的结构失真,而 GeoCalib 驱动的路由提供了稳健的几何锚点;二是中间过渡帧数量 N 的消融——直接单步生成(N=0)在大幅视角变化下出现空间撕裂与身份丢失,把 N 增加到 4、8、12 则提供了更强的连续几何先验。
Qualitative results alongside geometric error maps. For each method, the right panels display PF-I (top) and PF-T (bottom) errors, where warmer colors indicate larger deviations.
Pipeline of LightBridge. The framework first uses Multi-Illumination Relighting Dataset for paired supervision, then extracts the 2D Visual Token with Latent Bridge Relighting Diffusion, and finally propagates it to the full 3DGS representation with Gaussian Propagation Transformer.
Qualitative comparison at viewpoints outside the relighting video trajectory. From left to right, each example shows the source 3DGS rendering, the result rendered from the representation optimized by GR3EN, and the corresponding rendering from the relit 3DGS predicted by LightBridge under the same camera pose and target lighting. % Note that GR3EN fails to relight the scene whether the target light is visible (top row), as well as when it is not directly visible (bottom row)
批判点评:这篇的取舍很清楚,也因此值得读:它不追求重光照质量上的领先,而是把「免逐场景优化」这一条做成硬指标。这个选择是对的,因为生成式重光照真正的落地障碍从来不是单张图好不好看,而是每换一个场景都要再优化一轮——在资产规模上这条成本曲线根本压不下来。两个部件都服务于这个目标:把重光照建模成潜空间里源到目标的传输,从而一步取出视觉 token、免掉迭代采样;再用先「图像到点」稀疏自注意力、后「点到图像」交叉注意力的方式把线索铺到完整 3DGS 上,规避全注意力的组合爆炸。数据集的构造尤其体现方法意识:固定几何与相机位姿、只变光源状态/强度/颜色,这种控制变量式的成对监督正是前向模型能学到「只改光、不改结构」的前提。评测设计里最有判断力的一处是选在重光照视频轨迹之外的视角做对比——这直接检验重光照是否真的落进了 3D 表示,而不是只在被监督到的视角上成立,很多同类工作恰恰在这里含糊。需要看清几处边界。第一,摘要自己用的是「有竞争力的质量」,也就是承认这是以效率换相当质量;如果下游对光照真实感的要求高于对吞吐的要求,逐场景优化路线仍可能更合适。第二,数据集是作者自建的合成室内场景(卧室、餐厅、厨房这类),室内合成数据的光照统计与真实采集差距不小,而重光照恰恰高度依赖材质与间接光的真实性,因此室外场景与真实数据上的表现完全未知。第三,前向模型的可控性边界值得追问:它能处理的是「现有光源的状态、强度与颜色」这类控制,是否支持新增光源、改变光源位置或换成完全不同的光照环境,摘要没有说清。第四,一步式的潜空间传输省掉了迭代采样,但也放弃了扩散采样在质量上的调节余地——没有「多花几步换更好结果」这个旋钮,在困难样例上没有退路。最后,代码与数据集要等录用后公开,当前无法独立复核。
Architecture of VibeVoice-ASR-Streaming. Speech chunks $X_k$ and speaker-attributed text chunks $Y_k$ are interleaved in a single autoregressive context, and each chunk is followed by a fixed $L=4$-frame (0.5 s) lookahead before its text is generated.
实验效果:识别精度方面,7B 模型在五个评测集上取得了最低的平均 WER/CER。说话人归属方面,在 13 个评测设定中的 12 个取得最佳或并列最佳。论文的对比图把范围说得更具体:与四个已部署的商用流式 ASR 系统在四个会议基准与 MLC-Challenge 上比较,其中 AliMeeting 与 AISHELL-4 用 CER 评估、AMI-SDM 与 AMI-IHM 用 WER 评估;MLC-Challenge 里日语和韩语用 CER、其余七种语言用 WER。为保证可比性,所有在线 API 的音频都按 2.9 秒的块流式送入,与本模型的块大小对齐。论文另外报告了单说话人识别结果,但明确说明这是作为一种类别检查而非竞争性主张来呈现的。
Recognition error of VibeVoice-ASR-Streaming-7B and four deployed streaming ASR systems on the four meeting benchmarks and MLC-Challenge. AliMeeting and AISHELL-4 are evaluated with CER, while AMI-SDM and AMI-IHM are evaluated with WER. For MLC-Challenge, Japanese and Korean are evaluated with CER and the remaining seven languages with WER; the reported value is the macro average over the nine evaluated languages.
本文贡献:Kirin 是一条从视频到可渲染动画资产的完整链路,包含重建、学习先验、生成三段。第一段是数据:利用大量野外动物视频重建出 3D 运动序列并配上字幕,构成 AiM3D——论文称其为首个为四足动物提供对齐的「视频-文本-运动」三元组的大规模数据集。重建这一步的做法是把问题拆开:借现成的 3D 四足动物重建方法和 3D 跟踪方法分别推断关节运动与全局位移,再把两者合起来得到最终的运动重建。这个拆分很关键,因为「身体怎么动」和「整体往哪走」在野外视频里的可观测性完全不同。第二段是模型:在 AiM3D 上训练一个视觉引导的运动生成模型,同时以文本和图像为条件——图像条件不是可选项而是有实质作用的,论文的消融显示在相同文本提示「一只动物在行走」下,不同输入图像会带来相应的骨架变化与步态差异,也就是说图像承担了「这是什么动物、身体比例如何」这部分信息。第三段是落地:借一个现成的图像到 3D 模型自动完成绑定与动画,把生成的运动直接套到 3D 网格上,产出可直接渲染的动物动画。该文已被 ECCV 2026 接收。
Left: Overview of Kirin generation pipeline. A text description and an image are provided as inputs. The text and image are used for motion generation, while the image is also used to generate a T-posed mesh. The animation module then rigs the generated motion onto the mesh to produce the final animated 3D model. Right: Overview of the motion generation architecture. Text features are extracted using a frozen DistilBERT encoder, and image features are extracted using a frozen DINOv3 encoder. The text, image, and denoising step embeddings are combined and fed into a transformer decoder with cross-attention to generate motion sequences. % annotate the notations in the figure, add subtitles for left and right parts, use darker gray color for arrows, texts are too small
实验效果:论文与两个基线做了定性对比,两组都直指对手的具体失效模式。与 Puppeteer 比的是网格动画:在相同的文本与图像输入下,Puppeteer 常常在输入网格上产生很小甚至没有运动,而本方法生成的动作会跟随文本提示。与 AniMo 比的是骨架运动:AniMo 只用文本,本方法同时以文本和图像为条件;AniMo 出现的失败包括运动不遵循提示词以及骨架形状不一致,而本方法产出的运动更真实。数据侧论文给出多层展示:动物关节运动的数据样例(每行左侧是两段运动的文本描述,随后是视频帧与对应的重建 3D 运动)、全局位移的可视化(把根关节轨迹投影到地面以显示整体位移)、以及数据集统计——各动物类别的视频数与帧数分布、以及全数据集上的运动类型分布(每个视频被赋予一到三个运动标签)。另有一组补充结果专门展示全局运动更显著的例子。
Visual comparison of generated mesh animation with the baseline. The left columns show the input text and image, which are used for both pipelines. Puppeteer often produces little or no motion on the input mesh, whereas our method generates realistic movements that follow the text prompt.
批判点评:这篇解题的思路很值得学:面对「受控采集在动物身上不可行」这个硬约束,它没有去改进采集,而是把已经海量存在的野外视频当成数据源,用重建把无标注视频转成运动监督。这一步走通了,整个循环就被打开了。重建阶段把「关节运动」与「全局位移」分开处理、再合并,是有针对性的设计——野外视频里身体姿态和整体位移的可观测性差别很大,混在一起估计容易互相污染。图像条件的引入也不是堆砌模态:消融显示同一句「一只动物在行走」配不同图像会产出相应的骨架与步态差异,说明图像承担了物种与身体比例这部分文本说不清的信息,这正是动物域比人体域更需要视觉条件的原因。最后接上现成的图像到 3D 模型自动绑定,把产物一路推到「可直接渲染」,让这项工作对动画流程有实际可用性,而不是停在骨架序列。两组基线对比也选得准,各自指向具体失效模式:Puppeteer 几乎不动,AniMo 不遵循提示且骨架不一致。需要清楚的是证据形态与数据性质。第一,摘要层面给的主要是定性对比和数据集统计,缺少定量指标——运动生成有成熟的评价手段(关节误差、足部滑动、用户研究胜率),没有这些数字时「更真实」难以校准。第二,也是更根本的:AiM3D 的「真值」是用现成的 3D 重建与跟踪方法从野外视频推断出来的,不是动捕级真值,因此数据里必然带有上游工具的系统性偏差,而在这份数据上训练的模型会继承这些偏差;这不否定其价值,但意味着它衡量的是「与重建结果一致」而非「与真实运动一致」。第三,覆盖面限定在四足动物,鸟类、鱼类等形态差异大的类别未涉及,而这些恰恰是动画中同样常见的需求。第四,野外视频的运动分布天然偏向常见行为(行走、奔跑),罕见动作与高动态动作的样本很可能稀疏,数据集的运动类型分布图应能反映这一点。最后,自动绑定这一环依赖外部图像到 3D 模型,其网格质量与拓扑会直接决定最终动画的可用性,这部分不在本文的控制范围内。
10. RIG-BENCH:2000题测生成模型会不会推理
Thinking in Pictures: A Systematic Benchmark for Reasoning-driven Image Generation | 伊利诺伊大学厄巴纳-香槟分校、纽约大学 | arXiv:2609.02864
End-to-end pipeline of RIG-Bench. (1) Source Task Interface: items are drawn from established visual reasoning benchmarks and manually re-cast with an image-based output prompt (no options exposed to the model). (2) Unified Generation Protocol: selected models receive the context images plus the prompt and generate the answer image directly. (3) Scoring and Reporting: each generation is compared against the ground-truth image using both an LLM judge over a hand-written rubric and automatic perceptual metrics, with an additional human study for calibration.
Distribution of the $2{,}000$ samples in RIG-Bench. % The inner ring summarizes the four task families; the outer % ring breaks them down into eleven fine-grained subtasks.
评论 (0)