How Solaris runs. Solaris is organized as three levels. The user acts on the frame with mouse and keyboard, the language model holds the state of the session and specifies what each action should mean, and the world model executes that intent, generating both the appearance and the behavior of the result.[-0.05cm]
Interface world models produce more natural interactions. Participants preferred Solaris over coded interfaces both for following the requested interaction and, by an even larger margin, for behaving naturally within the scene, highlighting the ability of interface world models to generate interactions that remain coherent with the environment.[-0.3cm]
Overview of H3-World. (a) Latent-aligned action prompts are independently encoded and packed with the static semantic condition, initial observation, and video latents. (b) MiniMax-H3 processes the packed sequence through single-stream self-attention. (c) LoRA adapts the attention projections under single-egress routing, which connects each action span directly to its matched video latent and retains bidirectional attention among video latents.
Controlled action comparison. The initial observation, seed, and sampling configuration are fixed across rows. Changing the action produces distinct character and camera motion, including stronger fast pans.
Slow-Fast dual-system architecture. The Slow DiT processes observations, robot state, language instruction, along with noisy video and action tokens, to jointly predict future video latents and their corresponding actions (serves only as auxiliary supervision for video-action alignment). Concurrently, the Fast DiT ingests the updated observation and state, conditioned on the Slow DiT's video K/V caches, to generate fine-grained final actions for high-frequency closed-loop control.
Data pyramid. Three-stage training progresses from broad visual diversity to deployment-specific embodiment. romannumeral 1 (Pre-training) learns general visual dynamics from diverse video sources without action supervision. romannumeral 2 (Mid-training) introduces robot actions from multiple embodiments, grounding visual dynamics in cross-embodiment control. romannumeral 3 (Post-training) specializes the model on target embodiments and benchmarks for deployment.
批判点评:这篇把「用无动作视频扩机器人能力」从一个人人认同的方向变成了一条可以摆出来的曲线,36.1%→77.8% 这个数字的说服力在于它是真机零样本,不是仿真里刷出来的。三阶段课程的分工也清晰:预训练学动力学、中训练接地到动作、后训练特化部署,每一步要什么数据、解决什么问题都对得上。Slow-Fast 用 K/V 缓存做跨分支信息传递是个漂亮的工程解——不是简单地把大模型蒸馏成小模型,而是让小模型直接读大模型的中间表示,在单卡 4090 上拿到 30 Hz 是有实际部署价值的。但需要注意几点。第一,36.1%→77.8% 的对照是自身消融(只用目标机器人数据 vs 加 12 万小时视频),不是与其他 VLA 方法的横向比较,所以它证明的是「数据扩展有效」,不能直接读成「ZimaBlue 优于其他方法」。第二,Slow 分支输出的 K/V 缓存有个内在时滞——论文自己用 outdated 标注了这一点,也就是 Fast 分支正在依据一份基于过去观测的世界预测来行动,在场景快速变化时这份缓存的可靠性会下降,而这个失效边界论文没有量化。第三,12 万小时的数据规模本身构成了一道复现门槛,这个量级的第一人称视频不是多数团队能获取的,方法的可迁移性因此打折。第四,署名机构 Joy Future Academy 是个新出现的名字,源码里的作者上标标注的是任务分工(数据、预训练、中训练、后训练、双系统、加速、部署)而非机构隶属,实际的组织背景不透明,只有 Supervision 一栏出现了 Nan Duan。
4. PixSGR:粗尺度注意力指路稀疏精修
Advanced Pixel Diffusion Model with Guided Sparse Global Refinement | 电子科技大学 | arXiv:2609.00798
Overview of the proposed PixSGR framework. PixSGR starts with supervised low-channel bottleneck blocks, expands the channel dimensionality in subsequent Transformer blocks, and performs spatial refinement with sparse attention. Cross-scale token scoring transfers coarse attention patterns to guide fine-scale sparse global interactions. The bottleneck pathway provides an auxiliary clean-image prediction, while a convolutional upsampling head produces the final prediction.
Analysis of the key designs in our PixSGR framework. (a) Bottleneck feature visualization with and without the supervision of $\mathcal{L}_{\mathrm{bot}}$. Compared with the baseline, the supervised feature produces more structured and semantically meaningful coarse representations. (b) Attention map visualization across different Transformer blocks and timesteps. As the network depth increases, attention becomes increasingly concentrated on fewer regions, motivating sparse global refinement under a limited attention budget.
SinkPruner framework. % Left (Visual Sanitizer): As shown in the Attention Ranking (far left), high-norm outlier tokens (purple) dominate the top ranks despite being background noise. We aggregate these outliers into a single sink token (red hashed) and select informative reserved tokens (orange) via attention and similarity filtering. % Right (Text-Guided Pruner): The purified visual sequence further interacts with text tokens (green), utilizing accumulated text-to-vision attention to retain only semantically relevant visual tokens.
Distribution of Visual Attention Sinks in LLM Decoder. % We visualize the relationship between attention weights (y-axis) and massive activation values in sink dimensions (x-axis). % (Left) Standard text-guided methods retain a large cluster of sink tokens (red line, value $>5$), resulting in a high sink ratio of 14.23%. % (Right) Our approach, by filtering visual outliers upstream, suppresses these massive activations, significantly reducing the sink ratio to 3.85% and alleviating attention bias.
Uncovering Understanding-Generation Synergy in Native Unified Multimodal Models: From Representation, Task to System | 南洋理工大学 S-Lab + 商汤研究院 | arXiv:2609.01607
Illustration of different architectures in our study. (a) Dense sharing sends all tokens through the same LLM decoder. (b) Modality-decoupled MoT keeps text in the LLM and sends all visual tokens through a scratch-trained branch. (c) Task-decoupled MoT keeps Und-V language-anchored while specializing Gen-V. For simplicity, we omit the pre-buffer layers here.
实验效果:逐层探针给出的证据是具体的:联合训练确实强化了视觉表示,在语义与几何两类探测任务上都有提升,而且最明显的收益出现在早期与中间层。PCA 可视化从另一个角度佐证——把 patch 级特征投到共享的三维 PCA 基上做 RGB 图,联合训练产生的物体区域更连贯、空间结构更清晰,与冻结探针的结论一致。架构对比说明的是不对称退化的存在与任务解耦 MoT 的缓解作用。系统层对比则说明统一模型的价值不止于接口统一——在需要理解与生成配合的复杂任务上,端到端优化的收益是流水线拿不到的。
Layer-wise probing experiments of visual representations. Higher is better for classification and mIoU; lower is better for depth R-MSE. Joint training strengthens visual representations, with the clearest gains in early and middle layers across semantic and geometric probing tasks.
ReNFT pipeline. (a) Unconditional probes prioritize anti-hub prompts; (b) two policy-dominated mixed routes generate matched proposals from the same prompt and initial noise; (c) reward ranking and the adaptive flipping guard assign pull/push targets for joint-and-paired NFT repair. The VAE decoder is omitted for clarity.
Training dynamics and ablations. (a)(b): PickScore reward and DreamSim-Div; (c)(d): GenEval reward and DreamSim-Div. ReNFT branches from the hacked checkpoint $H$ (star) and repairs for 50 steps; the dashed line marks the base generator. (e)(f): mixed-route pattern and anti-hub ablations after 50 repair steps from the same $H$; the dashed line marks NFT at $H$.
PredErase training-free pipeline. Top: gray-filled $I_{\mathrm{vis}}$ lets frozen I-JEPA predict and cache hole tokens $\mathbf{E}_{\mathrm{target}}$. Bottom: a contact-band prior expands $M_{\mathrm{obj}}$ into the effect-aware Fill support $M_{\mathrm{flux}}$. Center: frozen FLUX.2-klein-4B follows its native trajectory; at $t\in\{4,2\}$, decoded features are aligned with $\mathbf{E}_{\mathrm{target}}$ (Eq.) and projected so that the additional I-JEPA update is zero outside the editable latent support (Eq.). No model weights are updated.
Overview of TimeSteer. (a) Given target intervals, TimeSteer localizes each utterance's source span and relocates its predicted audio-visual content. (b) Source Span Localization estimates source spans from text-to-audio cross-attention. (c) Region-Aware Latent Remapping relocates the content through its read map.
实验效果:在两个代表性骨干上的实验显示,TimeSteer 相对训练无关基线大幅提升区间可控性,同时保持有竞争力的整体生成质量。论文用 HR_0.2(命中率)、IoU 与 WER 三个指标交叉验证,并给出跨五个随机种子的样本级 Spearman 相关性:HR_0.2 与 IoU 正相关说明两个时间指标彼此一致,两者与 WER 负相关说明更好的时间控制往往伴随更低的转写错误率——这条负相关很重要,它排除了「靠牺牲语音清晰度换时间对齐」的可能。定性图展示了对照:不加 TimeSteer 时即使在提示里追加「from [start] to [end]」,语音仍停留在视频开头附近;加了之后台词能落进任意用户指定的目标区间,同时保持生成质量与文本对齐。
Qualitative speech scheduling. Without TimeSteer, a timing instruction of the form “from [start] to [end]” is appended to the original prompt, yet the speech remains near the video onset. TimeSteer places the utterance within any user-specified target interval while preserving generation quality and text alignment. Shading marks the target intervals; frames are sampled at 1,fps with aligned mel-spectrograms below.
批判点评:这篇的方法论姿态很典型也很有效:先在冻结模型内部找现成的可利用结构,再围绕它设计最小干预。两个发现都属于「模型其实已经有了,只是没暴露」——时间敏感的交叉注意力头暴露了隐含时间位置,干净 latent 已经组织好了音画耦合。这让方法不需要任何训练,也解释了为什么它能同时移动语音和口型:因为读取映射作用在时间索引上,耦合关系是被整体搬运的,而不是分别对齐的。映射的分段设计也考虑到了实际问题——目标语音段用仿射保持常速,避免把语速拉长,两侧用三次 Hermite 保证平滑。三指标交叉验证加上 WER 负相关是很扎实的验证方式,它主动排除了「用清晰度换对齐」这种取巧路径。但需要注意几点。第一,整套方法依赖「存在一个时间敏感的交叉注意力头」这个经验性结构,摘要没说清这个头是怎么选出来的、在不同骨干上是否稳定存在,若换模型需要重新定位这个头,方法的通用性就要打折。第二,SpeechShift 是自建基准,评测标准与难度分布都由作者定义,缺少第三方对照。第三,重映射搬运的是已有内容,因此目标区间的时长必须与源跨度基本匹配——若用户指定的区间明显短于或长于模型自然生成的语音长度,仿射段就得压缩或拉伸语速,这个失效边界摘要没有量化。第四,验证只在两个骨干上,且都属于联合音视频扩散这一类。
评论 (0)