Qualitative comparison with various methods on the streaming video editing task, where a sequence of editing instructions (here camera-movement and style edits) is applied continuously to a scenery source video. All methods share the same source video and editing instructions. Methods marked with $\dagger$ take no source video as input, while all other conditions are kept identical for fairness.
Overview of instruction-aware exact caching. Top: Static text anchors condition the reference branch once during cache construction. Only the resulting reference K and V are retained and reused across denoising steps. Bottom: Comparison of attention structures, where rows denote queries and columns denote keys and values.
Grounded-Exo2Ego: Structured Semantic Grounding for Robust Exocentric-to-Egocentric Video Generation | NVIDIA | arXiv:2608.20534
关键词:第一视角生成,视角转换,语义锚定,3D重建,NVIDIA
前序问题:从一段第三人称(exocentric)视频生成对应的第一人称(egocentric)视频,对 AR/VR 与具身智能都很关键——第一视角数据的采集成本远高于架好机位拍摄。但这个任务比常规的新视角合成难得多:视角变化极端,标准的几何条件(把场景重建成 3D 再从目标位姿渲染)在这种大跨度下变得高度不可靠,而且会出现大量在源视频里根本没被观察到的区域,几何重建对这些地方无能为力——渲染出来就是一片空洞。
本文贡献:作者在架构和数据两个层面同时下手。架构上是一个双分支的视频扩散模型:几何锚定分支把第三视角的颜色和深度抬升成 3D 重建、在目标 ego 位姿下渲染,用这个渲染结果条件化生成;语义锚定分支则跳出了主流的纯几何路线——用 captioning 提取逐物体的短语,经 LLM 转成物体 token,配合可学习的 register,通过带掩码的交叉注意力注入,让模型基于物体级上下文来合成那些几何搞不定的困难区域。另一个被忽视但影响巨大的发现是相机-重建失配问题:重建出的 3D 与相机位姿之间存在系统性错位,会严重拖累 exo-to-ego 的学习,他们为此设计了一个相机重定位算法,修正后全部指标都有实质提升。数据层面则做了一个全自动的合成数据引擎,在程序化生成的环境里生成并渲染绑定骨骼的 3D 角色。
Method overview. Geometric anchoring (top): Exocentric color and depth are lifted to a 3D reconstruction and rendered at the ego pose. Semantic grounding (bottom): Per-object context is extracted as segmentation masks and object tokens. The segmentation masks are reprojected into the ego-view as used as attention masks during cross-attention. The per-object masks encourages noisy tokens to attend to object information relevant to the specific region, providing spatial structure that grounds semantics onto pixels.
An overview of the Libra architecture. The core design is the switch attention and the switch FFN that decoupling self-modal modeling and cross-modal interaction. We build two variants, Libra-1 and Libra-2, to discuss several improvements in tokenization, positional encoding, and supervision.
Comparison on text-to-image generation. We highlight the key parts of the text prompts, where Libra-2 demonstrates strong prompt-following capacity. % shows several text-to-image generation results of Libra-2. More results are presented in Sec.. In Fig, we compare the text-to-image generation results of Libra-2 with two baselines, StableDiffusionXL and Show-O, where StableDiffusionXL is a widely-used strong text-to-image generation model, and Show-O is a unified MLLM that use the most closed training
MultiCube: Compositional 3D Generation With Part-Level Semantic and Spatial Control | Roblox、卡内基梅隆大学 (CMU)、斯坦福大学 | arXiv:2608.20448
关键词:组合式3D生成,部件级控制,布局适配器,两阶段扩散,Roblox
前序问题:游戏和动画里用的 3D 资产通常需要是组合式的——拆分成有语义意义的部件,这样才能单独换材质、做动画、改结构。最近的 3D 生成方法已经能从图像或文本 prompt 生成高质量的组合式物体,但这类全局条件缺少专业创作流程真正需要的部件级可控性:你可以说「一门带黄铜炮管和木轮底座的加农炮」,但没法指定炮管具体多长、轮子放在哪个位置。
本文贡献:MultiCube 的输入设计很直白:一个全局文本 prompt、一份指定期望部件的文本 schema、以及一个用包围盒标出各部件空间位置的布局。输出是由多个独立网格组成的 3D 物体,一个部件一个网格,且同时满足语义和空间条件。方法用两阶段扩散:第一阶段生成与 schema 和布局对齐的整体网格(monolithic shape),第二阶段把这个网格同时分解成各个部件——注意是同时而非逐个,因为部件之间存在相互约束。核心组件是 Part Layout Adapter,它独立于其他部件来编码每个部件的条件(文本标签走 Qwen-VL,包围盒走 Fourier + Linear),这个独立性保证了改动一个部件的条件不会牵动其他部件。第二阶段的多部件 DiT 之间插入 Cross-Part Attention,让各部件在生成时互相感知、保证接缝处的几何一致。
MultiCube uses a two-stage diffusion-based architecture. In Stage 1, it synthesizes a monolithic object conditioned on the input text prompt and part layout (a). In Stage 2, it takes the object latents produced in Stage 1 and decomposes them into multiple shapes, one for each specified part, guided by the part labels and bounding boxes (b). Spatial conditions are encoded with a Part Layout Adapter in Stage 1 and a lightweight embedding in Stage 2.
实验效果:实验显示方法能生成高质量的组合式 3D 物体并实现精确的部件级控制,包括那些仅靠文本或图像 prompt 难以达到的独特布局。定性结果覆盖面很广:婴儿车(车架/把手/座椅/顶篷/轮子五部件)、滑板(板面/桥/轮)、手电筒、树蛙、工具箱、摩托艇(船体/舷外机/螺旋桨/方向舵/挡风/座椅六部件)、黄铜提灯,以及喷气背包、加特林、火箭、军用水壶、飞碟、腕式弹弓、长剑等。每个例子里生成的形状都能看出与输入包围盒布局的对应关系——比如飞碟的「旋转外环」确实是一圈环绕结构、「诱捕光束发射器」在底部。
Full pipeline generation results. MultiCube generates compositional 3D objects from input text prompts and part layouts. Here, the part bounding box layouts are automatically generated using GPT-5.1; hence, the generation process from prompt to output shape is fully automatic.
Overview of Xemo-Talker. (a) The Geometry Motion Predictor learns emotion-agnostic audio-driven motion through diffusion reconstruction, producing the complete 70-D facial motion sequence. (b) With the predictor frozen, the Geometry Emotion Branch encodes the target emotion and injects multi-scale residual features through zero-initialized adapters to refine the motion prediction. (c) The subspace-aware Tri-Loss organizes global emotion representations using classification and prototype alignment, while applying contrastive discrimination to the less-principal PCA projection. The predicted motion is converted into deformed keypoints and rendered by the frozen LivePortrait renderer.
实验效果:在情绪分类准确率上达到 SOTA,同时保持有竞争力的唇形同步和高推理效率,性能接近在真实视频上测得的水平。定性对比里最有说服力的是 Fear 与 Surprised 这一对——这两种情绪在面部表情上高度相似(都是睁大眼睛、张嘴),是情绪控制里最容易混淆的一对。图中同一身份在两种标签下生成的结果确实呈现出可区分的差异:Fear 时眉毛内侧上扬、嘴角向下、整体表情收紧,Surprised 时眉毛整体上抬、嘴张得更圆、眼睛睁得更大。男女两个身份上这个区分都成立。
Visual comparison of Fear and Surprised expressions from MEAD.
Overview of OccluRank. OccluRank constructs spatially aligned instance features, incorporates ordinal rank embeddings, and performs location-wise cross-instance interaction before aggregating into an Instance Semantic Map (ISM). The masked ISM is residually injected into selected cross-attention layers of SDXL.
评论 (0)