Overview of LDR. Given three conditioning frames $I_0,I_1,I_2$, LDR predicts the future frames $\hat{I}_3,\hat{I}_4,\dots,\hat{I}_T$ in three stages. (A) LDR first encodes each input frame into a structured latent (SL). (B) LDR measures the low-order time derivatives of the given SL, then rolls out by regressing only the high-order residual with $f_\theta$ and numerically integrating the lower orders to the next latent $\hat{\boldsymbol{s}}_t$ ($t\in\{3,4,\cdots,T\}$). (C) LDR finally decodes each predicted latent to an RGB frame $\hat{I}_t$.
Stress test under severe OOD shift. Trained only on red balls but tested on an unseen object (e.g., “earth”), LDR still predicts the correct motion, while the others fail.
本文贡献:提出 Ex-Omni-2D,给定多模态查询、参考图像和参考音频,协同生成文本、个性化语音与参考条件视频三位一体的回复。模型先预测一个结构化的 Visual Thought Plan(VTP)描述场景、情绪与动作,再生成回复文本和原生多码本语音单元;这些语音单元构成共享的声学-时间接口,既解码为语音又与视频帧在线对齐。正是这个接口让回复通路与形象通路能各自从异质的语音、对话、形象视频数据中学习,从而绕开四元组大规模监督。全序列 Video Generator 作为主教师,进一步蒸馏成少步的块因果 Streaming Student,其 Prefix Streaming 机制把干净潜变量跨相邻块携带,抑制后段块的累积退化。
Overview of Ex-Omni-2D. Speech, text, and vision encoders provide multimodal context to the Large Language Model, which produces a non-user-facing Visual Thought Plan (VTP) and the user-facing response. The Speech Generator uses reference audio for voice conditioning and generates speech in a native multi-codebook space. The Video Generator encodes the reference image and VTP, and reuses the generated speech representation to synthesize the synchronized video response. It is instantiated as a full-sequence Teacher for primary realization and a distilled Prefix-Streaming Student for efficient incremental deployment.
Prefix-Streaming analysis. Top: representative later-chunk examples; each row shows the reference image, a no-prefix frame, and its Prefix-Streaming counterpart. Bottom: full-duration DINO curves over 200 paired CommonEval videos. Markers denote chunk means and shaded regions denote bootstrap uncertainty. Prefix is mixed over early chunks but degrades more slowly and is consistently better from chunk~9 onward, supporting a reduction in cumulative late-chunk subject drift rather than a uniform per-chunk improvement.
Overview of our adaptive training. (a) In the G-step, the generator is updated to minimize the static Fréchet objective together with the adaptive Fréchet discrepancy, while both the static and adversarial representations are frozen. (b) In the D-step, the generator is frozen, and the adversarial representation is updated to maximize the Fréchet discrepancy between real and generated distributions. The two steps alternately improve the generator and adapt the representation to discrepancies missed by the fixed encoder.
AdvFD mitigates Fréchet hacking in one-step generation. Top: Compared with JiT-L trained using the static FD loss, AdvFD generates cleaner and more coherent images under the same 1-NFE sampling budget. Bottom: AdvFD consistently reduces FD-r3 and FD-r6 across JiT-L and JiT-H, with relative improvements of 41.4%/38.0% and 34.0%/32.1%, respectively, showing that the gains persist as the generator scales up. These results demonstrate improved perceptual quality and generalization across evaluation representations. Lower is better for all metrics.
Overview of the curriculum training. Our method builds a continuous curriculum transition between independent training and inference-aligned training that balances full distribution coverage with consistency at inference time.
Overview of SparSTAR, a training-free block-sparse attention framework for video autoregressive generation that reduces attention latency while maintaining video quality.
Latent-to-4D training pipeline. A frozen video VAE encodes an observed video into a VAE-space latent. L4AR aligns the latent grid through a learned 3D convolution, reuses frozen camera and time tokens, and refines the representation through alternating frame-wise and global attention. The 4D decoder predicts cameras and dynamic world-space geometry.
Image-to-4D comparison. Each row shows the input condition and 4D results. RGB baselines reconstruct the decoded video, whereas Ours consumes its terminal latent directly.
Overview of the proposed MeanSR framework. Stage 1 estimates an LR-conditioned restoration trajectory by predicting its average velocity under the MeanFlow formulation, while Stage 2 further aligns the generated restoration trajectory with the real HR trajectory through distribution trajectory matching. The proposed Stage-Aware Temporal Sampling (SATS) assigns different temporal distributions to trajectory estimation and trajectory alignment, respectively.
Comparison of representative one-step super-resolution methods. The x-axis denotes CLIPIQA, the y-axis denotes inference runtime, and the bubble size represents FLOPs. MeanSR achieves a favorable trade-off between perceptual quality and computational efficiency.
Overview of ATDEdit for transforming a fox into a dog. (Top) Same-state source/target evaluations produce Token-wise Conditional Surprisal and an editable-token mask. (Bottom) Target-conditioned corrections are applied to editable tokens, while source key/value memory and hard projection constrain selected keep-token rows.
实验效果:在 PIE-Bench 上,ATDEdit 取得该基准上已报告的最强保持性指标,包括 27.44 dB PSNR 与 0.055 LPIPS,同时保持有竞争力的语义对齐。已被 ACM MM 2026 接收。
Comparison of edited results. Real images are in the first column.
批判点评:用「条件意外度」在推理期自动划定可编辑 token,是这篇最漂亮的一笔——它把「哪里该改」从需要人工掩码或额外分割模型的外部输入,转成模型自身条件响应差异的内生信号,免训练、免掩码这两个属性叠加起来带来的实用性相当高。更值得称赞的是学术诚实度:论文明确写出这些操作「促进背景保持但不构成像素级不变性保证」,在一个惯于把机制性偏好包装成理论保证的领域,这种自我设限反而提升了可信度,也准确点出了本方法的性质——它是软偏置而非硬约束。27.44 dB / 0.055 LPIPS 的保持性数字确实扎实。但需要警惕保持性与编辑力度的天然此消彼长:保持性做到基准最强,而语义对齐只被描述为「有竞争力」,这个措辞差异很可能意味着在需要大幅改动的编辑任务上它偏保守;意外度阈值 ρ(t) 是核心超参,其跨图像稳健性与失败模式摘要未交代,若阈值偏保守则编辑区域覆盖不全,偏激进则退回全局漂移;PIE-Bench 单一基准也不足以覆盖真实编辑意图的多样性,尤其是涉及大范围结构变化或多对象组合的情形。
9. CtrlSpeech:音素级控制音高音量与时长
CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis | 德克萨斯大学奥斯汀分校·Amazon | arXiv:2608.08362
Overview of the dataset. Top left: distribution of the expert-annotated examples over subjects across core disciplines. Remaining panels: representative prompt--video examples from each discipline, generated by Gemini-Omni-Flash, where faithful generation requires grounding in the underlying scientific mechanism.
评论 (0)