本文贡献:提出 Evoke,做法是把持久世界状态从模型里搬出去,同时重新设计教师。场景几何维护在一个外部的、按相机位姿索引的世界状态库(World State Bank)中,去噪器每步只检索与当前视角相关的信息,因此会话再长上下文也保持有界。教师不再被当成一个固定的生成器,而是专门为长程监督重新设计:它的稀疏注意力组合了块级分组、对选定远端帧的检索、以及一路线性注意力全局状态,让显存与计算随时长线性增长的同时支持长程监督。这种监督恰好能暴露那种「短窗口内看着合理」的内容漂移,而逐块条件化又让提示词切换和事件控制可以贯穿整个序列。最后用 30 秒的分布匹配目标在自强制 rollout 下训练,把这两项能力一起迁移给一个三步学生,且该学生不需要 classifier-free guidance。
With bounded retention, per-step cost does not grow with session length. The camera pose reads the world state bank into the target view; the student injects that render at the coarsest of its three coarse-to-fine evaluations, alongside a short parametric history and the current text condition. The emitted chunk then updates the store and the history.
An hour-long session stays bounded rather than degrading. One continuous $65.5$,min Evoke rollout, three steps per chunk without classifier-free guidance ($2619$ chunks; faint per-chunk, bold two-minute mean). (a) Color is flat over the hour. (b) Scene identity plateaus at cosine $0.523$, the level real footage scores against itself $60$,s apart. (c) Opening-window shift stays far below unrelated-scene drift. A stability claim, not fidelity ($n=1$).
Comparison of Audio Generation Paradigms. Decoupled Pipeline: Conventional methods synthesize soundscapes and speech independently, requiring post-hoc mixing. VoxAudio: A unified system for vocalized audio synthesis from single prompts.
Detailed architecture and attention mechanism of the Causal Diffusion Transformer. (a) The causal joint block facilitates interaction between audio latents and conditioning features; (b) The causal fused block refines integrated representations through deep self-attention; (c) The chunk-wise causal attention mask, where shaded regions indicate masked-out tokens, ensuring that each temporal chunk only attends to current and historical context.
V-RAE architecture. A frozen visual representation encoder produces temporally dense features, which are compressed by a lightweight temporal pooling module. The figure depicts the chunk-wise causal decoder used with image encoders.
Temporal Reconstruction Error Difference (TRED) for RAEv2, Wan2.2 VAE, OmniTokenizer, and V-RAE on K600 validation examples. The heatmaps visualize temporally averaged TRED within high-frequency regions of interest, while the curves show spatially averaged TRED over consecutive frame transitions.
Several challenges to physics-faithful scientific-diagram generation, and our approach. (a)~State-of-the-art multimodal generators produce diagrams that violate basic physics, such as mislabeled forces or incorrect circuit topology. (b)~The failure is generative rather than perceptual: a model that correctly describes the physics of a scene still generates it incorrectly. (c)~Appearance-based evaluation is unreliable in this setting, a CLIP-style scorer prefers the physically incorrect image (c.1), and an open-vocabulary detector cannot verify the geometric relations that determine correctness (c.2). (d)~Structured Physical Chain-of-Thought () addresses both problems by keeping data construction, annotation, training, and evaluation structured, finitely describable, and extensible.
Overview of : data construction, structured supervision, and training--evaluation. % (a)~Data-curation pipeline. Four complementary sources---a manually curated seed, OpenDataLab~opendatalab, public web archives, and a textbook-procurement stream---are aggregated into a raw pool, then de-duplicated, resolution-filtered, and routed by a subdiscipline classifier into a structurally annotated corpus of ${\sim}4.3$M physics diagrams; a high-quality expert-level subset is split into training and held-out test data, and a balanced subset of the test split forms . % (b)~Corpus composition. Distribution of the corpus across the six physics subdisciplines and their sub-categories, showing a strongly long-tailed profile in which mechanics and quantum mechanics dominate while optics is rare. % (c)~ and how it is used. From a diagram and its prompt, a fixed five-step schema (scene; parameters; force or process analysis; coordinate systems and governing laws; and synthesis) makes the underlying physics explicit. These structured annotations drive three stages: pre-training on the ${\sim}4.2$M corpus-level pairs, supervised fine-tuning on the $109$K human-verified expert-level pairs, and evaluation on , where each generation is scored by answering the binary questions compiled from its own annotation (Local, Global, and Strict accuracy).
Overview of the MotionCraft framework for high-fidelity visual upscaling. The system initiates with Robust Motion Encoding and Fusion, where external motion signals and internal image-driven estimates are gated by a Photometric Inconsistency confidence map to compensate for systematic biases. The core Latent World Transformer (LWT) processes low-resolution latents using Adaptive Sparse Attention, which employs a differentiable soft top-$k$ selection to maintain locality while capturing long-range dependencies. Temporal persistence is managed through a Key-Value Memory Cache ($\mathcal{M}_{t-1}$). The Controllability Interface modulates the latent state via global scalars $\gamma$ and spatial gating maps $g_t$ to balance temporal smoothness and reconstruction fidelity. Finally, a Tiny Conditional Decoder performs distillation-based refinement to produce the high-resolution output $\hat{X}_t$.
Control sensitivity. As $\gamma$ increases, motion smoothness rises monotonically while PSNR decreases gracefully, indicating a predictable trade-off between smoothness and fidelity.
VAE-only optimization. Column 1 shows the target image $X$. Columns 2--7 show optimized images $Z_i$ from two different random initializations, with their decoded latents $\dec(\enc(Z_i))$ shown in the bottom-right inset. Even without a diffusion model, optimizing only through the VAE encoder produces structured noise in $Z_i$, while the decoded latents remain clean.
DreamGaussian comparison. We apply in the second optimization stage of DreamGaussian~tang2024dreamgaussian. Compared with SDS, produces cleaner textures and fewer structured artifacts.
批判点评:这是一篇教科书式的诊断型论文,全程一个人完成。它的方法论比结论更值得学习:先用 VAE-only 优化把「latent 干净、像素脏」这个反直觉现象单独隔离出来复现——这一步很关键,因为它证明了伪影的来源与扩散先验、与渲染器都无关,纯粹是 VAE 编码器的欠约束逆映射造成的;再用简化分析给出数学上的解释;最后修法极轻,不动扩散模型、不动渲染器、不动 SDS 目标,只补一个 VAE 一致的干净方向。「latent 看起来完全正常所以问题很难被发现」这个洞察尤其有意思——它意味着任何只监控隐空间指标的优化流程都会对这类退化完全失明,这个教训可以外推到所有在隐空间做优化的任务,不限于文生 3D。修补的普适性也已用两个不同的 3D 框架(DreamGaussian 与 LucidDreamer)交叉验证。局限同样清楚。单作者、无机构支持,实验规模注定有限:文生 3D 只验证了两个开源框架,对闭源或更大规模的 3D 生成系统是否同样适用未知。效果全部以定性描述给出——「大幅减少」「更干净」,没有任何量化指标(如 CLIP 分数、用户偏好率、伪影强度度量),这在 3D 生成领域虽属常见,但让改进幅度无法与其他去伪影方法横向比较。前瞻步需要额外解码一次,带来的计算开销未报告;超参数 β 虽做了敏感性分析但最优值如何随任务变化没有指导。最后,这个方法治的是 VAE 欠约束这一个特定失效模式,而 SDS 还有过饱和、Janus 多面等其他老问题,它对那些完全无效——读者不应把它理解为 SDS 的通用改进。
8. GST:学整个画家而非几张名作
Through Van Gogh's Eyes: Global Style Transfer with Diffusion Model | KIST·成均馆大学 (ECCV 2026) | arXiv:2608.11546
Overall framework of Global Style Transfer (GST). (Left) Global Style Guidance (GSG) trains a lightweight Style Extraction Function $f_t$ on multiple artworks under the fixed prompt `A painting', learning a residual global style offset $\Delta \mathbf{h}_t$ that modulates the U-Net bottleneck representation $h_t$ in a text-independent manner. During reverse diffusion, $\Delta \mathbf{h}_t$ steers the denoising trajectory toward the artist's global style distribution. (Right) Content Alignment Guidance (CAG) first maps the content image $I_c$ into a noisy latent via DDIM inversion, then applies CLIP-based perceptual guidance at each timestep to align the semantic structure of the generated image with the content while permitting style-driven deformation, yielding artist-level stylization.
Global Style Transfer results across multiple artists for the same content image. Given the same content images $I_c$, our framework transfers the images into the global styles of eight different painters. All images are generated using a text-independent prompt (``A painting”) to eliminate stylistic bias in the diffusion model. For each content image $I_c$ (left column), our framework applies Global Style Guidance (GSG) to capture artist-specific global semantics and Content Alignment Guidance (CAG) to preserve flexible, style-based deformation of content. The results show that a single content image is rendered into distinctly different artistic styles across various artists, demonstrating our framework’s ability to synthesize artist-faithful images.
Motivation and overview of SketchSense. Fixed-sketch conditioning can propagate locally degraded strokes. SketchSense synchronously denoises RGB and structure, regulates sketch use through learned attention modulations and an optional signed prior, and exposes its interpretation as refined structure.
Bidirectional Attention Fusion with adapted QKV reuse. Native QKV provides the shared base, and a dedicated joint-attention QKV LoRA produces fusion-specific $\Delta QKV$. Directional messages pass through a zero-initialized projection, split into $\Delta H_x^l$ and $\Delta H_e^l$, and enter their receiving branches through $g_x^l$ and $g_e^l$.
批判点评:这篇抓的痛点很真实:用户画的草图从来不是均匀可靠的,通常整体意图对、局部笔画烂。而现有两条路线都建立在一个错误前提上——固定条件假设笔画一致可靠,先精修则要求在信息不足时提前把歧义决断掉。SketchSense 的解法在时序安排上更合理:让结构解读与外观生成同步演进、互相喂信息,去噪早期外观尚模糊时不强行定结构,随着 RGB 逐渐清晰再反过来修正结构理解。「精修后的结构暴露了模型对草图的动态理解」这一点还带来了额外的可解释性——用户能看到模型把哪一笔当成了什么,这在交互式创作工具里是实打实的价值,比一个黑箱输出有用得多。可选的带符号先验也务实,它承认「这笔要保留还是要改」有时只有用户知道,那就把这个意图开放给用户显式指定,而不是让模型猜。但这篇是单作者工作,实验规模与工程完备度需要留意。效果段只有「显著优于现有方法」这类定性表述,没有任何量化指标或基线名单,而复原质量与结构保真度都有成熟的客观度量(FID、LPIPS、边缘一致性等),不给数值让改进幅度无从核实。方法本身叠了四个组件——双向注意力融合、短语级目标、草图感知空间调控、可选带符号先验,摘要未提消融,无法判断收益主要来自哪一处,也存在过度设计的可能。绑定 FLUX 块族的 QKV 复用意味着与特定架构耦合,迁移到其他 DiT 骨干需要重新适配。双流同步去噪必然增加计算量,开销未报告。最后,「刻意不合常规的笔画」与「画错的笔画」在模型看来可能无法区分,这个歧义如何处理是这类方法的固有难题。
10. MuseCritic:先写乐评再打分准度飙升
MuseCritic: Learning Multi-Aspect Song Rewards through Natural-Language Aesthetic Critiques | 复旦大学 | arXiv:2608.11755
本文贡献:提出 MuseCritic,一个半标量奖励模型:它先生成一段覆盖五个审美维度的自然语言评论,再把这段评论作为中间表示来预测连续的奖励分数。训练分两阶段:先由教师模型(Gemini-3-Pro)针对歌曲、评分细则与专家评分产出高质量评论用于监督微调;随后让微调后的模型生成自己的评论再做奖励学习,以缓解训练与推理之间的分布偏移。推理时同一骨干挂两个头,LM head 输出五维评论,RM head 给出五维分数。
Overview of the two-stage training and inference pipeline of method. In Stage~I, Gemini-3-Pro generates aesthetic critiques from songs, evaluation rubrics, and expert mean ratings to construct $\mathcal{D}_{\mathrm{off}}$, which is used to fine-tune a pretrained large language model into an SFT critique generator. In Stage~II, the SFT model replaces the teacher critiques with one self-generated five-aspect critique per song to construct $\mathcal{D}_{\mathrm{on}}$; a scalar reward head is then initialized, and is trained using expert mean ratings as regression targets. At inference time, first generates a five-aspect critique and then predicts five continuous aesthetic scores conditioned on the song, rubric, and self-generated critique.
Effect of natural-language aesthetic critiques on five-dimensional score distributions. Curves show expert ratings, full predictions, and predictions from the no-critique variant on the fixed SongEval test set. Coherence, Musicality, Memorability, Clarity, and Naturalness correspond to the five aesthetic dimensions. Without critiques, predictions cluster in a narrow high-score region; the full model more closely matches the range of expert ratings.
评论 (0)