Text-to-image diffusion transformers follow predictable scaling laws and require ten times as much data as a comparably sized language model to reach compute optimality.
Iso-FLOP suboptimality penalty $(L_N - L^*)/L^*$ versus TPP for the family (left) and the Chinchilla / Hoffmann et al. data (right), the latter figure-digitized by~besiroglu2024chinchilla and binned into log-$N$ model-size bands. Markers are the measured points ( evaluated checkpoints; digitized Chinchilla points); curves interpolate through them. The green/red bands mark the $0.2\%$ penalty threshold.
批判点评:这是今天这批里信息密度最高的一篇,而且是那种一旦被验证就会改变整个领域默认做法的工作。「每参数 200 token、是 LLM 的十倍」这个数字本身就足够有价值——它意味着相当多现有的文生图模型可能都处于数据不足的状态,一直在把预算错配到参数上。IsoFLOP 的两个指数分别是 0.4951 和 0.5049,加起来几乎正好是 1,这种干净程度说明拟合质量确实高,不是硬凑出来的。「对过训练鲁棒」这条实践建议也很务实:它把一个需要精确调参的问题变成了单边保守选择,工程上极其友好。不过要清醒几点。第一,最大只训到 2B 参数,而当前前沿文生图模型早已在 10B 量级以上,外推距离仍然不短——论文自己的图里最优点已经外推到 10^23 FLOPs 处的红点,这已是纯预测而非实测。第二,「200 token per parameter」这个数字必然依赖 tokenizer/patch 配置与图像分辨率,换一个 VAE 压缩率或 patch size,token 的信息量就变了,200 这个绝对值能否跨配置迁移,是使用者最需要知道却没有明确交代的。第三,全程使用 flow-matching transformer 这一种架构,结论对 UNet 或其他扩散形式是否成立未验证。第四,Chinchilla 的对比曲线是从原论文图里数字化重建的(论文自己在图注里标注了 curves are noisier than real per-run data),这个对照的严谨性弱于自己跑的实验,而「十倍」这个最抢眼的结论恰恰建立在这个对照上。第五,Luma AI 作为商业公司发布这份研究,代码与检查点是否开放决定了社区能否独立复现这条曲线。
2. EditBridge:4K图像编辑61秒完成
EditBridge: Towards Faithful and Efficient Ultra-High-Resolution Image Editing | 上海交通大学·阿里 Qwen·新加坡国立大学 | arXiv:2608.18063
关键词:超高分辨率编辑,扩散桥,分块稀疏注意力,4K保真,上交·阿里Qwen
⚠️ 前序问题:专业修图工作流里 4K 起步是常态,但扩散编辑模型基本卡在 1K 以下——注意力的二次复杂度加上显存墙,让高分辨率直接变成不可能。业界的通行绕法是两段式:先在低分辨率上编辑,再独立跑一次超分。这个做法看似合理,实际上埋了两个坑。一是信息发散:超分模型是在「猜」细节,它猜出来的东西与原始高分辨率图里真实存在的细节经常对不上——原图眼睛是棕色,低分编辑后超分给你补出一只绿色的眼睛,且补得很自信。二是纹理退化:超分要么把纹理磨得过平,要么锐化过头产生假质感。根子在于超分环节完全看不到原始 HR 图,它只能从 LR 结果凭空重建。
本文贡献:EditBridge 的做法是重新定义这一步在做什么。传统扩散是从噪声重新生成,而这里把细化过程形式化为一次结构化的数据到数据翻译:起点是 LR 编辑结果,终点是它的 HR 对应版本,整个过程显式地以原始 HR 源图为条件。这样一来,真实细节是「继承」来的而不是「猜」来的,信息发散问题从形式上就被堵住了。但引入 HR 源图作为条件会带来新的算力问题——两张超高分辨率图之间做全交叉注意力是不现实的。为此他们设计了先验引导的分块稀疏注意力:利用第一阶段编辑过程中已经产生的语义对应关系(用注意力图取 argmax 得到位置映射 π),把跨图交互限制在空间上真正对齐的区域内。也就是说,目标图的某个 chunk 只需要去看源图中与它对应的那几个 chunk,而不是全图。注意力被拆成两条路径:Path A 是域内自注意力,Path B 是先验引导的跨 chunk 稀疏注意力。
Overview of our proposed EditBridge. The upper section illustrates the inference process of the diffusion bridge, which transports the low-resolution edit to its high-resolution counterpart. The lower section details the construction of the correspondence prior and the proposed Prior-Guided Sparse Attention (PG-BSA) mechanism.
GRNEdit: Efficient General Video Editing from a New Binary-Evidence Perspective in Generative Refinement Networks | 西安电子科技大学·小米 MiLM Plus | arXiv:2608.16328
本文贡献:GRNEdit 从 GRN(生成式细化网络)用 bit 组合编码视觉语义这一特性出发,把编辑语义重新表述成对每个 bit 的「保留还是翻转」的局部二元决策。这个视角转换是全文的支点:源信息不再是需要被重新编码的内容,而是支持当前二元状态的「坐标级证据」,而 GRN 主干继续负责把这些局部决策组合成全局连贯的生成语义。第一阶段用一个紧凑编码器把离散源码翻译为连续证据信号,在二元细化过程中持续注入。他们还借鉴 classifier-free guidance 的空提示训练,赋予空条件一个编辑专属含义——空指令即「不编辑」,由源重建来监督。这条恒等通路一石二鸟:既隐式强化了证据利用与内容保持,又在同一表示空间里产生了一个「源保持态」。于是第二阶段可以把每个编辑态与它对应的源保持态直接相减,用差异去修正那些第一阶段没定下来的目标 bit——这是一个相当巧妙的自我对照机制。
Overview of GRNEdit. Given a source video and an editing instruction, Stage I injects prompt-modulated source evidence into the GRN chunks through lightweight per-chunk projectors, each containing only about 3 M parameters for GRNEdit-2B and 7 M for GRNEdit-8B. Stage II refines the binary predictions by treating source evidence as support in keep regions and as counter-evidence in edit regions.
Qualitative comparison across diverse editing tasks. GRNEdit performs more accurate edits while better preserving unedited content, whereas competing methods often miss the intended change or disrupt scene consistency.
Capability-specific data construction in our framework. Shared collection and wrangling feed three specialized but interoperable engines for visual expression, editing association, and knowledge-grounded reasoning.
Capability-gap-driven active feedback loop. Capability-aware evaluation identifies failure cases, which seed neighboring-data retrieval and expert-driven construction. Gap-aware resampling then increases the weights of persistent failures and down-weights resolved gaps in subsequent training.
Illustration of MoE-ViE architecture and training pipeline. MoE-ViE adopts a fine-grained MoE design with optimized MoE kernels for latency reduction. We first perform contrastive pretraining on large-scale image-text pairs. During video finetuning, we introduce frame-level distillation and expert freezing to preserve the pretrained image understanding capabilities.
Overview of WildMoir'e dataset construction. We first collect real screen-captured images containing complex moir'e and localize the display regions. Five image-conditioned generative models then produce candidate GT images. The candidates are processed at matched spatial locations through spatial and color alignment, synchronized patch extraction, pre-filtering, and optimal GT selection. This offline procedure yields 6.8K moir'e-GT pairs at 1024()1024 resolution. The plots on the right summarize the improvements obtained by training ESDNet, SDXL, and Qwen-Image-Edit with constructed WildMoir'e dataset.
Comparison of qualitative results. The first example emphasizes logo and text fidelity under broad color interference, while the second contains severe mixed chromatic patterns over fine scene textures. Nano-Banana-2 is relatively content-faithful but may retain moir'e, whereas GPT-Image-2 removes artifacts aggressively while replacing scene content and producing an overly digital-image appearance. Using WildMoir'e suppresses residual bands more effectively across all 3 trainable models.
Overview of the benchmark construction pipeline. Candidate images are first collected and filtered from multiple sources, then annotated through structured tagging. The tagged pool is then balanced through source-wise sampling and augmented with synthetic data. It is subsequently used for formula-template sampling and prompt construction.
Overview of the ourmodel evaluation pipeline. Given a case with its prompt and evaluation setting, the agentic planner routes the case to applicable skills from the skill library, recording an evidence-grounded reason for every activated and skipped skill (e.g., the Offscreen Evolution Verifier is skipped because all requested actions remain visible). Each activated skill spawns sub-agents that inspect the world model rollout for specific sub-questions, such as target visibility and final state validity; here, the Intentional Change Verifier detects an unrelated human intervention picking up the cube and zeroes the no-extra-event score. The collected evidence is finally aggregated into per-dimension scores and an interpretable final score.
Human alignment of ourmodel and its comparison with WBench. (a) Model-level human alignment: each point is one of the evaluated models; we fit a linear curve for both Intentional and Physical Transition. (b) Controlled comparison with the closest WBench protocols on the same data: we report pairwise accuracy (higher is better), draw rate, and Brier score (both lower is better).
评论 (0)