每日 arXiv 论文简报
今日 arXiv 共收录 66 篇论文,扩散模型(Diffusion)依然占据绝对主导地位(52篇),自回归模型(Autoregressive)专注于多模态生成与具身智能(13篇),图像压缩仅 1 篇。
- 整体趋势:扩散模型正向高效推理、跨模态控制与物理约束方向深入发展,出现了量子扩散、扩散遗忘、噪声序列交叉等新范式;自回归模型则聚焦于统一视觉-语言-动作(VLA)智能体与三维场景生成的端到端框架。
- 交叉联系:语音处理(UniSAE)同时出现在自回归与扩散两赛道;自动驾驶场景成为两方向的共同焦点,涵盖感知-决策-生成全链路。
- 关键亮点:三维网格生成与Mesh生成取得突破性进展(Mesh BDF、PolyFlow);视频生成与Avatar实时化成为生成模型新战场;材料设计与分子扩散模型将AI4Science推向新深度。
重点推荐论文:
- Quantum Flow Matching — 将扩散模型与量子计算结合,开辟量子机器学习新范式
- InteractiveAvatar — 实现实时流式视频生成,支持一致性与意图感知的Avatar交互
- ForgeDrive — 双向交叉条件化统一视觉-动作生成,推动自动驾驶端到端可解释性
- OTCache — 引入最优传输实现几何感知的扩散模型缓存,兼顾生成质量与效率
- AR-CoPO — 对比策略优化对齐自回归视频生成,强化视觉常识推理能力
Autoregressive 类别论文概述(2025年6月)
今日 Autoregressive 相关论文呈现两个主要方向:一是图像/视频生成,包括布局条件图像生成和视频生成的策略优化;二是解码加速,如推理感知的推测解码用于自动驾驶场景。整体趋势显示自回归模型正在向更高效、更可控的方向发展,尤其在视觉-语言-动作(VLA)领域有显著应用。
- Layout-Conditioned Autoregressive Text-to-Image Generation via Structured Masking (2509.12046):提出结构化掩码方法,实现布局感知的自回归图像生成,提升生成可控性。
- AR-CoPO: Align Autoregressive Video Generation with Contrastive Policy Optimization (2603.17461):将对比策略优化引入自回归视频生成,提升动作一致性和质量。
- Reasoning-aware Speculative Decoding for Efficient Vision-Language-Action Models in Autonomous Driving (2606.31160):针对自动驾驶中 VLA 模型的推理感知推测解码,显著提升推理效率。
- UniSAE: Unified Speech Attribute Editing on Speaker, Emotion and Low-Level Content via Discrete Phonetic Posteriorgram Modelling (2606.31128):利用离散音素后验图实现多属性语音编辑,具有自回归生成特性。
Modeling Cell-Cycle-Aware Single-Cell Drug Perturbation Responses
中文标题:构建细胞周期感知的单细胞药物扰动响应模型
作者:Dingping Zhao, Jie Lin
Single-cell drug perturbation models should predict not only transcriptional response magnitude, but also whether a treatment alters the proliferative state of a cell. This is challenging because cell-cycle variation is often treated as nuisance variation, and benchmark pipelines rarely treat drug-induced phase changes as a primary prediction target. We introduce scCycleMol, a cell-cycle-aware perturbation prediction framework built on a curated 24-hour SciPlex3 benchmark with standardized molecule identities, dose and cell-line metadata, and gene expression with cell-cycle supervision derived from treated states. Instead of using cell-cycle state as an input covariate, scCycleMol derives supervision from predicted treated expression and propagates it through a learnable full-expression cell-cycle head with circular G1/S/G2M phase targets. We evaluate marker-based supervision, molecular representations, and pretraining strategies to isolate sources of improvement. Across a SciPlex3 benchmark with over 600k cells, 186 perturbation conditions, multiple cancer cell lines, and thousands of genes, scCycleMol improves out-of-distribution expression prediction compared with conditional perturbation baselines. The best LINCS-pretrained circular model achieves 0.9093 expected all-gene r squared and 0.6843 expected differentially expressed gene r squared, compared with 0.6800 and 0.5400 for LINCS-pretrained ChemCPA. Closed-loop cell-cycle supervision improves phase accuracy by about 0.5 to 0.6 points while maintaining nearly unchanged expression prediction. A Tahoe-pretrained variant reaches 0.9609 phase accuracy, highlighting the benefit of explicit cell-cycle-aware supervision in perturbation modeling.
单细胞药物扰动模型不仅应预测转录响应强度,还应预测治疗是否改变细胞的增殖状态。这具有挑战性,因为细胞周期变异通常被视为干扰变异,基准流程很少将药物诱导的时相变化作为主要预测目标。我们提出了scCycleMol,一个细胞周期感知的扰动预测框架,建立在经过整理的24小时SciPlex3基准数据集之上,该数据集具有标准化的分子身份、剂量和细胞系元数据,以及来源于处理状态的细胞周期监督基因表达。scCycleMol不使用细胞周期状态作为输入协变量,而是从预测的处理后表达中导出监督,并通过具有环状G1/S/G2M时相目标的可学习全表达细胞周期头将其传播。我们评估了基于标记物的监督、分子表示和预训练策略,以分离改进来源。在包含超过60万个细胞、186个扰动条件、多个癌细胞系和数千个基因的SciPlex3基准数据集上,scCycleMol相比条件扰动基线改善了分布外表达预测。最佳的LINCS预训练环状模型实现了0.9093的预期全基因R方和0.6843的预期差异表达基因R方,而LINCS预训练ChemCPA分别为0.6800和0.5400。闭环细胞周期监督将时相准确率提高约0.5至0.6个百分点,同时保持表达预测基本不变。Tahoe预训练变体达到0.9609的时相准确率,突出了显式细胞周期感知监督在扰动建模中的益处。
UniSAE: Unified Speech Attribute Editing on Speaker, Emotion and Low-Level Content via Discrete Phonetic Posteriorgram Modelling
中文标题:UniSAE:通过离散语音后验图建模实现说话人、情感和低级内容的统一语音属性编辑
作者:Chuanbo Zhu, Wuyou Zhou, Rongxiu Zhong, Shilei Zhang, Kun Qian, Yike Guo, Wei Xue
Speech editing aims to modify specific portions of an utterance while preserving the remaining speech. Existing approaches primarily focus on word-level content modification and typically treat content, speaker, and emotion editing as separate tasks, limiting both editing granularity and flexibility. We propose UniSAE, a unified speech attribute editing framework which supports composable speaker, emotion and content editing from sub-phoneme to word level within a single architecture. UniSAE introduces a Discrete Phonetic PosteriorGram (DPPG) representation that factorizes speech content into discrete tokens encoding phoneme identity, pronunciation variants, and duration, enabling direct phoneme- and sub-phoneme-level editing. For higher-level modifications, an autoregressive content transformer predicts edited DPPG sequences for word-level content editing. The edited sequences are rendered into speech by a diffusion-based acoustic decoder, conditioned on disentangled speaker and emotion representations. Experimental results demonstrate that the proposed unified framework supports precise speaker and emotion control, content editing at multiple granularities, and joint modification of all three attributes within a single framework.
语音编辑旨在修改语音中特定部分的同时保留其余内容。现有方法主要关注词级内容修改,并且通常将内容、说话人和情感编辑视为独立任务,这限制了编辑的粒度和灵活性。我们提出了UniSAE,一个统一的语音属性编辑框架,在一个架构中支持从亚音素级到词级的可组合说话人、情感和内容编辑。UniSAE引入了一种离散语音后验图(DPPG)表示,将语音内容分解为编码音素身份、发音变体和时长的离散标记,从而实现直接的音素级和亚音素级编辑。对于更高级别的修改,自回归内容变换器预测编辑后的DPPG序列,用于词级内容编辑。编辑后的序列通过基于扩散的声学解码器进行渲染,并以解耦的说话人和情感表示为条件。实验结果表明,所提出的统一框架支持精确的说话人和情感控制、多粒度的内容编辑,以及在单一框架内对所有三个属性的联合修改。
MIRTH: Mutual-Information Reasoning with Temporal Hubs for Vision-Language-Action Agents
中文标题:MIRTH:基于时间枢纽的视觉-语言-动作智能体互信息推理方法
作者:Hao Sun, Yu Song, Shiyu Teng, Ziwei Niu, Yen-Wei Chen
VLA models have emerged as a powerful paradigm for transferring semantic knowledge from web-scale data to physical robotic control. However, current single-frame architectures suffer from intrinsic limitations: temporal myopia that discards historical dynamics, reasoning gaps between high-level instructions and low-level motor commands, and inference inefficiency due to autoregressive scalar decoding. In this work, we propose MIRTH, a unified framework designed to address these challenges. MIRTH augments a pretrained VLA backbone with three key innovations: (1) dual-scale temporal memory hubs that compress long-term scene evolution and short-term motion trends into compact embeddings; (2) latent reasoning tokens optimized via a mutual-information objective carving out a semantic plan space to align multimodal context with action trajectories; and (3) a parallel action decoding scheme that replaces autoregressive generation with vector-wise prediction to maximize control throughput. Extensive evaluations on the LIBERO simulation benchmark and a real-world LeRobot platform demonstrate that MIRTH achieves state-of-the-art performance and exhibiting emergent error recovery capabilities. The codes and collected datasets are released at http://github.com/kiva12138/mirth.
视觉-语言-动作(VLA)模型已成为一种强大的范式,能够将来自网络规模数据的语义知识迁移到物理机器人控制中。然而,当前的单帧架构存在固有的局限性:丢弃历史动态的时序短视问题、高级指令与低级电机命令之间的推理差距,以及由于自回归标量解码导致的推理效率低下问题。本工作提出MIRTH,一个统一框架,旨在应对这些挑战。MIRTH通过三项关键创新增强预训练的VLA主干网络:(1)双尺度时间记忆枢纽,用于将长期场景演化和短期运动趋势压缩为紧凑的嵌入表示;(2)通过互信息目标优化的潜在推理标记,挖掘语义规划空间以对齐多模态上下文与动作轨迹;(3)并行动作解码方案,用向量级预测替代自回归生成,以最大化控制吞吐量。在LIBERO仿真基准测试和真实世界的LeRobot平台上的广泛评估表明,MIRTH达到了最先进的性能,并展现出涌现的错误恢复能力。代码和采集的数据集已发布于http://github.com/kiva12138/mirth。
Bridging Local Observation and Global Simulation in Closed-Loop Traffic Modeling
中文标题:桥接闭环交通建模中的局部观测与全局仿真
作者:Ziyan Wang, Tan Xiang, Peng Chen, Xintao Yan
A local-to-global context mismatch arises when autoregressive traffic simulators trained on ego-centric driving logs are deployed in globally observable closed-loop environments. In such logs, the ego vehicle has rich local observations, while surrounding agents are only partially observed due to perception limits and occlusions. As a result, simulators may learn incomplete context--action mappings that remain hidden in log-based training but emerge during closed-loop rollouts, leading to unrealistic behaviors such as abnormal stops, unsafe interactions, and rule violations. We propose CRAFT, a Contextual pReference Alignment Framework for Traffic Simulation, to mitigate this mismatch via self-supervised failure discovery and preference-guided test-time alignment. CRAFT treats the base simulator as a globally observable sandbox, generating diverse what-if rollouts from logged initial states to expose context-induced failures. These failures are grounded with human-aligned driving priors and converted into preference supervision for training a Contextual Preference Evaluator (CPE). At inference time, CPE acts as a plug-in alignment module that scores candidate actions under complete scene context and reweights autoregressive decoding toward globally coherent behaviors. CRAFT mitigates this local-to-global contextual bias, reducing collisions by 31.2\% and traffic violations by 33.2\% without retraining the base simulator.
当在全局可观测的闭环环境中部署基于自我中心驾驶日志训练的自回归交通仿真器时,会出现局部到全局的上下文不匹配问题。在这些日志中,自车具有丰富的局部观测,而由于感知限制和遮挡,周围交通参与者仅能被部分观测。因此,仿真器可能学习到不完整的上下文-动作映射,这种映射在基于日志的训练中隐藏,但在闭环推演过程中显现,导致异常停止、不安全交互和违规行为等不现实的行为。我们提出CRAFT(交通仿真的上下文偏好对齐框架),通过自监督失败发现和偏好引导的测试时对齐来缓解这种不匹配。CRAFT将基础仿真器视为全局可观测的沙盒,从日志记录的初始状态生成多样化的反事实推演,以暴露上下文诱导的失败。这些失败基于人类对齐的驾驶先验进行验证,并转化为偏好监督信号,用于训练上下文偏好评估器(CPE)。在推理时,CPE作为一个即插即用的对齐模块,在完整场景上下文下对候选动作进行评分,并重新加权自回归解码以产生全局一致的行为。CRAFT缓解了这种局部到全局的上下文偏差,在不重新训练基础仿真器的情况下,将碰撞减少31.2%,交通违规减少33.2%。
Layout-Conditioned Autoregressive Text-to-Image Generation via Structured Masking
中文标题:基于结构化掩码的布局条件自回归文本到图像生成
作者:Zirui Zheng, Takashi Isobe, Tong Shen, Xu Jia, Jianbin Zhao, Xiaomin Li, Mengmeng Ge, Baolu Li, Qinghe Wang, Dong Li, Dong Zhou, Yunzhi Zhuge, Huchuan Lu, Emad Barsoum
Although autoregressive (AR) models have demonstrated remarkable success in image generation, extending these models to layout-conditioned generation remains challenging due to the sparse nature of layout conditions and the risk of feature entanglement. We present \textbf{S}tructured \textbf{M}asking for \textbf{AR}-based \textbf{L}ayout-to-\textbf{I}mage (SMARLI), a novel framework that effectively integrates spatial layout constraints into the AR generation process. To equip AR models with layout control, a structured masking strategy is applied to the attention computation to govern the interaction among the global prompt, layout, and image tokens. This design prevents the misassociation of different regions with their corresponding descriptions while enabling the sufficient injection of layout constraints into the generation process. To alleviate the exposure bias of AR models and further enhance generation quality and layout accuracy, we incorporate a Group Relative Policy Optimization (GRPO) post-training scheme. We adapt it to the next-set-based paradigm and introduce a specifically designed layout reward, which is coordinated with an image quality reward to guide policy optimization in a balanced manner. Experimental results demonstrate that SMARLI seamlessly integrates layout tokens with text and image tokens without compromising generation quality, and the proposed masking strategy and post-training scheme can also be transferred to standard next-token-based AR models. The proposed framework achieves superior layout control while maintaining the structural simplicity and generation efficiency of AR models.
尽管自回归(AR)模型在图像生成方面已展现出显著的成功,但将这些模型扩展到布局条件生成仍然具有挑战性,这主要源于布局条件的稀疏性以及特征纠缠的风险。本研究提出SMARLI(基于自回归的布局到图像的结构化掩码),这是一个将空间布局约束有效整合到自回归生成过程中的新框架。为了赋予自回归模型布局控制能力,我们采用结构化掩码策略来调节注意力计算,以管理全局提示、布局和图像token之间的交互。该设计能够在防止不同区域与其对应描述错误关联的同时,将布局约束充分注入生成过程。为了缓解自回归模型的暴露偏差并进一步提升生成质量和布局准确性,我们引入了一种分组相对策略优化(GRPO)后训练方案。我们将其适配到基于下一集合的范式,并引入专门设计的布局奖励,该奖励与图像质量奖励相协调,以平衡的方式引导策略优化。实验结果表明,SMARLI能够将布局token与文本和图像token无缝集成,且不损害生成质量,所提出的掩码策略和后训练方案也可迁移到标准的基于下一token的自回归模型。所提框架在保持自回归模型结构简洁性和生成效率的同时,实现了卓越的布局控制能力。
InfiniVerse: Occupancy Guided Unbounded Scene Generation for Autonomous Driving
中文标题:InfiniVerse:面向自动驾驶的占用引导无界场景生成
作者:Xiaoyu Ye, Leheng Li, Xinyu Ji, Yingjie Cai, Hongda He, Xu Yan, Guanyi Zhao, Ying-Cong Chen, Bingbing Liu, Shuguang Cui, Zhen Li
Generating realistic, controllable, and temporally coherent urban environments is a critical yet unresolved challenge in the autonomous driving community. In this paper, we introduce InfiniVerse, a unified pipeline for long-range, 2D-3D-aligned, and controllable synthesis of dynamic urban scenes from a single frame. In practice, our approach first reconstructs a 3D occupancy representation from the input multi-view frame. This representation serves as a foundation for autoregressive scene extension along arbitrary trajectories. Subsequently, a video diffusion model translates the coarse occupancy grid into realistic, spatiotemporally consistent video sequences. Moreover, we propose a hierarchical sketch-and-refine paradigm, in which the generated videos are re-projected as image-conditioned feedback to enhance the 3D occupancy representation, establishing cross-modal alignment and mutual enhancement between the visual and spatial domains. Extensive evaluations on the Waymo Open Dataset and nuScenes demonstrate that InfiniVerse achieves state-of-the-art performance, with a FID of 6.4 and FVD of 67.97, significantly outperforming existing benchmarks in both duration and stability.
生成逼真、可控且时间一致的城市环境是自动驾驶领域一项关键但尚未解决的挑战。本文提出了InfiniVerse,一个从单帧图像进行长距离、2D-3D对齐且可控的动态城市场景合成的统一流程。具体而言,我们的方法首先从输入的多视角图像帧重建3D占用表示,以此为基础沿任意轨迹进行自回归场景扩展。随后,视频扩散模型将粗糙的占用网格转化为逼真且时空一致的视频序列。此外,我们提出了一种分层式草图-细化范式,将生成的视频重新投影为图像条件反馈,以增强3D占用表示,从而在视觉域与空间域之间建立跨模态对齐与相互增强。在Waymo Open Dataset和nuScenes数据集上的广泛评估表明,InfiniVerse达到了最先进的性能,FID得分为6.4,FVD得分为67.97,在持续时间和稳定性方面均显著优于现有基准方法。
Reasoning-aware Speculative Decoding for Efficient Vision-Language-Action Models in Autonomous Driving
中文标题:自动驾驶中视觉-语言-动作模型的高效推理感知推测解码
作者:Anh Dung Dinh, Simon Khan, Flora Salim
Modern Vision-Language-Action (VLA) planners for autonomous driving emit a chain-of-causation (CoC) reasoning step \emph{before} producing a trajectory. The reasoning is autoregressive and dominates inference latency, while the trajectory head is parallel and cheap. Latency is an operational constraint in autonomous driving, so accelerating the reasoning step is the central problem we address. We observe that CoC reasoning has two qualitatively different needs: most tokens continue routine setup that follows naturally from the ego-trajectory history, and a small fraction encode commitments that require fresh visual evidence about an unexpected situation. We split this reasoning into two specialized paths: a \emph{routine reasoner} that handles the predictable continuation by attending to trajectory history, and a \emph{deliberative reasoner} (the unmodified VLA target) that handles novel cases by attending to current visual evidence, using the speculative decoding framework as the architectural template for how the two paths cooperate. Unlike standard speculative decoding, our routine reasoner is not a smaller replica of the target; the two reasoners are deliberately specialized to read different parts of the prompt. We propose two techniques to realize this. First, we introduce \textbf{FlatRoPE}, a 1D rotary positional embedding in the draft that breaks the rotational symmetry of the target's 3D M-RoPE, redirecting attention away from visual tokens and onto trajectory-history tokens. Second, we introduce \textbf{Action-aware RL (AARL)}, a post-training stage that uses an action-quality reward together with a static-reference KL anchor. Together, our two-reasoner system reduces the reasoning-step running time by approximately $4\times$ relative to the original Alpamayo planner.
现代自动驾驶的视觉-语言-动作(VLA)规划器在生成轨迹之前会输出一个因果链(CoC)推理步骤。该推理是自回归的且主导推理延迟,而轨迹头部是并行的且计算成本较低。延迟是自动驾驶的操作约束条件,因此加速推理步骤是我们解决的核心问题。我们观察到CoC推理有两种本质上不同的需求:大多数标记延续着从自我轨迹历史自然跟随的常规设置,而一小部分标记编码了需要对意外情况进行新鲜视觉证据的承诺。我们将这种推理分成两条专门路径:一条是处理可预测延续的常规推理器,通过关注轨迹历史来实现;另一条是处理新颖情况的审慎推理器(未经修改的VLA目标),通过关注当前视觉证据来实现,采用推测解码框架作为两条路径协作的架构模板。与标准推测解码不同,我们的常规推理器不是目标的较小副本;两个推理器被特意设计为读取提示的不同部分。我们提出两项技术来实现这一点。首先,我们引入FlatRoPE,这是一种草稿中的1D旋转位置编码,打破了目标的3D M-RoPE的旋转对称性,将注意力从视觉标记重定向到轨迹历史标记。其次,我们引入动作感知强化学习(AARL),这是一个使用动作质量奖励结合静态参考KL锚的后训练阶段。我们的双推理器系统将推理步骤的运行时间相对于原始Alpamayo规划器减少了约4倍。
ForgeDrive: Bidirectional Cross-Conditioning for Unified Visual-Action Generation in Autonomous Driving
中文标题:ForgeDrive:用于自动驾驶中统一视觉-动作生成的双向交叉条件方法
作者:Xuchang Zhong, He Zheng, Chenxu Zhao, Tianxiong Lv, Hangqi Fan, Bohua Wang, Yushan Liu, Zhihao Liao, Leigang Luo, Congyang Zhao, Yang Cai
World-model-based autonomous driving endows the model with the ability to understand scene evolution. Yet this promise is undermined by the prevailing imagine-then-act paradigm, which allows errors from the more challenging visual generation stage to cascade into action planning. We introduce ForgeDrive, a unified autoregressive diffusion framework with visual-action cross-conditioning that closes this gap through act-then-imagine paradigm. ForgeDrive factorizes the future as a sequence of per-timestep frame-action pairs, intertwining each action with its corresponding visual observation. During training, we decouple the diffusion timesteps of the two modalities and introduce a UniDiffuser-style noise scheduler to get the ability to infer either modality from its counterpart and deepen understanding of relationships between images and actions. At inference, we propose a novel act-then-imagine inference paradigm, and find that at each step, action generation is a capability internalized during training, requiring no clean future frame as a prerequisite at inference time; instead, the generated action can improve the accuracy of future frame generation, which in turn enhances the quality of the next action. Additionally, we augment each step with future ego-status prediction, further sharpening planning ability. Extensive experiments on NAVSIM demonstrate that ForgeDrive not only unifies driving simulation, planning, and visual odometry into a single model, but also outperforms existing strong planners without any post-training strategy.
基于世界模型的自动驾驶赋予模型理解场景演变的能力。然而,这一愿景受到盛行的"先想象后行动"范式的阻碍,该范式允许更具挑战性的视觉生成阶段的错误级联传播到动作规划中。我们提出了 ForgeDrive,这是一个统一的自回归扩散框架,采用视觉-动作交叉条件,通过"先行动后想象"范式来弥补这一差距。ForgeDrive 将未来分解为一系列每时间步的帧-动作对,将每个动作与其对应的视觉观察交织在一起。在训练过程中,我们解耦了两种模态的扩散时间步,并引入了 UniDiffuser 风格的噪声调度器,以获得从一种模态推断另一种模态的能力,并加深对图像与动作之间关系的理解。在推理过程中,我们提出了一种新颖的"先行动后想象"推理范式,并发现每个步骤中,动作生成是训练过程中内化的能力,推理时不需要清晰未来帧作为先决条件;相反,生成的动作可以提高未来帧生成的准确性,从而增强下一个动作的质量。此外,我们在每个步骤中增加了未来自车状态预测,进一步提升了规划能力。在 NAVSIM 上的广泛实验表明,ForgeDrive 不仅将驾驶仿真、规划和视觉里程计统一到单个模型中,而且无需任何后训练策略就优于现有的强大规划器。
Bridging Video Understanding and Generation in a Unified Framework
中文标题:在统一框架中桥接视频理解与生成
作者:Yuqi Wang, Runyi Li, Ruoyu Feng, Renjie Chen, Wenfeng Lin, Mingyu Guo
Recently, unified image generation and understanding have been extensively explored. However, extending such unified modeling paradigms to the video domain remains largely underexplored. A central challenge is that video understanding favors compact, discriminative semantic representations, whereas video generation requires dense signals that preserve visual details and temporal coherence. Videos naturally capture both spatial semantics and temporal dynamics, making them a more suitable modality for unified multimodal modeling compared to static images. In this paper, we propose Vega, a unified framework that bridges video understanding and generation. Vega leverages a shared vocabulary to jointly model text and visual representations and employs a hybrid architecture combining autoregressive (AR) prediction with diffusion-based rendering. Specifically, the AR model focuses on predicting semantically meaningful visual tokens for keyframes, providing a structured representation that guides the diffusion module in rendering dense, high-resolution video frames. Extensive experiments demonstrate that Vega achieves strong performance on video generation benchmarks such as VBench and video understanding benchmarks like VideoMME.
最近,统一图像生成与理解已得到广泛探索。然而,将这种统一建模范式扩展到视频领域仍基本未被深入研究。一个核心挑战在于,视频理解倾向于紧凑的、具有判别性的语义表示,而视频生成则需要保留视觉细节和时间一致性的密集信号。与静态图像相比,视频自然地同时捕捉空间语义和时间动态,使其成为统一多模态建模的更合适模态。本文提出Vega,一个桥接视频理解与生成的统一框架。Vega利用共享词汇表联合建模文本和视觉表示,并采用结合自回归(AR)预测与基于扩散的渲染的混合架构。具体而言,AR模型专注于预测关键帧的语义有意义的视觉token,提供一种结构化表示,引导扩散模块渲染密集的高分辨率视频帧。大量实验表明,Vega在视频生成基准(如VBench)和视频理解基准(如VideoMME)上均取得了强劲性能。
Mesh BDF: Barycentric Dominance Field for 3D Native Mesh Generation
中文标题:Mesh BDF:用于3D原生网格生成的重心支配场
作者:Gaochao Song, Haohan Weng, Luo Zhang, Zibo Zhao, Shenghua Gao
Autoregressive (AR) modeling has recently achieved remarkable progress in native 3D mesh generation, largely due to its natural ability to handle variable-length, discrete data structures. However, the inherent constraints of the AR paradigm severely restrict the generated meshes, leading to limited face counts, bounded vertex resolutions, and difficulties in supporting textures. To overcome these bottlenecks, we propose the Barycentric Dominance Field (BDF), a continuous representation defined on triangular mesh surfaces that elegantly encodes vertex topological connectivity. BDF bridges the fundamental gap between discrete mesh topology and continuous diffusion-based generative modeling by transforming connectivity into a continuous surface signal. As an intrinsic mesh property, BDF shares strong similarities with texture maps, enabling its seamless integration into existing 3D diffusion pipelines without requiring architectural modifications. Extensive experiments demonstrate that BDF empowers diffusion models to generate native meshes with significantly higher quality, greater scalability, and stronger robustness compared to state-of-the-art autoregressive methods.
自回归(AR)建模在原生3D网格生成领域近年来取得了显著进展,这主要归功于其处理变长离散数据结构的天然能力。然而,自回归范式的固有约束严重限制了生成网格的质量,导致面数有限、顶点分辨率受限,以及纹理支持的困难。为克服这些瓶颈,我们提出了重心支配场(BDF),这是一种定义在三角网格表面上的连续表示方法,能够优雅地编码顶点拓扑连接性。BDF通过将连接性转换为连续表面信号,架起了离散网格拓扑与基于连续扩散的生成式建模之间的根本桥梁。作为一种内在网格属性,BDF与纹理映射具有很强的相似性,使其能够无缝集成到现有的3D扩散管道中,而无需对架构进行修改。大量实验表明,与最先进的自回归方法相比,BDF使扩散模型能够生成质量显著更高、可扩展性更强、鲁棒性更强的原生网格。
PolyFlow: Continuous Topology Embedding Flow Matching for Artist-style Mesh Generation
中文标题:PolyFlow:面向艺术风格网格生成的连续拓扑嵌入流匹配
作者:Chunshi Wang, Haohan Weng, Junliang Ye, Biwen Lei, Yang Li, Zibo Zhao, Zeqiang Lai, Kaiyi Zhang, Yunhan Yang, Zhuo Chen, Chunchao Guo, Yawei Luo
Autoregressive Transformers dominate high-quality mesh generation by producing artist-worthy topologies, yet their inherent sequential decoding induces substantial computational overhead, falling orders of magnitude slower than parallel generative models. On the other hand, while continuous diffusion and flow-matching methods support efficient parallel synthesis across a variety of domains, they cannot be directly applied to meshes: mesh connectivity is inherently discrete and incompatible with standard continuous noise injection and denoising operations. To resolve this fundamental incompatibility, we introduce a compact topology embedder that projects discrete mesh vertex positions and normals into continuous per-vertex embeddings, where the original discrete adjacency information can be faithfully recovered via spacetime distance thresholding. After pretraining and freezing this embedder, any raw mesh can be fully converted into a continuous per-vertex state space unifying position, normal, and implicit topological attributes. Built upon this novel continuous mesh representation, we present PolyFlow, a Transformer-based flow-matching framework that achieves fully parallel vertex state denoising conditioned on extracted point-cloud features. During inference, our model completes generation rapidly via an ODE solver, and supports explicit, precise control over output mesh resolution by directly specifying the target vertex count. Extensive evaluations on the Toys4K benchmark demonstrate that PolyFlow surpasses state-of-the-art autoregressive baselines in both Chamfer Distance and Hausdorff Distance.
自回归Transformer在生成艺术级拓扑结构的高质量网格方面占据主导地位,但其固有的顺序解码机制带来了巨大的计算开销,比并行生成模型慢数个数量级。另一方面,尽管连续扩散和流匹配方法在多个领域支持高效的并行合成,但它们无法直接应用于网格:网格连接性本质上是离散的,与标准的连续噪声注入和去噪操作不兼容。为解决这一根本不兼容性,我们引入了一个紧凑的拓扑嵌入器,将离散的网格顶点位置和法线投影到连续的逐顶点嵌入中,其中原始的离散邻接信息可通过时空距离阈值精确恢复。预训练并冻结该嵌入器后,任何原始网格都可以完全转换为统一的连续逐顶点状态空间,融合位置、法线和隐式拓扑属性。基于这一新颖的连续网格表示,我们提出了PolyFlow,一个基于Transformer的流匹配框架,在提取的点云特征条件下实现完全并行的顶点状态去噪。推理过程中,我们的模型通过ODE求解器快速完成生成,并支持通过直接指定目标顶点数来精确控制输出网格分辨率。在Toys4K基准上的广泛评估表明,PolyFlow在Chamfer距离和Hausdorff距离上均超越了最先进的自回归基线方法。
AR-CoPO: Align Autoregressive Video Generation with Contrastive Policy Optimization
中文标题:AR-CoPO:基于对比策略优化的自回归视频生成对齐方法
作者:Dailan He, Guanlin Feng, Xingtong Ge, Yi Zhang, Bingqi Ma, Guanglu Song, Yu Liu, Hongsheng Li
Streaming autoregressive (AR) video generators combined with few-step distillation achieve low-latency, high-quality synthesis, yet remain difficult to align via reinforcement learning from human feedback (RLHF). Existing SDE-based GRPO methods face challenges in this setting: few-step ODEs and consistency model samplers deviate from standard flow-matching ODEs, and their short, low-stochasticity trajectories are highly sensitive to initialization noise, rendering intermediate SDE exploration ineffective. We propose AR-CoPO (AutoRegressive Contrastive Policy Optimization), a framework that adapts the Neighbor GRPO contrastive perspective to streaming AR generation. AR-CoPO introduces chunk-level alignment via a forking mechanism that constructs neighborhood candidates at a randomly selected chunk, assigns sequence-level rewards, and performs localized GRPO updates. We further propose a semi-on-policy training strategy that complements on-policy exploration with exploitation over a replay buffer of reference rollouts, improving generation quality across domains. Experiments on Self-Forcing demonstrate that AR-CoPO improves both out-of-domain generalization and in-domain human preference alignment over the baseline, providing evidence of genuine alignment rather than reward hacking.
流式自回归(AR)视频生成器结合少步蒸馏实现了低延迟、高质量的合成,但仍然难以通过强化学习从人类反馈(RLHF)进行对齐。现有的基于SDE的GRPO方法在此场景下面临挑战:少步ODE和一致性模型采样器偏离标准流匹配ODE,且其短低随机性轨迹对初始化噪声高度敏感,使得中间SDE探索失效。我们提出AR-CoPO(自回归对比策略优化),一个将邻域GRPO对比视角适应于流式AR生成的框架。AR-CoPO通过分叉机制引入分块级对齐,在随机选择的分块处构建邻域候选,分配序列级奖励,并执行局部GRPO更新。我们进一步提出一种半在线策略训练策略,通过对参考rollout的回放缓冲区进行利用来补充在线探索,从而在多个领域提升生成质量。在Self-Forcing数据集上的实验表明,AR-CoPO在域外泛化和域内人类偏好对齐方面均优于基线,提供了真正对齐而非奖励黑客的证据。
InteractiveAvatar: Real-Time Streaming Video Generation for Consistent and Intent-Aware Avatars
中文标题:InteractiveAvatar:用于一致性和意图感知虚拟形象的实时流视频生成
作者:Quanyue Song, Yishan He, Yanfei Zhang, Shihao Cheng, Zhixiang He, Zhizhi Guo, Chi Zhang, Xuelong Li, Caigui Jiang
Recent diffusion-based models have enabled realistic audio-driven avatar generation in real-time streaming. However, existing approaches struggle to maintain visual temporal consistency and fail to explicitly perceive user intent in complex interactive streaming scenarios. To address these challenges, we propose InteractiveAvatar, a real-time infinite-streaming video generation framework that supports visually consistent avatar video generation and intent-aware interactions. With autoregressive distillation, InteractiveAvatar achieves real-time str-eaming generation of human avatars over arbitrarily long durations. For visual consistency, we introduce a Long-Short Visual Memory (LSVM) mechanism that flexibly compresses historical visual information into compact tokens, preserving both short-range coherence and long-term consistency. To generate avatars with speeches and actions aligned with user intent, we propose a Reasoning-Reaction Module (RRM), which incorporates a State-Cycling strategy and a Cache-Switching mechanism. Extensive experimental results over diverse scenarios demonstrate that our method achieves state-of-the-art visual consistency in long-duration generation, while enabling complex user-avatar interaction in real time.
近年来,基于扩散的模型已能够实现实时流式传输下的逼真音频驱动虚拟形象生成。然而,现有方法在复杂的交互式流媒体场景中难以保持视觉时间一致性,且无法明确感知用户意图。为解决这些挑战,我们提出了 InteractiveAvatar,一个实时无限流视频生成框架,支持视觉一致的虚拟形象视频生成和意图感知交互。通过自回归蒸馏,InteractiveAvatar 可在任意时长下实现人类虚拟形象的实时流式生成。为实现视觉一致性,我们引入了长短期视觉记忆(LSVM)机制,将历史视觉信息灵活压缩为紧凑的令牌,同时保留短期一致性和长期一致性。为了生成与用户意图对齐的语音和动作虚拟形象,我们提出了推理-反应模块(RRM),该模块整合了状态循环策略和缓存切换机制。在多种场景下的大量实验结果表明,我们的方法在长时间生成中实现了最先进的视觉一致性,同时能够支持实时的复杂用户-虚拟形象交互。
2025年6月 Diffusion 论文 Daily Overview
今日 Diffusion 相关论文呈现多元化应用趋势,涵盖图像/视频生成、3D 重建、语音合成、自动驾驶、物理学模拟等多个领域。整体来看,研究重点集中在效率优化(如一步生成、知识蒸馏、缓存机制)、可控性增强(如属性编辑、prompt-free 生成、unlearning)以及跨模态融合(视频-图像-动作统一框架)。值得关注的是,Flow Matching 作为新兴范式持续受到关注,而物理约束与 Diffusion 的结合(如材料正则化、逆渲染)成为新亮点。
重点论文推荐:
- SwiftAudio(2606.31259):数据高效的纯字幕蒸馏一步式文本到音频 Diffusion,显著降低训练成本且保持生成质量,对实际部署意义重大。
- Preserve the Hard(2606.31603):提出不确定性引导的合成训练数据增强策略,通过 Diffusion 模型提升模型在困难样本上的鲁棒性,方法新颖且实用。
- Entropy-Controlled Flow Matching(2602.22265):从熵角度重新审视 Flow Matching,提供更灵活的可控生成范式,为扩散模型理论贡献新视角。
- InteractiveAvatar(2606.22905):实时流式视频生成一致意图感知的 Avatar,实现高质量实时数字人生成,应用前景广阔。
- Editing Everything Everywhere(2606.31278):实现任意地点、任意时间的统一编辑框架,推动高效图像编辑进入新阶段。
Surrogate-Gated Generation and Foundation-Model Embeddings for Bayesian Materials Design
中文标题:面向贝叶斯材料设计的代理门控生成与基础模型嵌入
作者:Sk Md Ahnaf Akif Alvi, Jan Janssen, Danny Perez, Douglas Allaire, Raymundo Arroyave
Closed-loop materials discovery iterates between proposing candidate structures and evaluating their properties, and property evaluation dominates the cost. In the generative variant, a learned prior proposes candidate crystals and a property oracle scores them; we ask whether a cheap probabilistic surrogate can triage the generator's output, and what such a surrogate must do well. Across three architecturally distinct pretrained diffusion priors (MatterGen, CrystalFlow, ADiT) and two targets (room-temperature heat capacity and bulk modulus), we insert a Gaussian process acquisition gate between structure generation and the oracle in an RL-steered generative workflow. The gate matches or exceeds ungated fine-tuning of the generative model while capping oracle calls at a fixed per-cycle budget. Budget-matched ablations isolate the mechanism. At an identical four-call budget, ranking-based selection outperforms arbitrary selection, confirming that the gain comes from the surrogate&x27;s choice; the gate comes within $\sim$9\% of exhaustive oracle spending at roughly one-fifth of the calls. A density-functional-theory check of the bulk-modulus discoveries confirms the learned oracle to within 2.5\% on average and the surrogate's ranking of the generated structures at Spearman $\rho = 0.94$. A cross-factorial benchmark of surrogate performance spanning mechanical, electronic, and vibrational properties identifies pretrained ORB embeddings with a Gaussian process as the most reliable combination, which we adopt as the building blocks of the proposed workflow. The complete pipeline is released as open-source software.
闭环材料发现通过迭代提出候选结构并评估其性能来进行,其中性能评估占据了主要成本。在生成式变体中,学习先验提出候选晶体,属性预言机对其进行评分;我们探讨是否可以采用廉价的概率代理模型对生成器的输出进行分类筛选,以及该代理模型需要具备哪些能力。在三种架构不同的预训练扩散先验(MatteGen、CrystalFlow、ADiT)和两个目标属性(室温热容和体积模量)上,我们在结构生成和预言机之间插入了高斯过程获取门,形成强化学习引导的生成工作流。该门在将预言机调用限制在每轮固定预算内的同时,达到了或超过了生成模型无门控微调的性能。预算匹配的消融实验隔离了具体机制。在相同的四次调用预算下,基于排序的选择优于任意选择,证实了收益来自代理模型的选择;该门在约五分之一的调用次数下达到了穷举式预言机消耗约9%的性能水平。对体积模量发现的密度泛函理论验证确认,学习预言机的平均误差在2.5%以内,代理模型对生成结构的排序斯皮尔曼相关系数ρ = 0.94。跨越机械、电子和振动性能的代理模型交叉因子基准测试确定,预训练ORB嵌入配合高斯过程是最可靠的组合,我们将其作为所提工作流的基本构建块。完整流程已作为开源软件发布。
Unsupervised Thermodynamics of Molecular Diffusion Models: Action-Operator Semantics and Auditable Free-Energy Readout
中文标题:分子扩散模型的无监督热力学:动作-算符语义学与可审计自由能读出
作者:Wenjie Xi
Diffusion models are increasingly utilized for modeling molecular structures and conformational ensembles, yet the thermodynamic meaning of their learned representations and scores remains elusive. To resolve this ambiguity, we introduce a mathematically consistent action-operator framework natively compatible with diffusion models. By defining a fixed molecular environment as a base action $S_0(x)$ and an alchemical perturbation as an operator $O(x)$, standard diffusion noising induces effective noised actions and operators whose gradients and alchemical derivatives are directly represented by the model's learned fields. This rigorous self-consistency enables a ``noisy operator bridge&x27;' capable of reading out free-energy differences ($\Delta F$) from endpoint ensembles and per-frame evaluations. In controlled experiments on alanine dipeptide systems, we show that incorporating physical inductive biases enables partial recovery of the base action and perturbation operator. When applied to a challenging C6-H to C6-F ligand-pocket nonbonded perturbation (185L/IND) with negligible phase-space overlap, our supervised bridge estimates the alchemical $\Delta F$ within approximately $1\ k_\mathrm{B}T$ of a stable 19-state MBAR reference. Finally, we demonstrate that endpoint coordinates and binary labels alone are sufficient to partially recover the operator shape and a centered free-energy scale without any force or action supervision. This work provides a rigorous path toward transforming generative molecular diffusion models from black-box coordinate samplers into auditable thermodynamic estimators.
扩散模型越来越多地被用于建模分子结构和构象集合,然而其学习到的表示和分数的热力学意义仍然难以捉摸。为了解决这一模糊性,我们引入了一个数学上一致的动作-算符框架,该框架与扩散模型原生兼容。通过将固定分子环境定义为基态动作$S_0(x)$,并将化学变化扰动定义为算符$O(x)$,标准扩散噪声处理会诱导出有效的噪声动作和算符,其梯度以及化学变化导数直接由模型学习到的场表示。这种严格的自我一致性使得一种“噪声算符桥”能够从端点集合和逐帧评估中读出自由能差($\Delta F$)。在丙氨酸二肽系统的对照实验中,我们表明加入物理归纳偏置能够部分恢复基态动作和扰动算符。当应用于一个具有可忽略相空间重叠的挑战性C6-H到C6-F配体口袋非键合扰动(185L/IND)时,我们的监督桥在约$1\ k_\mathrm{B}T$的范围内估算了化学变化$\Delta F$,与稳定的19态MBAR参考值一致。最后,我们证明仅凭端点坐标和二元标签就足以部分恢复算符形状和居中的自由能尺度,无需任何力或动作监督。这项工作为将生成式分子扩散模型从黑盒坐标采样器转化为可审计的热力学估计器提供了一条严谨的路径。
DSIP: A Dynamic Coordination Planner for Signal-Free Intersections using Diffusion-Model-Based Multi-Agent Motion Planning
中文标题:DSIP:基于扩散模型的多智能体运动规划无信号灯交叉口动态协调规划器
作者:Qian Hu, Haoyang Peng, Songan Zhang, Ming Yang, Hongtei Eric Tseng
Traffic signal control at urban intersections inherently introduces stop-and-go behavior, resulting in increased delays and reduced traffic efficiency, especially under high traffic demand. With the emergence of connected and automated vehicles (CAVs), trajectory-level coordination has emerged as a high-potential strategy to augment or transcend conventional phase-based management. This paper proposes DSIP (Diffusion-model-based Signal-free Intersection Planner), a multi-agent motion planning framework driven by a generative diffusion process. DSIP shifts the intersection management paradigm from discrete temporal phasing to continuous multi-vehicle trajectory optimization. This work evaluates the theoretical upper-bound performance of this coordination strategy under idealized communication and execution conditions to isolate the core benefits of the diffusion-driven approach. Using the SUMO platform, we evaluate DSIP across diverse four-leg intersection configurations. Experimental results demonstrate that DSIP significantly reduces average delay and maintains higher average speed compared to both fixed-time signal control and state-of-the-art reinforcement-learning-based controllers, particularly in medium- to high-density traffic. These findings suggest that diffusion-based trajectory planning provides a scalable and robust foundation for future autonomous intersection management. By unlocking latent intersection capacity through software-defined coordination, this approach offers a cost-effective pathway to improve urban traffic flow efficiency without requiring physical infrastructure expansion.
城市交叉口的交通信号控制本质上会引发停停走走的行为,导致延误增加和交通效率降低,尤其是在高交通需求条件下。随着车联网与自动驾驶车辆(CAVs)的出现,轨迹级协调作为一种高潜力策略应运而生,用以增强或超越传统的基于相位的管理方式。本文提出DSIP(基于扩散模型的无信号灯交叉口规划器),一个由生成式扩散过程驱动的多智能体运动规划框架。DSIP将交叉口管理范式从离散时序相位转变为连续多车轨迹优化。本工作在理想通信和执行条件下评估了该协调策略的理论上界性能,以分离扩散驱动方法的核心优势。使用SUMO平台,我们在多种四岔路口配置下对DSIP进行了评估。实验结果表明,与固定配时信号控制和基于强化学习的最先进控制器相比,DSIP显著降低了平均延误并保持了更高的平均速度,尤其是在中高密度交通条件下。这些发现表明,基于扩散模型的轨迹规划为未来自主交叉口管理提供了可扩展且稳健的基础。通过软件定义协调释放交叉口的潜在容量,这种方法为提高城市交通流效率提供了一条无需物理基础设施扩展的经济有效途径。
OTCache: Optimal Transport for Geometry-Aware Caching in Diffusion Models
中文标题:OTCache:扩散模型中几何感知缓存的最优传输方法
作者:Huanlin Gao, Fang Zhao, Qiang Hui, Fuyuan Shi, Shaoan Zhao, Yantao Li, Chao Tan, Ting Lu, Yuren You, Kai Wang, Shiguo Lian
We propose OTCache, a training-free framework for accelerating diffusion sampling via caching schedule prediction. Existing graph-based caching methods reduce redundant computation by optimizing shortest-path objectives, but rely on an additive independence assumption, which often breaks down in the low NFE regime. To address this issue, OTCache models caching schedules across inference budgets as a smooth evolution in policy space, inspired by Optimal Transport (OT). The framework consists of three stages: (1) obtaining a high-fidelity \textbf{reference schedule} using a graph-based caching method under a conservative budget; (2) performing a lightweight anchor search under an extreme low-budget setting via Optuna optimization with an end-to-end perceptual objective; and (3) predicting schedules for target budgets via quantile interpolation between the reference and anchor policies using continuous warping representations. Experiments on FLUX.1 [dev], Qwen-Image, and HunyuanVideo show that OTCache achieves 4.5x, 4.7x, and 3.66x acceleration, respectively, while consistently improving generation fidelity over state-of-the-art caching baselines. This work provides a new perspective on accelerating diffusion models through Optimal-Transport-inspired schedule modeling. Code:https://github.com/UnicomAI/OTCache
我们提出了OTCache,这是一个无需训练的框架,通过缓存调度预测来加速扩散采样。现有的基于图的缓存方法通过优化最短路径目标来减少冗余计算,但依赖于加性独立假设,这在低NFE(函数评估次数)设置下往往不成立。为解决这一问题,OTCache受最优传输(Optimal Transport, OT)启发,将跨推理预算的缓存调度建模为策略空间中的平滑演变。该框架包含三个阶段:(1)在保守预算下使用基于图的缓存方法获取高保真度参考调度;(2)在极端低预算设置下通过Optuna优化进行轻量级锚点搜索,使用端到端感知目标;(3)通过连续扭曲表示在参考策略和锚点策略之间进行分位数插值,预测目标预算的调度。在FLUX.1 [dev]、Qwen-Image和HunyuanVideo上的实验表明,OTCache分别实现了4.5倍、4.7倍和3.66倍的加速,同时在生成保真度方面始终优于最先进的缓存基线方法。本工作为通过最优传输启发的调度建模加速扩散模型提供了新视角。代码:https://github.com/UnicomAI/OTCache
UniSAE: Unified Speech Attribute Editing on Speaker, Emotion and Low-Level Content via Discrete Phonetic Posteriorgram Modelling
中文标题:UniSAE:通过离散语音后验图建模实现说话人、情感和低级内容的统一语音属性编辑
作者:Chuanbo Zhu, Wuyou Zhou, Rongxiu Zhong, Shilei Zhang, Kun Qian, Yike Guo, Wei Xue
Speech editing aims to modify specific portions of an utterance while preserving the remaining speech. Existing approaches primarily focus on word-level content modification and typically treat content, speaker, and emotion editing as separate tasks, limiting both editing granularity and flexibility. We propose UniSAE, a unified speech attribute editing framework which supports composable speaker, emotion and content editing from sub-phoneme to word level within a single architecture. UniSAE introduces a Discrete Phonetic PosteriorGram (DPPG) representation that factorizes speech content into discrete tokens encoding phoneme identity, pronunciation variants, and duration, enabling direct phoneme- and sub-phoneme-level editing. For higher-level modifications, an autoregressive content transformer predicts edited DPPG sequences for word-level content editing. The edited sequences are rendered into speech by a diffusion-based acoustic decoder, conditioned on disentangled speaker and emotion representations. Experimental results demonstrate that the proposed unified framework supports precise speaker and emotion control, content editing at multiple granularities, and joint modification of all three attributes within a single framework.
语音编辑旨在修改语音中特定部分的同时保留其余内容。现有方法主要关注词级内容修改,并且通常将内容、说话人和情感编辑视为独立任务,这限制了编辑的粒度和灵活性。我们提出了UniSAE,一个统一的语音属性编辑框架,在一个架构中支持从亚音素级到词级的可组合说话人、情感和内容编辑。UniSAE引入了一种离散语音后验图(DPPG)表示,将语音内容分解为编码音素身份、发音变体和时长的离散标记,从而实现直接的音素级和亚音素级编辑。对于更高级别的修改,自回归内容变换器预测编辑后的DPPG序列,用于词级内容编辑。编辑后的序列通过基于扩散的声学解码器进行渲染,并以解耦的说话人和情感表示为条件。实验结果表明,所提出的统一框架支持精确的说话人和情感控制、多粒度的内容编辑,以及在单一框架内对所有三个属性的联合修改。
SwiftAudio: Data-Efficient Caption-Only Distillation for One-Step Text-to-Audio Diffusion-based Generation
中文标题:SwiftAudio:面向一步式文本到音频扩散生成的高效数据纯字幕蒸馏方法
作者:Binh Mai, Tran Quoc Bao Le, Hung Dinh, Cong Tran
Diffusion-based text-to-audio (TTA) models achieve impressive synthesis quality but suffer from high inference latency due to iterative multi-step denoising. Existing one-step approaches alleviate this issue but still rely on paired text--audio data during distillation. To address these limitations, we propose SwiftAudio, a one-step TTA framework that performs audio-free distillation from a pretrained diffusion teacher using only text captions. Specifically, we adapt Variational Score Distillation (VSD) to the audio domain and introduce a temporal smoothness regularization objective to encourage coherent latent audio representations. This design enables the student model to inherit the teacher's generative prior without requiring paired audio supervision and allows effective training with only approximately 45K captions. Experiments on AudioCaps and Clotho demonstrate that SwiftAudio achieves state-of-the-art performance among strict one-step methods and substantially narrows the gap to multi-step diffusion systems. Project page: https://swiftaudio.org/
基于扩散的文本到音频(TTA)模型在合成质量上表现优异,但因迭代多步去噪过程而面临高推理延迟问题。现有的一步式方法虽能缓解这一困境,但蒸馏过程中仍依赖成对的文本-音频数据。为解决这些局限性,我们提出SwiftAudio,一个一步式TTA框架,可仅使用文本字幕从预训练扩散教师模型进行无音频蒸馏。具体而言,我们将变分分数蒸馏(VSD)适配到音频领域,并引入时间平滑正则化目标以促进连贯的潜在音频表示。该设计使学生模型能够继承教师的生成先验,而无需成对音频监督,并可仅使用约45K条字幕进行有效训练。在AudioCaps和Clotho数据集上的实验表明,SwiftAudio在严格的一步式方法中达到最优性能,并显著缩小了与多步扩散系统的差距。项目主页:https://swiftaudio.org/
Temperature Field Reconstruction of Tungsten Monoblock Divertor on EAST using Physics-aware Neural Operator Transformer
中文标题:基于物理感知神经算子Transformer的EAST装置钨单块偏滤器温度场重建
作者:Zikang Yan, Xiao Wang, Qingquan Yang, Zhendong Yang, Gaoting Chen, Zehua Chen, Bo Jiang, Jin Tang, Guosheng Xu
Accurate modeling of the divertor temperature field is essential for preventing material melting and damage and for extending the service life of fusion devices. However, conventional numerical methods, such as the Finite Element Method (FEM), are computationally expensive and therefore unsuitable for real-time applications. Therefore, a fast and generalizable method is required for real-time reconstruction of the divertor temperature field and subsequent real-time control. To address the above issue, we propose a Physics-aware Neural Operator Transformer (PNOT) to characterize the spatiotemporal evolution of the divertor temperature field. It models boundary heat-flux relations as a structured graph and employs graph attention to explicitly capture spatial physical dependencies. Inspired by physics-aware attention, we further develop a physics-aware neural operator module to aggregate query points with similar physical conditions via slicing and model heat diffusion, while a gradient-constrained Sobolev regularization loss enforces consistency between function values and their derivatives. Experimental results show that these physical constraints improve prediction accuracy while preserving physical consistency. The source code of this paper will be released on https://github.com/Event-AHU/OpenFusion
准确建模偏滤器温度场对于防止材料熔化和损伤以及延长聚变装置使用寿命至关重要。然而,有限元方法(FEM)等传统数值方法计算成本较高,不适用于实时应用。因此,需要一种快速且通用性强的偏滤器温度场实时重建方法,以实现后续的实时控制。针对上述问题,我们提出了一种物理感知神经算子Transformer(PNOT)来表征偏滤器温度场的时空演化。该方法将边界热通量关系建模为结构化图,并采用图注意力机制显式捕获空间物理依赖关系。受物理感知注意力启发,我们进一步开发了物理感知神经算子模块,通过切片方法聚合具有相似物理条件的查询点并建模热扩散过程,同时引入梯度约束的Sobolev正则化损失来保证函数值及其导数之间的一致性。实验结果表明,这些物理约束在保持物理一致性的同时提高了预测精度。本文的源代码将发布于https://github.com/Event-AHU/OpenFusion
Preserve the Hard, Regenerate the Rest: Uncertainty-Guided Synthetic Training Data Augmentation with Diffusion Models
中文标题:保留困难样本,重建其余部分:基于扩散模型的不确定性引导合成训练数据增强
作者:Nikolai R\"ohrich, Julian Glei{\ss}ner, Ahmed H. A. Ibrahim, Silvan Mertes, Tobias Huber
Semantic segmentation models struggle with data sparsity and rare or visually diverse regions, e.g., dense regions or small objects in aerial or autonomous mobility data. While synthetic augmentation is an appealing solution, directly generating new labeled data risks misalignment of labels and generated pixels. Existing solutions to this problem often rely on external models, or employ coarse heuristics such as indiscriminately augmenting all foreground objects or entire backgrounds, which wastes capacity on uninformative pixels. To address this, we propose an uncertainty-guided synthetic context augmentation strategy that strictly preserves label validity and efficiently maximizes pixel informativeness per synthetic sample - no external guardrails required. Using a baseline segmenter's predictive entropy, we identify uncertain semantic regions and inpaint only the complementary visual context. When fine-tuning the segmenter on this synthetic data, we compute the loss only over the original pixels, excluding inpainted regions. This focuses learning on the unmodified, uncertain regions while presenting them in novel contexts. We demonstrate substantial mIoU gains on Cityscapes, UAVID, and BDD100K with the largest gains on rare and difficult classes such as buses, trains, or (from the aerial perspective) cars. Our results demonstrate that uncertainty-guided context augmentation is a highly effective lever to improve segmentation performance on complex datasets, with code provided at https://github.com/XITASO/Preserve-the-Hard-Regenerate-the-Rest.
语义分割模型在处理数据稀疏以及稀有或视觉多样化区域(如航空数据或自动驾驶数据中的密集区域或小目标)时面临挑战。虽然合成增强是一种可行的解决方案,但直接生成新的标注数据存在标签与生成像素错位的风险。现有的解决方案通常依赖外部模型,或采用粗略的启发式方法,如无差别地增强所有前景对象或整个背景,这会导致在无信息像素上浪费资源。为解决这一问题,我们提出了一种不确定性引导的合成上下文增强策略,该策略严格保留标签有效性,并在每个合成样本中最大化像素信息量——无需外部约束条件。我们利用基线分割器的预测熵来识别不确定的语义区域,并仅对互补的视觉上下文进行修复。在使用该合成数据微调分割器时,仅对原始像素计算损失,排除修复区域。这使得学习聚焦于未修改的不确定区域,同时将它们置于新颖的上下文中。我们在 Cityscapes、UAVID 和 BDD100K 数据集上展示了显著的 mIoU 提升,在稀有且困难的类别(如公交车、火车,或从航空视角下的汽车)上提升最为明显。我们的结果表明,不确定性引导的上下文增强是提升复杂数据集分割性能的有效手段,代码已发布于 https://github.com/XITASO/Preserve-the-Hard-Regenerate-the-Rest。
Histogram-constrained Image Generation
中文标题:基于直方图约束的图像生成
作者:Haoming Liu, Yuanhe Guo, Yijia Cao, Shenji Wan, Hongyi Wen
Diffusion models have emerged as a dominant paradigm in generative modeling, enabling high-fidelity sampling from complex data distributions. Despite impressive capabilities, controlling diffusion models to produce outputs aligned with user intent remains an open challenge, especially when balancing global coherence with local precision. Existing control mechanisms vary in the granularity of their conditioning signals. For example, textual prompts guide generation globally through high-level semantics, while ControlNet-like approaches secure precise local structure via dense conditions. In this work, we introduce Histogram-constrained Image Generation (HIG), a novel control mechanism that falls into the middle ground of control granularity. Our framework enforces user-specified distributional constraints (e.g., color histograms or latent token distributions) during the generation process with exact precision. We model such control as an optimal transport (OT) problem and apply explicit guidance transformations during sampling, thereby driving the diffusion trajectory to align with the desired histogram. We demonstrate the versatility of HIG across diverse applications, including constrained generation via color/latent histograms and high-capacity information embedding through histogram-level encoding. Our findings underscore the promise of distributional control, a flexible and interpretable control scheme that is fully compatible with existing control mechanisms, diversifying the hybrid strategies for controllable image generation. Our project page is available at: https://maps-research.github.io/hig/.
扩散模型已成为生成式建模的主流范式,能够从复杂数据分布中进行高保真采样。尽管能力出色,控制扩散模型生成符合用户意图的输出仍然是一个开放性挑战,尤其是在平衡全局一致性与局部精度方面。现有控制机制的条件信号粒度各异。例如,文本提示通过高级语义全局引导生成,而类似ControlNet的方法通过密集条件确保精确的局部结构。在本工作中,我们提出了直方图约束图像生成(HIG),这是一种处于控制粒度中间层次的新型控制机制。我们的框架在生成过程中以精确精度强制执行用户指定的分布约束(如颜色直方图或潜在token分布)。我们将这种控制建模为最优传输(OT)问题,并在采样过程中应用显式引导变换,从而驱动扩散轨迹与期望的直方图对齐。我们展示了HIG在多种应用中的通用性,包括通过颜色/潜在直方图的约束生成以及通过直方图级别编码的高容量信息嵌入。我们的发现强调了分布控制的潜力——一种完全兼容现有控制机制的灵活且可解释的控制方案,为可控图像生成的混合策略提供了更多选择。我们的项目页面见:https://maps-research.github.io/hig/.
Look But Don't Touch with Sparse Autoencoders for Unlearning in Diffusion Models
中文标题:稀疏自编码器在扩散模型反学习中的应用:可远观而不可亵玩
作者:Enrico Cassano, Riccardo Renzulli, Rayyan Ahmed, Marco Grangetto, Stephan Alaniz
Sparse autoencoders (SAEs) have recently been proposed as interpretable tools for concept-level manipulation, under the assumption that isolated features can serve as controllable intervention points. In this work, we systematically evaluate this assumption in the context of object erasure and steering in diffusion models. We show that while SAEs reliably detect and localize semantic concepts within diffusion model activations, direct intervention in their latent space frequently induces out-of-distribution activations, resulting in severe visual artifacts. To disentangle detection from intervention, we use SAE activations purely as semantic detectors to identify image regions containing the target object, and replace those patch embeddings with the ones that do not contain it. This detection-based replacement preserves the diffusion model's activation statistics and produces significantly cleaner erasure results than latent steering. Our findings reveal a fundamental gap between concept detection and concept intervention in diffusion models: monosemantic or sparse features are not inherently suitable as control knobs for steering. These results position SAEs as powerful interpretability tools for analyzing generative models, but highlight important limitations when used for direct manipulation, such as unlearning.
稀疏自编码器(SAE)最近被提出作为可解释性工具,用于概念层面的操作,其基本假设是分离的特征可以作为可控制的干预点。本工作系统地评估了这一假设在扩散模型中对象擦除和引导方面的有效性。我们发现,虽然SAE能够可靠地检测和定位扩散模型激活中的语义概念,但直接干预其潜在空间常常诱导分布外激活,导致严重的视觉伪影。为了将检测与干预分离,我们仅将SAE激活用作语义检测器,以识别包含目标对象的图像区域,并用不包含该对象的patch嵌入替换这些区域。这种基于检测的替换方式保持了扩散模型的激活统计量,产生了比潜在引导显著更干净的擦除结果。我们的发现揭示了扩散模型中概念检测与概念干预之间的根本性差距:单语义或稀疏特征并不天然适合作为引导的控制旋钮。这些结果将SAE定位为分析生成模型的有力可解释性工具,但同时也强调了将其用于直接操作(如反学习)时存在的重要局限性。
Diffusion Crossover: Defining Evolutionary Recombination in Diffusion Models via Noise Sequence Interpolation
中文标题:扩散交叉:通过噪声序列插值在扩散模型中定义进化重组
作者:Chisato Kumada, Satoru Hiwa, Tomoyuki Hiroyasu
Interactive Evolutionary Computation (IEC) provides a powerful framework for optimizing subjective criteria such as human preferences and aesthetics, yet it suffers from a fundamental limitation: in high-dimensional generative representations, defining crossover in a semantically consistent manner is difficult, often leading to a mutation-dominated search. In this work, we explicitly define crossover in diffusion models. We propose Diffusion crossover, which formulates evolutionary recombination as step-wise interpolation of noise sequences in the reverse process of Denoising Diffusion Probabilistic Models (DDPMs). By applying spherical linear interpolation (Slerp) to the noise sequences associated with selected parent images, the proposed method generates offspring that inherit characteristics from both parents while preserving the geometric structure of the diffusion process. Furthermore, controlling the time-step range of interpolation enables a principled trade-off between diversity (exploration) and convergence (exploitation). Experimental results using PCA analysis and perceptual similarity metrics (LPIPS) demonstrate that Diffusion crossover produces perceptually smooth and semantically consistent transitions between parent images. Qualitative interactive evolution experiments further confirm that the proposed method effectively supports human-in-the-loop image exploration. These findings suggest a new perspective: diffusion models are not only powerful generators, but also structured evolutionary search spaces in which recombination can be explicitly defined and controlled.
交互式进化计算(IEC)为优化人类偏好和美学等主观标准提供了强大的框架,但它存在一个根本性限制:在高维生成表示中,以语义一致的方式定义交叉是困难的,往往导致搜索被变异所主导。在这项工作中,我们明确地在扩散模型中定义交叉。我们提出了扩散交叉方法,将进化重组表述为去噪扩散概率模型(DDPM)逆向过程中的噪声序列逐步插值。通过对所选父图像关联的噪声序列应用球面线性插值(Slerp),该方法生成的后代继承了双亲的特征,同时保留了扩散过程的几何结构。此外,控制插值的时间步范围能够实现多样性(探索)与收敛性(利用)之间的原则性权衡。使用主成分分析(PCA)和感知相似度度量(LPIPS)的实验结果表明,扩散交叉在父图像之间产生了感知平滑且语义一致的过渡。定性交互式进化实验进一步证实,该方法有效地支持了人机交互式图像探索。这些发现提供了一个新视角:扩散模型不仅是强大的生成器,也是结构化的进化搜索空间,其中重组可以被明确定义和控制。
Quantum Flow Matching
中文标题:量子流匹配
作者:Zidong Cui, Pan Zhang, Ying Tang
The flow matching has rapidly become a dominant paradigm in classical generative modeling, offering an efficient way to interpolate between two complex distributions. We extend this idea to the quantum realm and introduce the Quantum Flow Matching (QFM), a quantum-circuit realization that offers efficient interpolation between two density matrices. QFM offers systematic preparation of density matrices and generation of samples for accurately estimating observables, and can be realized on quantum computers without the need for costly circuit redesigns. We validate its versatility on a set of applications: (i) generating target states with prescribed magnetization and entanglement entropy, (ii) estimating nonequilibrium free-energy differences to test the quantum Jarzynski equality, and (iii) expediting the study on superdiffusion. These results position QFM as a unifying and promising framework for generative modeling across quantum systems.
流匹配已成为经典生成式建模的主导范式,提供了一种在两个复杂分布之间进行插值的有效方法。我们将这一思想拓展到量子领域,提出了量子流匹配(QFM)——一种量子电路实现,能够在两个密度矩阵之间实现高效插值。QFM提供了对密度矩阵的系统化制备以及用于准确估计可观测量的样本生成,并且可以在量子计算机上实现,无需进行代价高昂的电路重新设计。我们在一系列应用场景中验证了其通用性:(i)生成具有指定磁化强度和纠缠熵的目标态;(ii)估计非平衡自由能差以检验量子Jarzynski等式;(iii)加速超扩散研究。这些结果表明,QFM作为量子系统生成式建模的统一框架具有良好的应用前景。
CharDiff-LP: A Diffusion Model with Character-Level Guidance for License Plate Image Restoration
中文标题:CharDiff-LP:一种基于字符级引导的车牌图像修复扩散模型
作者:Kihyun Na, Gyuhwan Park, Injung Kim
License plate image restoration is important not only as a preprocessing step for license plate recognition but also for enhancing evidential value, improving visual clarity, and enabling broader reuse of license plate images. We propose a novel diffusion-based framework with character-level guidance, CharDiff-LP, which effectively restores and recognizes severely degraded license plate images captured under realistic conditions. CharDiff-LP leverages fine-grained character-level priors extracted through external segmentation and Optical Character Recognition (OCR) modules tailored for low-quality license plate images. For precise and focused guidance, CharDiff-LP incorporates a novel Character-guided Attention through Region-wise Masking (CHARM) module, which ensures that each character's guidance is restricted to its own region, thereby avoiding interference with other regions. In experiments, CharDiff-LP significantly outperformed baseline restoration models in both restoration quality and recognition accuracy, achieving a 28.3% relative reduction in character error rate (CER) on the Roboflow-LP dataset compared with the best-performing baseline.
车牌图像修复不仅是车牌识别的重要预处理步骤,还能够增强证据价值、提升视觉清晰度并实现车牌图像更广泛的复用。本文提出了一种基于扩散模型的新型字符级引导框架 CharDiff-LP,该框架能够有效修复在现实条件下拍摄的严重退化车牌图像。CharDiff-LP 利用针对低质量车牌图像定制的外部分割和光学字符识别(OCR)模块提取细粒度字符级先验知识。为实现精确且聚焦的引导,CharDiff-LP 引入了一种新颖的字符引导区域掩码注意力(CHARM)模块,该模块确保每个字符的引导仅限制在其自身区域内,从而避免对其他区域产生干扰。在实验中,CharDiff-LP 在修复质量和识别准确率方面均显著优于基线修复模型,在 Roboflow-LP 数据集上相较于最佳基线模型实现了 28.3% 的字符错误率(CER)相对降低。
Not Every Time and Frequency Need to Be Forgotten in Diffusion Unlearning
中文标题:并非所有时间和频率都需在扩散模型反学习中遗忘
作者:Jinseong Park, Mijung Park
Data unlearning aims to remove the influence of specific training samples from a trained model. In fine-tuning methods, data unlearning relies primarily on loss maximization over forget samples, which often leads to quality degradation or incomplete forgetting. Existing methods perform unlearning uniformly across diffusion stages, ignoring diffusion dynamics from noise to data. Our systematic study of diffusion phases shows that forgetting in diffusion models is uneven across time and frequency, with theoretical justification of distributive distortion and forgetting-utility trade-off. By selectively forgetting time and frequency in diffusion models, we achieve both higher unlearning success rates and improved generation quality across diverse settings, including both conditional and unconditional scenarios. We also introduce an improved SSCD metric that measures dissimilarity using a normalized perturbation distance. Together, we provide practical insights for understanding and improving data unlearning in diffusion models.
数据反学习旨在消除特定训练样本对已训练模型的影响。在微调方法中,数据反学习主要依赖于对遗忘样本进行损失最大化,这往往导致质量下降或遗忘不完整。现有方法在扩散阶段统一执行反学习,忽略了从噪声到数据的扩散动态。我们对扩散阶段的系统研究表明,扩散模型中的遗忘在时间和频率上是不均匀的,并从理论上论证了分布畸变和遗忘-效用权衡。通过在扩散模型中选择性地遗忘时间和频率,我们实现了更高的反学习成功率和更好的生成质量,涵盖条件和无条件等多种场景。我们还提出了一种改进的SSCD指标,该指标使用归一化扰动距离来衡量差异性。综上,我们为理解和改进扩散模型中的数据反学习提供了实践洞察。
Rethinking Garment Conditioning in Diffusion-based Virtual Try-On: Decouple, Don't Denoise
中文标题:基于扩散的虚拟试穿中服装条件化的再思考:解耦而非去噪
作者:Kihyun Na, Jinyoung Choi, Injung Kim
Virtual Try-On (VTON) synthesizes realistic images of a person wearing a target garment, with broad applications in e-commerce and fashion. Diffusion-based dual-UNet methods achieve strong results but double the parameters by dedicating a separate network to garment conditioning. Spatial concatenation offers a simpler single-network alternative, yet both UNet- and DiT-based instantiations report that full fine-tuning is ineffective, and the community has settled for attention-only training. We ask: why does full fine-tuning fail, and can this be resolved? Through what is, to our knowledge, the first visualization study of dual-UNet reference network behavior, we identify a unifying insight: garment conditioning must be decoupled from the denoising process. Spatial concatenation violates this by embedding the garment within the denoising target, causing three conflicts: guidance leakage, gradient competition, and train-test discrepancy. We derive three design principles to restore this decoupling and implement them as a pure recipe atop a standard architecture with no modification. The resulting model, DeCo-VTON (860M params), achieves single-network state of the art, matching the dual-UNet state of the art at half the cost while being preferred in human evaluation.
虚拟试穿(VTON)可合成人物穿着目标服装的真实图像,在电子商务和时尚领域具有广泛应用。基于扩散的双UNet方法取得了优异成绩,但通过为服装条件化配置专用网络而使参数量翻倍。空间拼接提供了更简单的单网络替代方案,然而无论是UNet还是DiT架构的实例化,都报告全参数微调效果不佳,学界因而退而求其次采用仅注意力训练。我们不禁追问:全参数微调为何失效,能否加以解决?据我们所知,通过对双UNet参考网络行为的首次可视化研究,我们揭示了一个统一见解:服装条件化必须与去噪过程解耦。空间拼接违反了这一原则,因为它将服装嵌入去噪目标中,导致三类冲突:引导泄漏、梯度竞争和训练-测试差异。我们据此推导出三条设计原则以恢复解耦,并将其作为纯配方实现于标准架构之上,无需任何修改。所得模型DeCo-VTON(8.6亿参数)实现了单网络最先进水平,以一半成本达到双UNet最先进水平的性能,且在人类评估中更受青睐。
SyncCache: Exploiting Asymmetric Dynamics for Fast Audio-Driven Portrait Animation
中文标题:SyncCache:利用非对称动力学实现快速音频驱动人像动画
作者:Juncheng Ma, Yuxuan Du, Yanan Sun, Zhening Xing, Changlin Li, Zhenyu Tang, Bo Li, Peng-Tao Jiang, Li Yuan, Daquan Zhou, Yonghong Tian
Diffusion Transformers (DiTs) have significantly advanced audio-driven portrait animation, but their high computational cost leads to substantial inference latency. Although training-free diffusion caching accelerates inference significant, existing methods are primarily developed for text-conditioned generation and overlook the spatial and modality imbalances inherent in audio-driven portrait animation. In this paper, we propose SyncCache, a training-free caching acceleration method tailored for DiT-based portrait animation that explicitly exploits asymmetric dynamics. Specifically, high-frequency dynamics driven by audio conditions and concentrated in human regions are more challenging and critical to cache and reuse than the low-frequency visual background in portrait animation. First, we introduce Spatially-Asymmetric Probing to prioritize error sensitivity in dynamic human region. Second, through Modality-Decoupled Caching, we bypass heavy DiT block by reusing stable inter-block residuals, while continuously recomputing lightweight audio blocks to preserve precise lip synchronization. Furthermore, we introduce a cache ratio to control cache capacity and formulate memory-adaptive cache selection as an offline dynamic programming problem without online overhead. Extensive experiments demonstrate that SyncCache achieves superior speed-quality trade-offs, delivering up to 4.12x acceleration on HunyuanVideo-Avatar and 3.75x on Wan-S2V with near-lossless visual fidelity and precise audio alignment.
扩散变换器(Diffusion Transformers,DiTs)显著推动了音频驱动人像动画的发展,但其高计算成本导致推理延迟显著增加。尽管无训练扩散缓存能够大幅加速推理,但现有方法主要针对文本条件生成,未能充分考虑音频驱动人像动画中固有的空间和模态不平衡问题。本文提出SyncCache,这是一种针对基于DiT的人像动画量身定制的无训练缓存加速方法,能够明确利用非对称动力学。具体而言,在人像动画中,由音频条件驱动的高频动力学集中在人脸区域,比低频视觉背景更具挑战性,更值得缓存和重用。首先,我们引入空间非对称探测,以优先考虑动态人脸区域的误差敏感性。其次,通过模态解耦缓存,我们通过重用稳定的块间残差来绕过繁重的DiT块,同时持续重新计算轻量级音频块以保持精确的唇同步。此外,我们引入缓存比率来控制缓存容量,并将内存自适应缓存选择表述为一个无需在线开销的离线动态规划问题。广泛实验表明,SyncCache实现了卓越的速度-质量权衡,在HunyuanVideo-Avatar上提供高达4.12倍的加速,在Wan-S2V上提供3.75倍的加速,同时保持近乎无损的视觉保真度和精确的音频对齐。
PhotoQuilt: Training-Free Arbitrary-Resolution Photomosaics via Bootstrapped Tiled Denoising
中文标题:PhotoQuilt:基于引导式分块去噪的任意分辨率照片马赛克生成方法
作者:Koorosh Roohi, Javad Rajabi, Andrew Fleet, Babak Taati
Photomosaics are large images whose local regions are seen as independent tiles while their overall arrangement forms a coherent scene. Generating them at high resolution, with every tile convincing in its own right, is computationally expensive, since the canvas must hold many detailed tiles at once. We present PhotoQuilt, a training-free framework that generates photomosaics at arbitrary resolution. Diffusion models struggle to satisfy both scales at once, as direct high-resolution generation is costly and tends toward one smooth image rather than a mosaic, while patch-based tiling keeps local detail but loses global structure. PhotoQuilt resolves this with a bootstrapped tiled denoising procedure. We first produce a global composition at low resolution to fix the layout, then upscale it in latent space and re-inject noise to restore generative capacity. Denoising proceeds within fixed tiles, so each forms its own image while the shared global structure holds them in one layout. Because tile generation is handled separately, PhotoQuilt scales to large canvases without quadratic attention cost. Experiments show that PhotoQuilt outperforms current baselines on both global structure and local realism.
照片马赛克是一类大型图像,其局部区域被视为独立图块,而整体布局则形成连贯的场景。以高分辨率生成此类图像且每个图块都具有令人信服的细节非常耗费计算资源,因为画布必须同时容纳许多细节丰富的图块。本研究提出 PhotoQuilt,一个无需训练的任意分辨率照片马赛克生成框架。扩散模型难以同时满足两个尺度的需求,因为直接高分辨率生成成本高昂且倾向于生成单一平滑图像而非马赛克效果,而基于图块的拼接方式虽能保留局部细节却会丢失全局结构。PhotoQuilt 通过引导式分块去噪方法解决这一问题。首先在低分辨率下生成全局构图以确定布局,然后在潜空间中将其放大并重新注入噪声以恢复生成能力。去噪过程在固定的图块内进行,因此每个图块形成独立的图像,而共享的全局结构则将它们保持在同一布局中。由于图块生成是独立处理的,PhotoQuilt 能够扩展到大型画布而无需承担二次方注意力开销。实验表明,PhotoQuilt 在全局结构和局部真实感方面均优于现有基线方法。
WarpI2I: Image Warping for Image-to-Image Translation
中文标题:WarpI2I:用于图像到图像迁移的图像扭曲方法
作者:Shen Zheng, Anurag Ghosh, Gaurav Parmar, Srinivasa Narasimhan
Image-to-image (I2I) translation has achieved strong results in tasks like human relighting and driving scene translation using latent diffusion models (LDMs). However, compact LDMs often struggle to preserve fine-grained structures because the encoder compresses high-resolution inputs into a spatially downsampled latent space. To address this issue, we propose a simple saliency-guided warp-unwarp framework that reallocates spatial representation toward salient regions before encoding, enabling better preservation of structural details without increasing latent resolution. The warped image is processed by the original diffusion model and then mapped back via an inverse warp. In addition, we propose a simple and efficient outpainting-based synthetic data generation pipeline to produce high-quality paired data for image relighting. Our method is model-agnostic, requires no architectural modification, and introduces negligible computational overhead. Experiments on human relighting, driving scene relighting, and translation demonstrate improved structural preservation, lighting faithfulness, and image quality, with our framework extending naturally to video via frame-by-frame application with good temporal stability. Project Webpage: https://shenzheng2000.github.io/WarpI2I.github.io
图像到图像(I2I)迁移在使用潜在扩散模型(LDMs)的任务中(如人物重光照和驾驶场景迁移)已取得了显著成效。然而,紧凑型LDMs通常难以保留细粒度结构,因为编码器将高分辨率输入压缩到空间下采样的潜在空间中。为解决这一问题,我们提出了一个简单的显著性引导的扭曲-反扭曲框架,在编码前将空间表示重新分配到显著区域,从而在无需增加潜在分辨率的情况下更好地保留结构细节。扭曲后的图像由原始扩散模型处理,然后通过逆扭曲映射回来。此外,我们提出了一个简单高效的外绘式合成数据生成流程,用于生成高质量的配对数据用于图像重光照。我们的方法与模型无关,无需修改架构,且引入的计算开销可忽略不计。在人物重光照、驾驶场景重光照和迁移任务上的实验表明,本方法在结构保持、光照保真度和图像质量方面均有提升,且该框架可通过逐帧应用自然扩展到视频领域,并展现出良好的时间稳定性。项目网页:https://shenzheng2000.github.io/WarpI2I.github.io
Diffusion-Based Material Regularization for Physics-Based Inverse Rendering
中文标题:基于扩散的材质正则化用于基于物理的逆渲染
作者:Jingwang Ling, Lifan Wu, Feng Xu, Shuang Zhao
Reconstructing physics-based 3D assets -- geometry, materials, and illumination -- from multi-view images is a core problem in computer graphics and vision, and a prerequisite for realistic relighting and editing. Physics-based inverse rendering offers an accurate image-formation model, but is severely underconstrained: without strong priors, illumination is baked into materials, and reconstructions generalize poorly to novel views and lighting. Data-driven diffusion models, in contrast, predict visually plausible materials, yet their predictions rarely satisfy the rendering equation and are not directly usable for physics-based rendering. We bridge these two paradigms rather than replacing either. Our key idea is to treat the predictions of a state-of-the-art diffusion model not as target material values but as a similarity kernel for optimization: we introduce a regularization loss that penalizes deviations in the optimized material over surface regions where the diffusion predictions are near-constant, while leaving the optimization free to match the input images. Built on this regularizer, our end-to-end pipeline jointly reconstructs geometry, materials, and illumination, yielding high-quality assets that drop into standard rendering pipelines and relight faithfully. On the Synthetic4Relight, Stanford-ORB, and DTC-Synthetic datasets, our method significantly outperforms state-of-the-art baselines in both reconstruction accuracy and relighting quality.
从多视角图像重建基于物理的3D资产(几何、材质和光照)是计算机图形学与视觉领域的核心问题,也是实现真实感重光照和编辑的前提条件。基于物理的逆渲染提供了准确的成像模型,但存在严重的欠约束问题:若无强先验知识,光照会被烘焙到材质中,导致重建结果在新视角和新光照条件下的泛化能力较差。相比之下,数据驱动的扩散模型能够预测视觉上可信的材质,但其预测结果往往无法满足渲染方程,无法直接用于基于物理的渲染。本研究并未取代其中任何一种范式,而是将两者桥接。我们的核心思想是将最先进的扩散模型预测结果而非作为目标材质值,而是作为优化的相似性核:引入一种正则化损失,在扩散预测结果接近恒定的表面区域惩罚优化后材质的偏差,同时允许优化过程自由匹配输入图像。基于该正则化器,我们的端到端管线联合重建几何、材质和光照,生成可直接嵌入标准渲染管线并实现真实重光照的高质量资产。在Synthetic4Relight、Stanford-ORB和DTC-Synthetic数据集上,本方法在重建精度和重光照质量方面均显著优于最先进的基线方法。
AnyMatch: Supercharging Universal Multi-Modal Image Matching with Large-Scale Single-View Images
中文标题:AnyMatch:利用大规模单视角图像增强通用多模态图像匹配
作者:Meng Yang, Zizhuo Li, Linfeng Tang, Fan Fan, Jiayi Ma
Multi-modal image matching is essential for visual localization and multi-sensor fusion, but it is hindered by the scarcity of large-scale training data with precise geometric annotations. Existing real-world datasets suffer from prohibitive costs, limited scene diversity, and errors in SfM-MVS pipelines, while synthetic methods struggle to maintain 3D geometric consistency or achieve photorealistic appearance. To address this, we propose AnyMatch, a novel framework that leverages abundant, easily accessible single-view images at minimal cost to generate rich multi-modal training data. AnyMatch integrates monocular depth estimation, 3D reprojection, diffusion-based inpainting, and crossmodal image translation to synthesize multi-view, multi-modal image pairs with 3D geometric fidelity. Crucially, our method provides annotations that strictly adhere to 3D geometric consistency through explicit 3D reprojection, avoiding SfM-MVS error accumulation. Furthermore, AnyMatch offers strong scalability, enabling controllable scene diversity and annotation difficulty via adjustable input and camera parameters. We construct Any-syn, a large-scale synthetic multi-modal dataset using AnyMatch. Experimental results show that matching networks (e.g., LoFTR, EDM, RoMa) fine-tuned on Any-syn achieve substantial performance gains on multi-modal benchmarks, exhibiting superior generalization and robustness compared to models trained on existing data.
多模态图像匹配对于视觉定位和多传感器融合至关重要,但受限于带有精确几何标注的大规模训练数据稀缺问题。现有的真实世界数据集存在成本高、场景多样性有限以及SfM-MVS pipeline误差累积等问题,而合成方法难以保持3D几何一致性或实现照片级真实感外观。针对这一问题,我们提出AnyMatch,这是一种新颖的框架,利用丰富且易于获取的单视角图像,以极低成本生成大量多模态训练数据。AnyMatch集成了单目深度估计、3D重投影、基于扩散的图像修复和跨模态图像翻译技术,以合成具有3D几何保真度的多视角、多模态图像对。关键在于,我们的方法通过显式3D重投影提供严格遵循3D几何一致性的标注,避免了SfM-MVS误差累积。此外,AnyMatch具有强大的可扩展性,能够通过可调整的输入和相机参数控制场景多样性和标注难度。我们使用AnyMatch构建了Any-syn,这是一个大规模合成多模态数据集。实验结果表明,在Any-syn上微调的匹配网络(如LoFTR、EDM、RoMa)在多模态基准测试中实现了显著的性能提升,与在现有数据上训练的模型相比表现出更强的泛化能力和鲁棒性。
InfiniVerse: Occupancy Guided Unbounded Scene Generation for Autonomous Driving
中文标题:InfiniVerse:面向自动驾驶的占用引导无界场景生成
作者:Xiaoyu Ye, Leheng Li, Xinyu Ji, Yingjie Cai, Hongda He, Xu Yan, Guanyi Zhao, Ying-Cong Chen, Bingbing Liu, Shuguang Cui, Zhen Li
Generating realistic, controllable, and temporally coherent urban environments is a critical yet unresolved challenge in the autonomous driving community. In this paper, we introduce InfiniVerse, a unified pipeline for long-range, 2D-3D-aligned, and controllable synthesis of dynamic urban scenes from a single frame. In practice, our approach first reconstructs a 3D occupancy representation from the input multi-view frame. This representation serves as a foundation for autoregressive scene extension along arbitrary trajectories. Subsequently, a video diffusion model translates the coarse occupancy grid into realistic, spatiotemporally consistent video sequences. Moreover, we propose a hierarchical sketch-and-refine paradigm, in which the generated videos are re-projected as image-conditioned feedback to enhance the 3D occupancy representation, establishing cross-modal alignment and mutual enhancement between the visual and spatial domains. Extensive evaluations on the Waymo Open Dataset and nuScenes demonstrate that InfiniVerse achieves state-of-the-art performance, with a FID of 6.4 and FVD of 67.97, significantly outperforming existing benchmarks in both duration and stability.
生成逼真、可控且时间一致的城市环境是自动驾驶领域一项关键但尚未解决的挑战。本文提出了InfiniVerse,一个从单帧图像进行长距离、2D-3D对齐且可控的动态城市场景合成的统一流程。具体而言,我们的方法首先从输入的多视角图像帧重建3D占用表示,以此为基础沿任意轨迹进行自回归场景扩展。随后,视频扩散模型将粗糙的占用网格转化为逼真且时空一致的视频序列。此外,我们提出了一种分层式草图-细化范式,将生成的视频重新投影为图像条件反馈,以增强3D占用表示,从而在视觉域与空间域之间建立跨模态对齐与相互增强。在Waymo Open Dataset和nuScenes数据集上的广泛评估表明,InfiniVerse达到了最先进的性能,FID得分为6.4,FVD得分为67.97,在持续时间和稳定性方面均显著优于现有基准方法。
WaterGen: Decoupling Scene and Medium in Underwater Image Generation
中文标题:WaterGen:水下图像生成中场景与介质的解耦
作者:Jiayi Wu, Tianfu Wang, Tianyi Xiong, Dehao Yuan, Xiaomin Lin, Md Jahidul Islam, Cornelia Fermuller, Christopher Metzler, Yiannis Aloimonos
Underwater computer vision tasks, such as detection, restoration, and segmentation, are limited by the scarcity of large-scale and diverse training data. We introduce WaterGen, a method for generating large-scale, realistic, and diverse underwater images that provides independent control of the scene and water medium conditions. Our approach treats underwater image generation as the decoupled control of two factors: realistic and diverse scene content (what is in the image), and accurate and controllable water medium effects (what the water does to the image). Existing methods generally achieve only part of this objective: they either provide controllability with limited realism or diversity, or generate realistic scenes without accurately and independently modeling water-medium effects. Our key insight, that allows us to avoid this compromise, is that scene generation and medium modeling can be decoupled within a latent diffusion framework, enabling diverse scene generation together with accurate and controllable underwater appearance. To do this, we decompose underwater image synthesis into two stages. First, we fine-tune the latent diffusion U-Net using degradation-free underwater images so that it learns to generate diverse and realistic latent embeddings of underwater scene content without medium-induced degradation. Second, we formulate the physically accurate medium degradation synthesis as a conditional decoding process applied to these latent embeddings. This decoupled design allows our model to generate diverse scenes with full control of underwater appearance. We leverage WaterGen to build large-scale synthetic underwater datasets that are diverse in scene structures and accurate in water effects and pseudo-labels. We demonstrate that our synthetic data consistently improve downstream performance in underwater restoration and semantic segmentation.
水下计算机视觉任务,如检测、恢复和分割,受限于大规模和多样化训练数据的匮乏。我们提出WaterGen,一种用于生成大规模、真实、多样化水下图像的方法,能够独立控制场景和水介质条件。我们的方法将水下图像生成视为两个因素的解耦控制:真实且多样化的场景内容(图像中的内容),以及准确且可控的水介质效果(水对图像的影响)。现有方法通常只能实现这一目标的局部:它们要么在可控性方面实现有限,要么在生成真实场景时无法准确独立地建模水介质效果。我们的关键见解是,场景生成和介质建模可以在潜在扩散框架内解耦,从而实现多样化场景生成以及准确可控的水下外观。为此,我们将水下图像合成分解为两个阶段。首先,我们使用无退化的水下图像对潜在扩散U-Net进行微调,使其学习生成多样化且真实的水下场景内容潜在嵌入,而不受介质引起的退化影响。其次,我们将物理精确的介质退化合成表述为应用于这些潜在嵌入的条件解码过程。这种解耦设计使我们的模型能够生成多样化场景,同时对水下外观实现完全控制。我们利用WaterGen构建了大规模合成水下数据集,这些数据集在场景结构上多样化,在水效果和伪标签上准确。我们证明,我们的合成数据在水下恢复和语义分割任务中持续提升了下游性能。
ForgeDrive: Bidirectional Cross-Conditioning for Unified Visual-Action Generation in Autonomous Driving
中文标题:ForgeDrive:用于自动驾驶中统一视觉-动作生成的双向交叉条件方法
作者:Xuchang Zhong, He Zheng, Chenxu Zhao, Tianxiong Lv, Hangqi Fan, Bohua Wang, Yushan Liu, Zhihao Liao, Leigang Luo, Congyang Zhao, Yang Cai
World-model-based autonomous driving endows the model with the ability to understand scene evolution. Yet this promise is undermined by the prevailing imagine-then-act paradigm, which allows errors from the more challenging visual generation stage to cascade into action planning. We introduce ForgeDrive, a unified autoregressive diffusion framework with visual-action cross-conditioning that closes this gap through act-then-imagine paradigm. ForgeDrive factorizes the future as a sequence of per-timestep frame-action pairs, intertwining each action with its corresponding visual observation. During training, we decouple the diffusion timesteps of the two modalities and introduce a UniDiffuser-style noise scheduler to get the ability to infer either modality from its counterpart and deepen understanding of relationships between images and actions. At inference, we propose a novel act-then-imagine inference paradigm, and find that at each step, action generation is a capability internalized during training, requiring no clean future frame as a prerequisite at inference time; instead, the generated action can improve the accuracy of future frame generation, which in turn enhances the quality of the next action. Additionally, we augment each step with future ego-status prediction, further sharpening planning ability. Extensive experiments on NAVSIM demonstrate that ForgeDrive not only unifies driving simulation, planning, and visual odometry into a single model, but also outperforms existing strong planners without any post-training strategy.
基于世界模型的自动驾驶赋予模型理解场景演变的能力。然而,这一愿景受到盛行的"先想象后行动"范式的阻碍,该范式允许更具挑战性的视觉生成阶段的错误级联传播到动作规划中。我们提出了 ForgeDrive,这是一个统一的自回归扩散框架,采用视觉-动作交叉条件,通过"先行动后想象"范式来弥补这一差距。ForgeDrive 将未来分解为一系列每时间步的帧-动作对,将每个动作与其对应的视觉观察交织在一起。在训练过程中,我们解耦了两种模态的扩散时间步,并引入了 UniDiffuser 风格的噪声调度器,以获得从一种模态推断另一种模态的能力,并加深对图像与动作之间关系的理解。在推理过程中,我们提出了一种新颖的"先行动后想象"推理范式,并发现每个步骤中,动作生成是训练过程中内化的能力,推理时不需要清晰未来帧作为先决条件;相反,生成的动作可以提高未来帧生成的准确性,从而增强下一个动作的质量。此外,我们在每个步骤中增加了未来自车状态预测,进一步提升了规划能力。在 NAVSIM 上的广泛实验表明,ForgeDrive 不仅将驾驶仿真、规划和视觉里程计统一到单个模型中,而且无需任何后训练策略就优于现有的强大规划器。
Editing Everything Everywhere All at Once
中文标题:编辑《瞬息全宇宙》:实现单次前向传播的多实例图像编辑
作者:Fabio Quattrini, Carmine Zaccagnino, Enis Simsar, Marta Tintor\'e Gazulla, Rita Cucchiara, Alessio Tonioni, Silvia Cascianelli
Editing multiple elements of an image in a single forward pass is a practical alternative to multi-turn image manipulation, offering improved efficiency and potentially better harmonization. However, when several instructions target different regions, semantic interference often leads to attribute leakage and poor edit disentanglement, especially as the number of edits increases. In this work, we propose MICE (Multi-Instance Concurrent Editing), a training-free strategy for scalable multi-instance image editing with Multimodal Diffusion Transformers. MICE modifies the additive bias of joint attention to regulate interactions between instance-specific edit instructions, latent, and context tokens identified via user-provided segmentation masks. Specifically, MICE allows intra-instance attention, penalizes interactions between neighboring region tokens, and suppresses unrelated cross-instance attention. As a result, our method enforces attribute binding while preserving global visual consistency. We evaluate MICE on LoMOE-Bench and introduce MICE-Bench, a more challenging benchmark with an average of 8.5 concurrent edits per image. The experiments demonstrate that our approach outperforms strong baselines and recent competitors in terms of visual quality preservation and faithfulness to the editing instructions.
在单次前向传播中编辑图像的多个元素是一种实用的多轮图像编辑替代方案,能够提高效率并可能实现更好的视觉协调。然而,当多条指令针对不同区域时,语义干扰常常导致属性泄漏和编辑解耦困难,尤其是在编辑数量增加的情况下。在本工作中,我们提出了MICE(Multi-Instance Concurrent Editing,多实例并发编辑),一种基于多模态扩散变换器的可扩展多实例图像编辑的无需训练策略。MICE通过修改联合注意力的附加偏置来调节用户提供的分割掩码所识别出的实例特定编辑指令、潜在token和上下文token之间的交互。具体而言,MICE允许实例内注意力,惩罚相邻区域token之间的交互,并抑制无关的跨实例注意力。因此,我们的方法在保持全局视觉一致性的同时实现了属性绑定。我们在LoMOE-Bench上评估MICE,并引入MICE-Bench,这是一个更具挑战性的基准测试,平均每张图像包含8.5个并发编辑。实验表明,我们的方法在视觉质量保持和忠实于编辑指令方面优于强基线方法和最近的竞争方法。
Wavelet-Optimized Pseudo-3D Accelerated Diffusion Model for Truncated Computed Laminography
中文标题:用于截断计算层析成像的小波优化伪3D加速扩散模型
作者:Genyuan Zhang, Junyao Wang, Chuandong Tan, Fenglin Liu, Yongning Zhou
Computed Laminography (CL) is a key technology for the nondestructive testing of large plate-shaped objects. However, field-of-view (FOV) limitations inevitably lead to truncation of projected data, an ill-posed inverse problem that causes severe reconstruction artifacts. Existing deep learning methods typically rely on 2D architectures that lack rigorous data consistency constraints. Furthermore, they conventionally confine artifact removal strictly to the FOV, discarding potentially recoverable information outside it. To overcome these limitations, we first introduce a comprehensive CL FOV analysis, categorizing the space into data-complete, data-incomplete, and data-free regions. By extending our reconstruction target to encompass the data-incomplete region, we significantly expand the effective imaging range and enhance scanning efficiency. To achieve this, we propose a novel wavelet-optimized pseudo-3D accelerated diffusion model for CL truncation reconstruction (CL-DM). Our method utilizes a standard 2D diffusion model for slice aggregation, combined with a 3D model-based iterative reconstruction (MBIR) method to ensure strict data consistency. To mitigate inter-slice discontinuities, we introduce wavelet regularization along the z-direction, paired with a translation-invariant (TI) mechanism and a low-frequency preservation strategy. Finally, we introduce a 3D fast sampling architecture, significantly accelerating inference speed. Extensive simulations and real-world experiments demonstrate that CL-DM is superior in effectively eliminating truncation artifacts and restoring high-fidelity, continuous 3D structures.
计算层析成像(Computed Laminography, CL)是大型板状物体无损检测的关键技术。然而,视场(Field-of-View, FOV)限制不可避免地导致投影数据截断,这是一个病态逆问题,会产生严重的重建伪影。现有深度学习方法通常依赖缺乏严格数据一致性约束的2D架构。此外,传统方法将伪影去除严格限制在FOV内,丢弃了外部可能可恢复的信息。为克服这些限制,我们首先引入全面的CL FOV分析,将空间划分为数据完整区、数据不完整区和无数据区。通过将重建目标扩展至数据不完整区,我们显著扩展了有效成像范围并提高了扫描效率。为此,我们提出了一种用于CL截断重建的新型小波优化伪3D加速扩散模型(CL-DM)。我们的方法采用标准2D扩散模型进行切片聚合,并结合基于3D模型的迭代重建(Model-Based Iterative Reconstruction, MBIR)方法以确保严格的数据一致性。为缓解切片间不连续性,我们引入z方向小波正则化,并与平移不变(Translation-Invariant, TI)机制和低频保留策略相结合。最后,我们引入了一种3D快速采样架构,显著加速了推理速度。大量模拟和真实实验表明,CL-DM在有效消除截断伪影和恢复高保真连续3D结构方面具有优越性。
Accelerated Likelihood Maximization for Diffusion-based Versatile Content Generation
中文标题:用于扩散模型通用内容生成的加速似然最大化方法
作者:Hyunsoo Lee, Inwoo Hwang, Young Min Kim
Generating diverse, coherent, and plausible content from partially given inputs remains a fundamental challenge for diffusion models. Existing approaches face clear limitations: training-based approaches offer strong task-specific results but require costly computation, and they generalize poorly across tasks. Training-free approaches offer better efficiency, but they do not explicitly optimize over unobserved variables, leading to globally inconsistent results. To address these limitations, we introduce Accelerated Likelihood Maximization (ALM), a novel training-free sampling strategy integrated into the reverse diffusion process that significantly extends the applicability of diffusion models beyond simple generation tasks. Unlike previous methods that implicitly influence missing regions through pre-generated region constraints, we directly optimize the unobserved region during the sampling process, enabling globally coherent and plausible generation. Furthermore, we incorporate an acceleration strategy that significantly improves computational efficiency without sacrificing performance. Experimental results demonstrate that ALM consistently outperforms state-of-the-art methods in various data domains and tasks, establishing a powerful paradigm for versatile content generation.
从部分给定的输入生成多样、一致且合理的内容仍然是扩散模型面临的基本挑战。现有方法存在明显的局限性:基于训练的方法能够提供强大的任务特定结果,但需要昂贵的计算成本,且在不同任务间的泛化能力较差。无训练方法虽然具有更高的效率,但并未对未观察变量进行显式优化,导致结果全局不一致。为解决这些局限性,我们提出了加速似然最大化(Accelerated Likelihood Maximization,ALM)方法,这是一种新型的无训练采样策略,集成于逆向扩散过程中,显著扩展了扩散模型在简单生成任务之外的适用性。与先前通过预生成区域约束隐式影响缺失区域的方法不同,我们在采样过程中直接优化未观察区域,从而实现全局一致且合理的内容生成。此外,我们引入了一种加速策略,在不牺牲性能的前提下显著提高计算效率。实验结果表明,ALM在各种数据域和任务中始终优于最先进的方法,为通用内容生成建立了强大的范式。
Bridging Video Understanding and Generation in a Unified Framework
中文标题:在统一框架中桥接视频理解与生成
作者:Yuqi Wang, Runyi Li, Ruoyu Feng, Renjie Chen, Wenfeng Lin, Mingyu Guo
Recently, unified image generation and understanding have been extensively explored. However, extending such unified modeling paradigms to the video domain remains largely underexplored. A central challenge is that video understanding favors compact, discriminative semantic representations, whereas video generation requires dense signals that preserve visual details and temporal coherence. Videos naturally capture both spatial semantics and temporal dynamics, making them a more suitable modality for unified multimodal modeling compared to static images. In this paper, we propose Vega, a unified framework that bridges video understanding and generation. Vega leverages a shared vocabulary to jointly model text and visual representations and employs a hybrid architecture combining autoregressive (AR) prediction with diffusion-based rendering. Specifically, the AR model focuses on predicting semantically meaningful visual tokens for keyframes, providing a structured representation that guides the diffusion module in rendering dense, high-resolution video frames. Extensive experiments demonstrate that Vega achieves strong performance on video generation benchmarks such as VBench and video understanding benchmarks like VideoMME.
最近,统一图像生成与理解已得到广泛探索。然而,将这种统一建模范式扩展到视频领域仍基本未被深入研究。一个核心挑战在于,视频理解倾向于紧凑的、具有判别性的语义表示,而视频生成则需要保留视觉细节和时间一致性的密集信号。与静态图像相比,视频自然地同时捕捉空间语义和时间动态,使其成为统一多模态建模的更合适模态。本文提出Vega,一个桥接视频理解与生成的统一框架。Vega利用共享词汇表联合建模文本和视觉表示,并采用结合自回归(AR)预测与基于扩散的渲染的混合架构。具体而言,AR模型专注于预测关键帧的语义有意义的视觉token,提供一种结构化表示,引导扩散模块渲染密集的高分辨率视频帧。大量实验表明,Vega在视频生成基准(如VBench)和视频理解基准(如VideoMME)上均取得了强劲性能。
No Prompt, No Leaks: A Robust Generative Steganography Framework via Prompt-Free Diffusion
中文标题:无提示,无泄露:一种基于无提示扩散的鲁棒生成式隐写术框架
作者:Jingwen Cai, Fen Xiao, Shuhua Deng, Xieping Gao
Generative image steganography synthesizes stego images directly from secret information to achieve inherent security advantages. Latent Diffusion Models (LDMs) have recently emerged as a fundamental image steganography framework that modulates secret latent representations with text prompts. Limited by the inflexibility of text prompts, these methods still struggle to generate high-quality stego images and accurately recover secret images. In this work, we propose a prompt-free diffusion image steganography framework that integrates style semantic priors to control more robust and reliable stego image generation. Specifically, a Cascaded Affine Coupling Module (CACM) establishes a bijective, deterministic mapping between a secret image and its latent representation. Then, style semantics are integrated into the diffusion process to control latent representation and ensure visual imperceptibility in the generated stego images. To mitigate trajectory deviations stemming from the unconditioned reverse process, a predictor-corrector mechanism is introduced to iteratively refine the generation trajectory via feedback from the current and predicted next states. Extensive experimental results show that the proposed method achieves competitive performance compared to state-of-the-art methods in terms of security, secret image reconstruction accuracy and controllability.
生成式图像隐写术通过直接从秘密信息合成隐写图像来实现固有安全优势。潜在扩散模型(Latent Diffusion Models,LDMs)作为一种基础图像隐写框架,近年来引起了广泛关注,它通过文本提示来调制秘密潜在表示。然而,受限于文本提示的局限性,这些方法在生成高质量隐写图像和准确恢复秘密图像方面仍面临挑战。本工作提出了一种无提示扩散图像隐写框架,该框架通过整合风格语义先验来控制更加鲁棒可靠的隐写图像生成。具体而言,我们设计了级联仿射耦合模块(Cascaded Affine Coupling Module, CACM)在秘密图像与其潜在表示之间建立双射确定性映射。随后,将风格语义整合到扩散过程中,以控制潜在表示并确保生成的隐写图像的视觉不可感知性。为缓解由无条件反向过程导致的轨迹偏差,我们引入了一种预测-校正机制,通过当前状态和预测下一状态的反馈来迭代优化生成轨迹。大量实验结果表明,所提方法在安全性、秘密图像重建准确性和可控性方面与现有最先进方法相比具有竞争力的性能。
Mesh BDF: Barycentric Dominance Field for 3D Native Mesh Generation
中文标题:Mesh BDF:用于3D原生网格生成的重心支配场
作者:Gaochao Song, Haohan Weng, Luo Zhang, Zibo Zhao, Shenghua Gao
Autoregressive (AR) modeling has recently achieved remarkable progress in native 3D mesh generation, largely due to its natural ability to handle variable-length, discrete data structures. However, the inherent constraints of the AR paradigm severely restrict the generated meshes, leading to limited face counts, bounded vertex resolutions, and difficulties in supporting textures. To overcome these bottlenecks, we propose the Barycentric Dominance Field (BDF), a continuous representation defined on triangular mesh surfaces that elegantly encodes vertex topological connectivity. BDF bridges the fundamental gap between discrete mesh topology and continuous diffusion-based generative modeling by transforming connectivity into a continuous surface signal. As an intrinsic mesh property, BDF shares strong similarities with texture maps, enabling its seamless integration into existing 3D diffusion pipelines without requiring architectural modifications. Extensive experiments demonstrate that BDF empowers diffusion models to generate native meshes with significantly higher quality, greater scalability, and stronger robustness compared to state-of-the-art autoregressive methods.
自回归(AR)建模在原生3D网格生成领域近年来取得了显著进展,这主要归功于其处理变长离散数据结构的天然能力。然而,自回归范式的固有约束严重限制了生成网格的质量,导致面数有限、顶点分辨率受限,以及纹理支持的困难。为克服这些瓶颈,我们提出了重心支配场(BDF),这是一种定义在三角网格表面上的连续表示方法,能够优雅地编码顶点拓扑连接性。BDF通过将连接性转换为连续表面信号,架起了离散网格拓扑与基于连续扩散的生成式建模之间的根本桥梁。作为一种内在网格属性,BDF与纹理映射具有很强的相似性,使其能够无缝集成到现有的3D扩散管道中,而无需对架构进行修改。大量实验表明,与最先进的自回归方法相比,BDF使扩散模型能够生成质量显著更高、可扩展性更强、鲁棒性更强的原生网格。
Cross-Space Distillation: Teaching One-Step Students with Modern Diffusion Teachers
中文标题:跨空间蒸馏:用现代扩散模型教师指导一步式学生模型
作者:Anh Nguyen, Ngan Nguyen, Duc Vu, Trung Dao, Viet Nguyen, Quan Dao, Kien Nguyen, Chi Tran, Phong Nguyen, Khoi Nguyen, Cuong Pham, Dimitris Metaxas, Vishal M. Patel, Anh Tran
Modern one-step diffusion models achieve impressive quality through distribution-based timestep distillation. Yet, they rely on a critical assumption: Teacher and Student must inhabit the same latent space. This Shared-Space constraint prevents knowledge transfer from modern high-capacity Teachers (e.g., SD 3.5 and Flux) into compact, deployment-friendly Students such as SD 1.5, whose latent resolution and VAE parameterization differ from the Teacher. We formalize this overlooked regime as Cross-Space Distillation, where Teacher and Student differ in both latent resolution and VAE space. To enable distillation under this mismatch, we introduce the Bridge, a lightweight latent interface that maps Student latents into the Teacher space without modifying the Student backbone. Bridge combines a frozen Student VAE decoder as a spatial prior with a compact learnable projector, and is trained with latent reconstruction and attention fidelity objectives for stable Teacher-space alignment. Across diverse modern Teachers, Bridge enables substantial gains for compact one-step Students; for example, it improves SD 1.5 from 5.4 to 9.4 HPSv3 while preserving one-step inference, low latency, and broad ecosystem compatibility. These results show that heterogeneous large Teachers can be distilled into efficient, deployable backbones through a lightweight latent-space interface.
现代一步式扩散模型通过基于分布的时间步蒸馏实现了令人印象深刻的质量提升。然而,它们依赖于一个关键假设:教师模型和学生模型必须处于同一潜在空间。这种共享空间约束阻碍了知识从现代高容量教师模型(如 SD 3.5 和 Flux)向紧凑、适合部署的学生模型(如 SD 1.5)的迁移,因为学生模型的潜在分辨率和 VAE 参数化与教师模型不同。我们将这一被忽视的范式正式定义为跨空间蒸馏,即教师模型和学生模型在潜在分辨率和 VAE 空间上均存在差异。为在这种不匹配条件下实现蒸馏,我们引入了 Bridge(桥接模块),这是一个轻量级的潜在空间接口,可将学生模型的潜在表示映射到教师空间,而无需修改学生模型的骨干网络。Bridge 将冻结的学生 VAE 解码器作为空间先验与一个紧凑的可学习投影器相结合,并通过潜在重建和注意力保真度目标进行训练,以实现稳定的教师空间对齐。在多种现代教师模型上,Bridge 为紧凑的一步式学生模型带来了显著提升;例如,它将 SD 1.5 的 HPSv3 分数从 5.4 提升至 9.4,同时保持一步推理、低延迟和广泛的生态兼容性。这些结果表明,异构大型教师模型可以通过轻量级的潜在空间接口蒸馏成高效、可部署的骨干网络。
SpheRoPE: Zero-Shot Optimization-Free 360 Panorama Generation with Spherical RoPE
中文标题:SpheRoPE:基于球面RoPE的零样本无优化360度全景生成
作者:Or Hirschorn, Aaron Olender, Eli Alshan, Ianir Ideses, Lior Fritz, Sagie Benaim
We present a zero-shot, training-free and optimization-free framework for generating 360 panoramic images and videos by directly injecting spherical priors into pre-trained diffusion transformers. Existing methods either rely on costly fine-tuning on scarce panoramic data that limits generalization, or leverage multi-step optimization that incurs prohibitive inference latency. We observe that contemporary generative models natively exhibit some panoramic priors from large-scale training. However, these emergent capabilities are insufficient, as the models fundamentally fail to satisfy the rigorous topological constraints imposed by equirectangular projection (ERP). We introduce a zero-shot and optimization-free approach that resolves these constraints at inference time. Spherical RoPE replaces standard rotary position embeddings: low-frequency channels are re-parameterized as 3D Cartesian coordinates to natively encode the spherical manifold, while high-frequency channels are harmonically quantized to enforce exact periodicity. Coupled with complementary Semantic Distortion classifier-free guidance (CFG) that explicitly steers geometry, we avoid retraining and inherit the full creative breadth of state-of-the-art models. Our approach generalizes across diverse backbones and 360 generation modalities. We demonstrate this across text-to-panorama using Flux.1, Flux.2, and LTX-Video backbones, achieving competitive performance against baselines, all while remaining training-free. Project page: https://orhir.github.io/SpheRoPE
本文提出了一种零样本、无训练、无优化的框架,通过将球面先验直接注入预训练的扩散变换器来生成360度全景图像和视频。现有方法要么依赖于在稀缺的全景数据上进行成本高昂的微调(这限制了泛化能力),要么采用多步优化(导致推理延迟过高)。我们观察到,当代生成模型在大规模训练中天然地表现出一些全景先验。然而,这些涌现能力仍然不足,因为模型从根本上无法满足等距矩形投影(ERP)施加的严格拓扑约束。我们引入了一种在推理时解决这些约束的零样本无优化方法。球面RoPE替换了标准的旋转位置编码:低频通道被重新参数化为3D笛卡尔坐标以原生编码球面流形,而高频通道则经过谐波量化以实现精确的周期性。结合互补的语义失真无分类器引导(CFG)来显式引导几何结构,我们无需重新训练即可继承最先进模型的全部创造性。我们的方法在不同骨干网络和360度生成模态上具有泛化能力。我们使用Flux.1、Flux.2和LTX-Video骨干网络在文本到全景任务上进行了演示,实现了与基线方法相当的性能,同时保持无训练的特性。项目页面:https://orhir.github.io/SpheRoPE
PolyFlow: Continuous Topology Embedding Flow Matching for Artist-style Mesh Generation
中文标题:PolyFlow:面向艺术风格网格生成的连续拓扑嵌入流匹配
作者:Chunshi Wang, Haohan Weng, Junliang Ye, Biwen Lei, Yang Li, Zibo Zhao, Zeqiang Lai, Kaiyi Zhang, Yunhan Yang, Zhuo Chen, Chunchao Guo, Yawei Luo
Autoregressive Transformers dominate high-quality mesh generation by producing artist-worthy topologies, yet their inherent sequential decoding induces substantial computational overhead, falling orders of magnitude slower than parallel generative models. On the other hand, while continuous diffusion and flow-matching methods support efficient parallel synthesis across a variety of domains, they cannot be directly applied to meshes: mesh connectivity is inherently discrete and incompatible with standard continuous noise injection and denoising operations. To resolve this fundamental incompatibility, we introduce a compact topology embedder that projects discrete mesh vertex positions and normals into continuous per-vertex embeddings, where the original discrete adjacency information can be faithfully recovered via spacetime distance thresholding. After pretraining and freezing this embedder, any raw mesh can be fully converted into a continuous per-vertex state space unifying position, normal, and implicit topological attributes. Built upon this novel continuous mesh representation, we present PolyFlow, a Transformer-based flow-matching framework that achieves fully parallel vertex state denoising conditioned on extracted point-cloud features. During inference, our model completes generation rapidly via an ODE solver, and supports explicit, precise control over output mesh resolution by directly specifying the target vertex count. Extensive evaluations on the Toys4K benchmark demonstrate that PolyFlow surpasses state-of-the-art autoregressive baselines in both Chamfer Distance and Hausdorff Distance.
自回归Transformer在生成艺术级拓扑结构的高质量网格方面占据主导地位,但其固有的顺序解码机制带来了巨大的计算开销,比并行生成模型慢数个数量级。另一方面,尽管连续扩散和流匹配方法在多个领域支持高效的并行合成,但它们无法直接应用于网格:网格连接性本质上是离散的,与标准的连续噪声注入和去噪操作不兼容。为解决这一根本不兼容性,我们引入了一个紧凑的拓扑嵌入器,将离散的网格顶点位置和法线投影到连续的逐顶点嵌入中,其中原始的离散邻接信息可通过时空距离阈值精确恢复。预训练并冻结该嵌入器后,任何原始网格都可以完全转换为统一的连续逐顶点状态空间,融合位置、法线和隐式拓扑属性。基于这一新颖的连续网格表示,我们提出了PolyFlow,一个基于Transformer的流匹配框架,在提取的点云特征条件下实现完全并行的顶点状态去噪。推理过程中,我们的模型通过ODE求解器快速完成生成,并支持通过直接指定目标顶点数来精确控制输出网格分辨率。在Toys4K基准上的广泛评估表明,PolyFlow在Chamfer距离和Hausdorff距离上均超越了最先进的自回归基线方法。
Off the Rails: Hijacking the Scoring Head in Generative End-to-End Driving Planners with Safety-Violating Adversarial Perturbations
中文标题:脱轨:利用安全性违反对抗扰动劫持生成式端到端驾驶规划器的评分头
作者:Halima Bouzidi, Mboutidem Ekemini Mkpong, Haoyu Liu, Mohammad Abdullah Al Faruque
Generative models have recently seen rapid adoption in End-to-End (E2E) autonomous driving (AD), with diffusion-based denoising and vocabulary-based retrieval becoming the dominant trajectory-decoding paradigms. Despite their architectural diversity, current generative AD planners share a common inference pattern: a fixed set of candidate trajectories (anchors, vocabulary entries, or proposal queries) is scored by one or more learned heads conditioned on the Bird's-Eye-View (BEV) features, and the highest-scored candidate is returned as the final trajectory. Under this design, the scoring head is the only barrier between perception and the motion command, and its decision margins between competing candidates are often small. We introduce \textsc{Derail}, an adversarial framework that exploits this scoring-head attack surface. Evaluated on various generative planners, \textsc{Derail} flips the trajectory selection from a safe to an unsafe candidate, with score drops of $39$--$80\%$ and collision rates of up to $50\%$, consistently outperforming generic loss-maximization and feature-divergence attacks. Our analysis suggests that safety-violating objectives govern attack effectiveness against generative AD planners, and that the scoring-head inference pattern itself is a recurring attack surface worth explicit defensive consideration.
生成模型近年来在端到端(E2E)自动驾驶(AD)中得到快速采用,扩散去噪和基于词汇的检索已成为主导的轨迹解码范式。尽管架构多样,当前生成式自动驾驶规划器共享一种常见的推理模式:一组固定的候选轨迹(锚点、词汇条目或提议查询)由一个或多个基于鸟瞰图(BEV)特征条件化的学习评分头进行评分,得分最高的候选轨迹被返回作为最终轨迹。在这种设计下,评分头是感知与运动指令之间的唯一屏障,而其竞争候选之间的决策边界通常很小。本文提出了Derail,一个利用这一评分头攻击面的对抗框架。在各种生成式规划器上的评估表明,Derail能够将轨迹选择从安全候选翻转为不安全候选,得分下降39%至80%,碰撞率高达50%,始终优于通用的损失最大化和特征分歧攻击。我们的分析表明,安全性违反目标对生成式自动驾驶规划器的攻击效果起主导作用,而评分头推理模式本身是一个值得明确防御考虑的反复出现攻击面。
Distortion-Corrected Diffusion MRI Using Rotated-View EPI and Joint Field-Map/Image Estimation with Gaussian Primitives
中文标题:基于旋转视角EPI和联合场图/图像估计(高斯基元)的畸变校正扩散磁共振成像
作者:Wenqi Huang, Zhitao Li, Nan Wang, Yimeng Lin, Mengze Gao, Yurui Qian, Sevgi Gokce Kafali, Xiaozhi Cao, Kawin Setsompop, Daniel Rueckert, Congyu Liao
Echo Planar Imaging (EPI) is the standard acquisition technique for diffusion and functional neuroimaging, enabling rapid imaging but suffering from geometric distortions caused by B0 field inhomogeneities. Existing correction methods first reconstruct distorted images using parallel imaging, then estimate the B0 field and correct the distortion in the image domain. In this sequential process, reconstruction artifacts at high acceleration factors and low SNR at high diffusion b-values degrade B0 estimation and limit the overall correction quality. We propose a physics-informed framework that jointly estimates the B0 field and distortion-free image directly from k-space data, without depending on an intermediate parallel-imaging reconstruction for the correction. The image and the B0 field are each represented as a superposition of Gaussian primitives embedded within an MRI physics forward model. The explicit, continuous parameterization captures both smooth regions and tissue boundaries and supports rotated-view EPI acquisitions without interpolation. The diffusion-weighted image is modeled as real and non-negative, with the image phase absorbed into a per-shot phase factor. Rotated views distribute distortions across multiple phase-encoding orientations, improving point spread function isotropy and providing stronger constraints for B0 estimation. On in vivo brain diffusion EPI, the proposed method attains the closest brain-boundary agreement with a distortion-free structural reference, with the largest improvement over sequential methods at high b-value and high acceleration. Extensive visual comparisons further show improved detail fidelity and noise suppression.
回波平面成像(EPI)是扩散和功能神经成像的标准采集技术,能够实现快速成像,但会受到由B0场不均匀性引起的几何畸变影响。现有的校正方法首先使用并行成像重建畸变图像,然后在图像域估计B0场并进行畸变校正。在这一顺序过程中,高加速因子下的重建伪影以及高扩散b值下的低信噪比会降低B0估计的准确性,并限制整体校正质量。我们提出了一个物理信息框架,直接从k空间数据联合估计B0场和无畸变图像,而不依赖于中间并行成像重建进行校正。图像和B0场均表示为嵌入MRI物理前向模型中的高斯基元叠加。显式、连续的参数化方法能够同时捕捉平滑区域和组织边界,并支持无插值的旋转视角EPI采集。扩散加权图像被建模为实数且非负,图像相位被吸收到每次激发相位因子中。旋转视角将畸变分散到多个相位编码方向,提高了点扩散函数的各向同性,并为B0估计提供了更强的约束条件。在活体脑部扩散EPI实验中,所提出的方法与无畸变结构参考达到了最接近的脑边界一致性,在高b值和高加速因子下相比顺序方法获得了最大的改进。广泛的视觉比较进一步显示了细节保真度和噪声抑制的改善。
PSHuman: Photorealistic Single-image 3D Human Reconstruction using Cross-Scale Multiview Diffusion and Explicit Remeshing
中文标题:PSHuman:基于跨尺度多视图扩散和显式重网格化的单张照片级3D人体重建
作者:Peng Li, Wangguandong Zheng, Yuan Liu, Tao Yu, Yangguang Li, Xingqun Qi, Xiaowei Chi, Siyu Xia, Yan-Pei Cao, Wei Xue, Wenhan Luo, Yike Guo
Detailed and photorealistic 3D human modeling is essential for various applications and has seen tremendous progress. However, full-body reconstruction from a monocular RGB image remains challenging due to the ill-posed nature of the problem and sophisticated clothing topology with self-occlusions. In this paper, we propose PSHuman, a novel framework that explicitly reconstructs human meshes utilizing priors from the multiview diffusion model. It is found that directly applying multiview diffusion on single-view human images leads to severe geometric distortions, especially on generated faces. To address it, we propose a cross-scale diffusion that models the joint probability distribution of global full-body shape and local facial characteristics, enabling detailed and identity-preserved novel-view generation without any geometric distortion. Moreover, to enhance cross-view body shape consistency of varied human poses, we condition the generative model on parametric models like SMPL-X, which provide body priors and prevent unnatural views inconsistent with human anatomy. Leveraging the generated multi-view normal and color images, we present SMPLX-initialized explicit human carving to recover realistic textured human meshes efficiently. Extensive experimental results and quantitative evaluations on CAPE and THuman2.1 datasets demonstrate PSHumans superiority in geometry details, texture fidelity, and generalization capability.
详细且逼真的3D人体建模对于各种应用至关重要,并已取得巨大进展。然而,由于该问题的病态性质以及具有自遮挡的复杂服装拓扑结构,从单张RGB图像进行全身重建仍然具有挑战性。本文提出PSHuman,一个利用多视图扩散模型先验进行显式人体网格重建的新框架。研究发现,直接将多视图扩散应用于单张人体图像会导致严重的几何畸变,尤其是在生成的人脸上。为解决这一问题,本文提出一种跨尺度扩散方法,对全身形状与局部面部特征的联合概率分布进行建模,从而在无需任何几何畸变的情况下实现细节丰富且保持身份的新视角生成。此外,为增强不同人体姿态下的跨视角身体形状一致性,本文将SMPL-X等参数化模型作为生成模型的先验条件,提供身体先验并防止生成与人体解剖学不一致的不自然视角。利用生成的多视图法线和彩色图像,本文提出基于SMPLX初始化的显式人体雕刻方法,以高效恢复逼真的带纹理人体网格。在CAPE和THuman2.1数据集上的广泛实验结果和定量评估表明,PSHuman在几何细节、纹理保真度和泛化能力方面具有优势。
Filterless Snapshot Hyperspectral Imaging using Guided Patch Diffusion
中文标题:基于引导补丁扩散的无滤镜快照高光谱成像
作者:Dean Hazineh, Luca Sacchi, Davide Cassara, Federico Capasso, Todd Zickler
We consider the problem of reconstructing a HxWx31 hyperspectral image from a $H\times W$ grayscale snapshot measurement that is captured using only a single diffractive lens and a filterless panchromatic photosensor. This problem is severely ill-posed, but we present a model that produces high-quality results in simulation and experiment. We make efficient use of limited training data by creating a conditional denoising diffusion model that operates on small patches in a shift-invariant manner. During inference, we synchronize per-patch hyperspectral predictions using guidance by physical consistency with the system's optical point spread function. Our experiments reveal that the patch size can be as small as the point spread function, with local optical cues being the main source of information about complete spectra. Also, by drawing multiple samples, our model provides per-pixel uncertainty estimates that strongly correlate with reconstruction error.
本文研究从使用单个衍射透镜和无滤镜全色光电传感器捕获的H×W灰度快照测量重建HxWx31高光谱图像的问题。该问题严重欠定,但本文提出的模型在仿真和实验中均能产生高质量结果。我们通过创建以移位不变方式在小型补丁上运行的条件去噪扩散模型,高效利用有限的训练数据。推理过程中,我们通过与系统光学点扩散函数的物理一致性引导,同步每个补丁的高光谱预测。实验表明,补丁尺寸可以小至点扩散函数大小,光学线索是获取完整光谱信息的主要来源。此外,通过多次采样,我们的模型可提供与重建误差高度相关的每像素不确定性估计。
Improved Immiscible Diffusion: Accelerate Diffusion Training by Reducing Its Miscibility
中文标题:改进的不混溶扩散:通过降低混溶性加速扩散模型训练
作者:Yiheng Li, Feng Liang, Dan Kondratyuk, Masayoshi Tomizuka, Kurt Keutzer, Chenfeng Xu
The substantial training cost of diffusion models hinders their deployment. Immiscible Diffusion recently showed that reducing diffusion trajectory mixing in the noise space via linear assignment accelerates training by simplifying denoising. To extend immiscible diffusion beyond the inefficient linear assignment under high batch sizes and high dimensions, we refine this concept to a broader miscibility reduction at any layer and by any implementation. Specifically, we empirically demonstrate the bijective nature of the denoising process with respect to immiscible diffusion, ensuring its preservation of generative diversity. Moreover, we provide thorough analysis and show step-by-step how immiscibility eases denoising and improves efficiency. Extending beyond linear assignment, we propose a family of implementations including K-nearest neighbor (KNN) noise selection and image scaling to reduce miscibility, achieving up to >4x faster training across diverse models and tasks including unconditional/conditional generation, image editing, and robotics planning. Furthermore, our analysis of immiscibility offers a novel perspective on how optimal transport (OT) enhances diffusion training. By identifying trajectory miscibility as a fundamental bottleneck, we believe this work establishes a potentially new direction for future research into high-efficiency diffusion training. The code is available at https://github.com/yhli123/Immiscible-Diffusion.
扩散模型的高昂训练成本阻碍了其部署应用。Immiscible Diffusion近期研究表明,通过线性分配减少噪声空间中的扩散轨迹混合可以简化去噪过程,从而加速训练。为将不混溶扩散推广到高批量大小和高维度下低效的线性分配方法,我们将该概念进一步提炼为更通用的任意层和任意实现方式的混溶性降低。具体而言,我们实证证明了去噪过程相对于不混溶扩散的双射性质,确保其对生成多样性的保持。此外,我们提供了详尽分析,逐步阐明不混溶性如何简化去噪并提升效率。在线性分配之外,我们提出了一系列实现方法,包括K近邻(KNN)噪声选择和图像缩放来降低混溶性,在无条件/条件生成、图像编辑和机器人规划等多种模型和任务上实现了最高超过4倍的训练加速。进一步地,我们对不混溶性的分析为最优传输(OT)如何增强扩散训练提供了新的视角。通过识别轨迹混溶性作为根本瓶颈,我们相信这项工作为未来高效扩散训练研究确立了潜在的新方向。代码已开源于https://github.com/yhli123/Immiscible-Diffusion。
ISAC: Training-Free Instance-to-Semantic Attention Control for Multi-Instance Generation
中文标题:ISAC:面向多实例生成的免训练实例到语义注意力控制
作者:Sanghyun Jo, Wooyeol Lee, Ziseok Lee, Jonghyun Choi, Jaesik Park, Kyungsu Kim
Recent open-weight text-to-image (T2I) diffusion models still struggle with multi-instance prompts, often omitting or merging instances and mixing semantics among similar objects. We trace these failures to early denoising steps, before instance boundaries are reliably stabilized. Existing training-free guidance is largely driven by cross-attention or other token-conditioned semantic signals. Such guidance can separate concepts at the token level, but largely assumes that distinct instance regions have already emerged. In early denoising steps, it cannot reliably carve out these regions, so count failures and semantic mixing persist. By contrast, self-attention exposes class-agnostic instance layouts during early denoising. To exploit this asymmetry, we propose $\textbf{ISAC}$ ($\textbf{I}$nstance-to-$\textbf{S}$emantic $\textbf{A}$ttention $\textbf{C}$ontrol), a training-free, model-agnostic objective that first stabilizes self-attention layouts and then binds cross-attention semantics within them, without fine-tuning or external vision models. Across T2I-CompBench, HRS-Bench, and our newly curated IntraCompBench, ISAC consistently outperforms prior training-free methods. Furthermore, ISAC enhances layout-to-image controllers by refining coarse, overlapping bounding boxes into dense instance masks. Code and IntraCompBench are available at https://shjo-april.github.io/ISAC.
当前开放权重的文本到图像(T2I)扩散模型在处理多实例提示时仍存在困难,常常遗漏或合并实例,并在相似对象间产生语义混淆。我们将这些失败追溯到早期去噪步骤,此时实例边界尚未可靠稳定。现有的免训练引导方法主要由交叉注意力或其他token条件的语义信号驱动。此类引导可以在token层面分离概念,但很大程度上假设不同的实例区域已经出现。在早期去噪步骤中,它无法可靠地分割这些区域,因此计数失败和语义混淆问题依然存在。相比之下,自注意力在早期去噪过程中揭示了与类别无关的实例布局。为利用这种不对称性,我们提出了ISAC(Instance-to-Semantic Attention Control,免训练的实例到语义注意力控制),这是一种模型无关的目标函数,首先稳定自注意力布局,然后将交叉注意力语义绑定到其中,无需微调或外部视觉模型。在T2I-CompBench、HRS-Bench以及我们新构建的IntraCompBench上,ISAC始终优于现有的免训练方法。此外,ISAC还能通过将粗糙、重叠的边界框精炼为密集的实例掩码来增强布局到图像的控制能力。代码和IntraCompBench可访问 https://shjo-april.github.io/ISAC。
Few to Big: Prototype Expansion Network via Diffusion Learner for Point Cloud Few-shot Semantic Segmentation
中文标题:从少到多:基于扩散学习器的点云小样本语义分割原型扩展网络
作者:Qianguang Zhao, Dongli Wang, Yan Zhou, Jianxun Li, Richard Irampa
Few-shot 3D point cloud semantic segmentation aims to segment novel categories using a minimal number of annotated support samples. However, prototypes derived from the limited non-structural point cloud support set are often misaligned and have a small capacity, hindering effective gen eralization to novel categories. This stems from two core issues: i) the prototype possess limited representational capacity fails to cover the full intra-class diversity of a novel category, and ii) the prototypes suffer from misalignment with the query space due to the inter-set inconsistency between support and query sets. To address these issues, our work focuses on leveraging the few support samples to construct a well-aligned big-capacity prototype. Motivated by the powerful generative capabilities of diffusion models, we re-purpose its pre-trained conditional encoder to provide rich feature components for prototype ex pansion. Subsequently, a push-pull force aligns this expanded prototype towards the query feature space. Under this setup, we introduce the Prototype Expansion Network (PENet), a framework that constructs aligned big-capacity prototypes from two complementary feature sources. Specifically, PENet employs a dual-stream learner architecture: it retains a conventional fully supervised Intrinsic Learner (IL) to distill representative features, while introducing a novel Diffusion Learner (DL) to provide rich generalizable features. The resulting dual prototypes are then processed by a Prototype Assimilation Module (PAM), which adopts a push-pull attention block to align the prototypes with the query space. Furthermore, a Prototype Calibration Mechanism (PCM) regularizes the final big-capacity prototype to prevent semantic drift. Extensive experiments on the S3DIS and ScanNet datasets demonstrate that PENet outperforms state-of-the-art methods across various few-shot settings.
小样本3D点云语义分割旨在使用极少的带标注支持样本来分割新类别。然而,从有限的无结构点云支持集推导出的原型往往存在未对齐问题且容量较小,阻碍了向新类别的有效泛化。该问题源于两个核心挑战:i)原型表示能力有限,无法覆盖新类别的完整类内多样性;ii)由于支持集与查询集之间的集间不一致性,原型与查询空间存在未对齐问题。为解决这些问题,本工作致力于利用少量支持样本构建对齐的大容量原型。得益于扩散模型强大的生成能力,我们重新利用其预训练的条件编码器为原型扩展提供丰富的特征分量。随后,推拉力机制将该扩展后的原型对齐到查询特征空间。在此框架下,我们提出了原型扩展网络(PENet),该框架从两个互补的特征源构建对齐的大容量原型。具体而言,PENet采用双流学习器架构:保留传统的全监督内在学习器(IL)来提取代表性特征,同时引入新型扩散学习器(DL)提供丰富的可泛化特征。所得的双重原型随后由原型同化模块(PAM)处理,该模块采用推拉注意力块将原型与查询空间对齐。此外,原型校准机制(PCM)对最终的大容量原型进行正则化以防止语义漂移。在S3DIS和ScanNet数据集上的大量实验表明,PENet在各种小样本设置下均优于现有最先进方法。
FeRA: Frequency-Energy Constrained Routing for Effective Diffusion Adaptation Fine-Tuning
中文标题:FeRA:面向有效扩散适应微调的频率能量约束路由
作者:Bo Yin, Xiaobin Hu, Xingyu Zhou, Yu He, Peng-Tao Jiang, Yue Liao, Junwei Zhu, Jiangning Zhang, Ying Tai, Shuicheng Yan
Diffusion models have achieved remarkable success in generative modeling, yet how to effectively adapt large pretrained models to new tasks remains challenging. We revisit the reconstruction behavior of diffusion models during denoising to unveil the underlying frequency energy mechanism governing this process. Building upon this observation, we propose FeRA, a frequency driven fine tuning framework that aligns parameter updates with the intrinsic frequency energy progression of diffusion. FeRA establishes a comprehensive frequency energy framework for effective diffusion adaptation fine tuning, comprising three synergistic components: (i) a compact frequency energy indicator that characterizes the latent bandwise energy distribution, (ii) a soft frequency router that adaptively fuses multiple frequency specific adapter experts, and (iii) a frequency energy consistency regularization that stabilizes diffusion optimization and ensures coherent adaptation across bands. Routing operates in both training and inference, with inference time routing dynamically determined by the latent frequency energy. It integrates seamlessly with adapter based tuning schemes and generalizes well across diffusion backbones and resolutions. By aligning adaptation with the frequency energy mechanism, FeRA provides a simple, stable, and compatible paradigm for effective and robust diffusion model adaptation.
扩散模型在生成建模中取得了显著成功,然而如何有效地将大型预训练模型适应新任务仍具挑战性。本研究重新审视扩散模型在去噪过程中的重建行为,以揭示支配该过程的底层频率能量机制。基于这一观察,本研究提出FeRA,一个参数更新与扩散模型内在频率能量演进相一致的频率驱动微调框架。FeRA建立了一个面向有效扩散适应微调的全面频率能量框架,包含三个协同组件:(i)一个紧凑的频率能量指示器,用于表征潜在空间的频段能量分布;(ii)一个软频率路由器,用于自适应融合多个频段专属的适配器专家;(iii)一个频率能量一致性正则化器,用于稳定扩散优化并确保跨频段的一致性适应。路由机制在训练和推理阶段均可运行,其中推理时路由由潜在频率能量动态决定。该方法与基于适配器的微调方案无缝集成,并可很好地推广到不同的扩散骨干网络和分辨率。通过将适应过程与频率能量机制对齐,FeRA为实现有效且稳健的扩散模型适应提供了一个简单、稳定且兼容的范式。
Phase-Aligned RoPE for Mixed-Resolution Diffusion Transformer
中文标题:混合分辨率扩散变换器的相位对齐RoPE
作者:Haoyu Wu, Jingyi Xu, Qiaomu Miao, Dimitris Samaras, Hieu Le
Rotary positional embeddings (RoPE) are widely used in diffusion transformers (DiTs) to encode spatial relationships, yet their behavior with mixed-resolution tokens remains underexplored. A natural approach is to rescale token positions from different resolutions into a unified coordinate system before attention, but we show this fails. Our analysis shows that with RoPE, the attention similarity score is a highly structured and periodic function of token distance, so rescaling distances across resolutions moves token pairs to different regions of this periodic function, leading to incorrect attention scores. Motivated by this, we introduce Phase-Aligned Mixed-Resolution Attention (PMA), a training-free mechanism that stabilizes mixed-resolution attention. PMA modifies the RoPE position mapping to enforce a consistent positional scale for every query-key pair, ensuring that relative distances are evaluated under a single reference scale. To further improve local coherence near resolution transitions, we incorporate a lightweight boundary refinement module that softly exchanges features across adjacent scales. Experiments on image and video diffusion models validate our analysis and demonstrate consistent improvements in visual fidelity and computational efficiency.
旋转位置嵌入(RoPE)广泛应用于扩散变换器(DiTs)中以编码空间关系,然而其在混合分辨率令牌下的行为仍缺乏深入研究。一种自然的方法是在注意力计算前将不同分辨率的令牌位置重新缩放到统一坐标系中,但我们发现此方法效果不佳。我们的分析表明,使用RoPE时,注意力相似度分数是令牌距离的高度结构化且周期性的函数,因此跨分辨率重新缩放距离会将令牌对映射到该周期性函数的不同区域,从而导致注意力分数计算错误。基于此,我们提出了相位对齐混合分辨率注意力(PMA),这是一种无需训练的机制,用于稳定混合分辨率注意力。PMA修改了RoPE的位置映射,为每个查询-键对强制执行一致的位置尺度,确保相对距离在单一参考尺度下进行评估。为进一步提升分辨率转换处的局部一致性,我们引入了一个轻量级的边界细化模块,用于在相邻尺度之间进行软特征交换。在图像和视频扩散模型上的实验验证了我们的分析,并展示了视觉保真度和计算效率的一致性提升。
PA-VAD: Diffusion-Based Pseudo-Only Video Anomaly Detection via Domain-Aligned Memory Updates
中文标题:PA-VAD:基于扩散模型的纯伪视频异常检测方法及其域对齐记忆更新策略
作者:Satoshi Hashimoto, Yanan Wang, Hitoshi Nishimura, Mori Kurokawa
Deploying video anomaly detection (VAD) in the real world is often constrained by the scarcity, privacy, and cost of collecting real abnormal footage. We propose PA-VAD, a novel pseudo-only framework that trains an anomaly detector without using any real abnormal videos, by pairing real normal videos with diffusion-synthesized pseudo-abnormal videos generated from a small set of real normal images. Beyond proposing a generation-driven training pipeline, we make a key empirical discovery: pseudo anomalies exhibit a characteristic spatiotemporal magnitude bias in feature space, which can dominate Multiple Instance Learning and degrade generalization if left unaddressed. To counter this pseudo-induced bias, we introduce the Domain-Aligned Regularized Module (DARM), which combines domain alignment with usage-aware memory updates to balance prototype coverage and stabilize optimization under biased pseudo supervision. Extensive experiments demonstrate that PA-VAD achieves 98.2% AUC on ShanghaiTech, 82.5% on UCF-Crime, and 95.1% on XD-Violence, and further improves generalization to unseen anomaly classes in open-set evaluations. Notably, PA-VAD surpasses the best real-abnormal WVAD baselines on ShanghaiTech and XD-Violence by +0.6% and +0.9%, respectively, and improves over the UVAD state of the art on UCF-Crime by +1.9% -showing that high-accuracy VAD is attainable without collecting real abnormal videos.
在现实世界中部署视频异常检测(VAD)系统往往受到真实异常 footage 稀缺性、隐私问题及采集成本等因素的限制。我们提出了 PA-VAD,这是一种创新的纯伪样本框架,无需使用任何真实异常视频即可训练异常检测器。该方法通过将真实正常视频与基于少量真实正常图像由扩散模型生成的伪异常视频进行配对来实现。除了提出这种生成驱动的训练流程外,我们还发现了一个关键的实证问题:伪异常在特征空间中呈现出特征性的时空幅度偏差,如果不加以处理,该偏差会主导多示例学习(MIL)过程并导致泛化能力下降。为应对这种伪样本引入的偏差,我们引入了域对齐正则化模块(DARM),该模块结合域对齐与感知使用的记忆更新策略,以平衡原型覆盖并在有偏的伪样本监督下稳定优化。大量实验表明,PA-VAD 在 ShanghaiTech 数据集上达到 98.2% AUC,在 UCF-Crime 数据集上达到 82.5% AUC,在 XD-Violence 数据集上达到 95.1% AUC,并在开放集评估中进一步提升了 对 unseen 异常类别的泛化能力。值得注意的是,PA-VAD 在 ShanghaiTech 和 XD-Violence 数据集上分别超越了最佳真实异常 WVAD 基线方法 +0.6% 和 +0.9%,并在 UCF-Crime 数据集上比 UVAD 最优方法提升了 +1.9%——这表明在不采集真实异常视频的情况下也可实现高精度的 VAD。
Bridging Information Asymmetry: A Hierarchical Framework for Deterministic Blind Face Restoration
中文标题:弥合信息不对称:面向确定性盲人脸修复的分层框架
作者:Zhengjian Yao, Jiakui Hu, Kaiwen Li, Hangzhou He, Xinliang Zhang, Shuang Zeng, Lei Zhu, Yanye Lu
Blind face restoration remains a persistent challenge due to the inherent ill-posedness of reconstructing holistic structures from severely constrained observations. Current generative paradigms, while capable of synthesizing realistic facial details, remain limited by the under-constrained nature of blind restoration, where severely degraded inputs can be mapped to plausible yet identity-inconsistent outputs. To address this issue, we present Pref-Restore, a hierarchical framework for deterministic BFR. Our design is organized around three complementary principles: (1) Semantic Information Augmentation, where an auto-regressive semantic branch converts image and text cues into structured tokens that provide a stable high-level anchor; (2) Texture-level Fidelity Alignment, where the diffusion generator is trained under this anchor to recover identity-relevant details; and (3) Fidelity-constrained Preference Optimization, where a face-aware reward refines the diffusion trajectory while controlling the quality-fidelity trade-off. Extensive experiments on synthetic and real-world benchmarks show that Pref-Restore achieves state-of-the-art performance, with stronger identity-sensitive fidelity and lower restoration uncertainty across repeated sampling. Systematic ablations further attribute these gains to the proposed hierarchical design, showing the necessity of staged training, the robustness of the text pathway under deployment-faithful conditions, and the benefit of fidelity-constrained preference optimization.
盲人脸修复由于从严重受限的观测中重建整体结构所固有的不适定性,一直是一个持续性挑战。当前生成范式虽能合成逼真的人脸细节,但仍受限于盲修复的欠约束性质——严重退化的输入可能被映射为看似合理但身份不一致的输出。为解决这一问题,我们提出了Pref-Restore,一个面向确定性BFR的分层框架。我们的设计围绕三个互补原则展开:(1)语义信息增强,其中自回归语义分支将图像和文本线索转换为提供稳定高层锚点的结构化标记;(2)纹理级保真度对齐,其中扩散生成器在此锚点下进行训练以恢复身份相关细节;(3)保真度约束的偏好优化,其中人脸感知奖励在控制质量-保真度权衡的同时优化扩散轨迹。在合成和真实世界基准上的大量实验表明,Pref-Restore实现了最先进的性能,在重复采样中具有更强的身份敏感保真度和更低的修复不确定性。系统性的消融实验进一步将性能提升归因于所提出的分层设计,证明了分阶段训练的必要性、文本路径在部署忠实条件下的鲁棒性,以及保真度约束偏好优化的优势。
VS3R: Robust Full-frame Video Stabilization via Deep 3D Reconstruction
中文标题:VS3R: 基于深度3D重建的稳健全帧视频稳定化方法
作者:Muhua Zhu, Xinhao Jin, Xinping Wang, Yu Zhang, Yifei Xue, Tie Ji, Yizhen Lao
Video stabilization aims to mitigate camera shake but faces a fundamental trade-off between geometric robustness and full-frame consistency. While 2D methods suffer from aggressive cropping, 3D techniques are often undermined by fragile optimization pipelines that fail under extreme motions. Novel view synthesis models suffer from structural artifacts and scale blindness. To bridge this gap, we propose VS3R, a framework that synergizes feed-forward 3D reconstruction with generative video diffusion. Our pipeline jointly estimates camera parameters, depth, and masks to ensure all-scenario reliability, and introduces a Hybrid Stabilized Rendering (HSR) module that fuses semantic and geometric cues to preliminarily address parallax occlusions caused by pose transformations while maintaining dynamic-static consistency. Finally, a Video Stabilization-Driven Diffusion Model (VSDM) leverages contextual information to restore disoccluded regions, jointly optimizing texture and temporal consistency. Collectively, VS3R achieves high-fidelity, full-frame stabilization across diverse camera models and significantly outperforms state-of-the-art methods in robustness and visual quality.
视频稳定化旨在减轻相机抖动,但在几何稳健性与全帧一致性之间存在根本性权衡。2D方法遭受激进裁剪之困,3D技术则常因脆弱的优化流程而在极端运动下失效,新视图合成模型存在结构伪影和尺度盲问题。为弥补这一差距,我们提出VS3R框架,将前馈3D重建与生成式视频扩散协同融合。我们的流程联合估计相机参数、深度和掩码以确保全场景可靠性,并引入混合稳定渲染(HSR)模块,该模块融合语义和几何线索,初步解决姿态变换引起的视差遮挡问题,同时保持动态-静态一致性。最后,视频稳定驱动扩散模型(VSDM)利用上下文信息恢复遮挡暴露区域,联合优化纹理和时间一致性。总体而言,VS3R在多样化相机模型上实现了高保真全帧稳定化,并在稳健性和视觉质量方面显著优于现有最先进方法。
Cross-Resolution Distribution Matching for Diffusion Distillation
中文标题:用于扩散蒸馏的跨分辨率分布匹配
作者:Feiyang Chen, Hongpeng Pan, Haonan Xu, Xinyu Duan, Zhefeng Wang, Yang Yang
Diffusion distillation is central to accelerating image and video generation, yet existing methods are fundamentally limited by the denoising process, where step reduction has largely saturated. Partial timestep low-resolution generation can further accelerate inference, but it suffers noticeable quality degradation due to cross-resolution distribution gaps. We propose Cross-Resolution Distribution Matching Distillation (RMD), a novel distillation framework that bridges cross-resolution distribution gaps for high-fidelity, few-step multi-resolution cascaded inference. Specifically, RMD divides the timestep intervals for each resolution using logarithmic signal-to-noise ratio (logSNR) curves, and introduces logSNR-based mapping to compensate for resolution-induced shifts. Distribution matching is conducted along resolution trajectories to reduce the gap between low-resolution generator distributions and the teacher's high-resolution distribution. In addition, a predicted-noise re-injection mechanism is incorporated during upsampling to stabilize training and improve synthesis quality. Quantitative and qualitative results show that RMD preserves high-fidelity generation while accelerating inference across various backbones. Notably, RMD achieves up to 33.4X speedup on SDXL and 25.6X on Wan2.1-14B, while preserving high visual fidelity.
扩散蒸馏是加速图像和视频生成的核心技术,但现有方法在去噪过程方面存在根本性限制,时间步的减少已基本达到饱和。部分时间步低分辨率生成可进一步加速推理,但因跨分辨率分布差距而出现明显的质量下降。本文提出跨分辨率分布匹配蒸馏(RMD),这是一种新型蒸馏框架,能够弥合跨分辨率分布差距,实现高保真、少步的多分辨率级联推理。具体而言,RMD使用对数信噪比(logSNR)曲线划分各分辨率的时间步间隔,并引入基于logSNR的映射来补偿分辨率引起的变化。沿分辨率轨迹进行分布匹配,以缩小低分辨率生成器分布与教师高分辨率分布之间的差距。此外,在上采样过程中引入预测噪声重注入机制,以稳定训练并提升合成质量。定量和定性结果表明,RMD在加速各种骨干网络推理的同时保持了高保真生成能力。值得注意的是,RMD在SDXL上实现了高达33.4倍的加速,在Wan2.1-14B上实现了25.6倍的加速,同时保持了较高的视觉保真度。
GeoNVS: Geometry Grounded Video Diffusion for Novel View Synthesis
中文标题:GeoNVS:基于几何的新视角合成视频扩散模型
作者:Minjun Kang, Inkyu Shin, Taeyeop Lee, Myungchul Kim, In So Kweon, Kuk-Jin Yoon
Novel view synthesis requires strong 3D geometric consistency and the ability to generate visually coherent images across diverse viewpoints. While recent camera-controlled video diffusion models show promising results, they often suffer from geometric distortions and limited camera controllability. To overcome these challenges, we introduce GeoNVS, a geometry-grounded novel-view synthesizer that enhances both geometric fidelity and camera controllability through explicit 3D geometric guidance. Our key innovation is the Gaussian Splat Feature Adapter (GS-Adapter), which lifts input-view diffusion features into 3D Gaussian representations, renders geometry-constrained novel-view features, and adaptively fuses them with diffusion features to correct geometrically inconsistent representations. Unlike prior methods that inject geometry at the input level, GS-Adapter operates in feature space, avoiding view-dependent color noise that degrades structural consistency. Its plug-and-play design enables zero-shot compatibility with diverse feed-forward geometry models without additional training, and can be adapted to other video diffusion backbones. Experiments across 9 scenes and 18 settings demonstrate state-of-the-art performance, achieving 11.3% and 14.9% improvements over SEVA and CameraCtrl, with up to 2x reduction in translation error and 7x in Chamfer Distance.
新视角合成需要强大的3D几何一致性以及在多样化视角下生成视觉连贯图像的能力。尽管近期基于相机控制的视频扩散模型展现出良好效果,但它们通常面临几何畸变和相机可控性有限的问题。为克服这些挑战,我们提出了GeoNVS,一个基于几何的新视角合成器,通过显式的3D几何指导来增强几何保真度和相机可控性。我们的关键创新是Gaussian Splat Feature Adapter(GS-Adapter),它将输入视角的扩散特征提升到3D高斯表示,渲染几何约束的新视角特征,并自适应地将其与扩散特征融合以纠正几何不一致的表示。与先前在输入层面注入几何的方法不同,GS-Adapter在特征空间中运行,避免了损害结构一致性的视角依赖颜色噪声。其即插即用设计实现了与各种前馈几何模型的零样本兼容性,无需额外训练,并可适配其他视频扩散backbone。在9个场景和18个设置下的实验表明达到了最先进性能,相比SEVA和CameraCtrl分别实现了11.3%和14.9%的提升,平移误差降低至多2倍,Chamfer Distance降低至多7倍。
VIGOR: VIdeo Geometry-Oriented Reward for Temporal Generative Alignment
中文标题:VIGOR:面向视频几何的时间生成对齐奖励
作者:Tengjiao Yin, Jinglei Shi, Heng Guo, Xi Wang
Video diffusion models lack explicit geometric supervision during training, leading to inconsistency artifacts such as object deformation, spatial drift, and depth violations in generated videos. To address this limitation, we propose a geometry-based reward model that leverages pretrained geometric foundation models to evaluate multi-view consistency through cross-frame reprojection error. Unlike previous geometric metrics that measure inconsistency in pixel space, where pixel intensity may introduce additional noise, our approach conducts error computation in a pointwise fashion, yielding a more physically grounded and robust error metric. Furthermore, we introduce a geometry-aware sampling strategy that filters out low-texture and non-semantic regions, focusing evaluation on geometrically meaningful areas with reliable correspondences to improve robustness. We apply this reward model to align video diffusion models through two complementary pathways: post-training of a bidirectional model via SFT or Reinforcement Learning and inference-time optimization of a Causal Video Model (e.g., Streaming video generator) via test-time scaling with our reward as a path verifier. Experimental results validate the effectiveness of our design, demonstrating that our geometry-based reward provides superior robustness compared to other variants. By enabling efficient inference-time scaling, our method offers a practical solution for enhancing open-source video models without requiring extensive computational resources for retraining.
视频扩散模型在训练过程中缺乏显式的几何监督,导致生成的视频出现不一致的伪影,如物体变形、空间漂移和深度违规等问题。为解决这一局限性,我们提出了一种基于几何的奖励模型,该模型利用预训练的几何基础模型,通过跨帧重投影误差来评估多视角一致性。与先前在像素空间测量不一致性的几何指标不同(像素强度可能引入额外噪声),我们的方法以逐点方式进行误差计算,产生了更物理合理、更鲁棒的误差指标。此外,我们引入了一种几何感知的采样策略,通过过滤低纹理和非语义区域,将评估重点聚焦于具有可靠对应关系的几何有意义区域,以提高鲁棒性。我们通过两种互补途径将该奖励模型应用于视频扩散模型的对齐:一是采用监督微调或强化学习方法对双向模型进行后训练,二是对因果视频模型(如流式视频生成器)进行推理时缩放优化,将我们的奖励作为路径验证器。实验结果验证了我们设计的有效性,表明基于几何的奖励相比其他变体具有更优的鲁棒性。通过实现高效的推理时缩放,我们的方法为增强开源视频模型提供了一种实用的解决方案,无需耗费大量计算资源进行重新训练。
Physically Grounded 3D Generative Reconstruction under Hand Occlusion using Proprioception and Multi-Contact Touch
中文标题:基于本体感知和多接触触觉的手部遮挡下物理接地3D生成式重建
作者:Gabriele Mario Caddeo, Pasquale Marra, Lorenzo Natale
We propose a multimodal, physically grounded approach for metric-scale amodal object reconstruction and pose estimation under severe hand occlusion. Unlike prior occlusion-aware 3D generation methods that rely only on vision, we leverage physical interaction signals: proprioception provides the posed hand geometry, and multi-contact touch constrains where the object surface must lie, reducing ambiguity in occluded regions. We represent object structure as a pose-aware, camera-aligned signed distance field (SDF) and learn a compact latent space with a Structure-VAE. In this latent space, we train a conditional flow-matching diffusion model, pretraining on vision-only images and finetuning on occluded manipulation scenes while conditioning on visible RGB evidence, occluder/visibility masks, the hand latent representation, and tactile information. Crucially, we incorporate physics-based objectives and differentiable decoder-guidance during finetuning and inference to reduce hand--object interpenetration and to align the reconstructed surface with contact observations. Because our method produces a metric, physically consistent structure estimate, it integrates naturally into existing two-stage reconstruction pipelines, where a downstream module refines geometry and predicts appearance. Experiments in simulation show that adding proprioception and touch substantially improves completion under occlusion and yields physically plausible reconstructions at correct real-world scale compared to vision-only baselines; we further validate transfer by deploying the model on a real humanoid robot with an end-effector different from those used during training.
我们提出了一种多模态、物理接地的严重手部遮挡下度量尺度非模态物体重建与姿态估计方法。与以往仅依赖视觉的遮挡感知3D生成方法不同,我们利用物理交互信号:本体感知提供手部几何姿态,多接触触觉约束物体表面位置,从而减少遮挡区域的歧义。我们将物体结构表示为姿态感知的相机对齐有符号距离场(SDF),并通过结构变分自编码器(Structure-VAE)学习紧凑潜在空间。在该潜在空间中,我们训练了一个条件流匹配扩散模型,首先在纯视觉图像上进行预训练,然后在手部遮挡操作场景上进行微调,同时以可见RGB证据、遮挡物/可见性掩码、手部潜在表示和触觉信息为条件。关键在于,我们在微调和推理阶段融入了基于物理的目标函数和可微解码器引导,以减少手-物体穿透并使重建表面与接触观测对齐。由于我们的方法产生度量、物理一致的结构估计,它可以自然集成到现有的两阶段重建流程中,其中下游模块细化几何并预测外观。模拟实验表明,与纯视觉基线相比,添加本体感知和触觉显著提升了遮挡下的完整重建性能,并在真实世界尺度下产生物理可信的重建结果;我们进一步通过在训练时未使用过的末端执行器的真实人形机器人上部署模型,验证了模型的迁移能力。
Stable and Near-Reversible Diffusion ODE Solvers for Image Editing
中文标题:用于图像编辑的稳定且近似可逆的扩散ODE求解器
作者:Barbora Barancikova, Daniil Shmelev, Cristopher Salvi
The inversion of diffusion models plays a central role in image editing. Algebraically reversible ODE solvers provide an appealing approach to diffusion inversion for text-guided image editing, by eliminating the inversion error inherent in DDIM-based editing pipelines. However, empirical results indicate that reversibility alone is insufficient. As edits require larger semantic or visual changes, reversible diffusion solvers often exhibit instabilities and suffer sharp drops in output quality. In this paper, we show that the trade-off between exact reversibility and numerical stability manifests empirically as a trade-off between background preservation and prompt alignment in image editing. We then investigate the use of near-reversible Runge-Kutta methods as a more stable alternative to exactly reversible diffusion schemes. When combined with a vector-field smoothing strategy, the resulting approach improves edit fidelity, remains stable under large edits, and largely retains the background-preservation benefits of reversible solvers.
扩散模型的逆过程在图像编辑中起着核心作用。代数可逆的ODE求解器为文本引导图像编辑的扩散逆过程提供了一种吸引人的方法,它消除了基于DDIM编辑流程中固有的逆过程误差。然而,实践经验表明,仅有可逆性是不够的。当编辑需要更大的语义或视觉变化时,可逆扩散求解器经常出现不稳定现象,并遭受输出质量的急剧下降。本研究表明,精确可逆性与数值稳定性之间的权衡在经验上表现为图像编辑中背景保持与提示对齐之间的权衡。随后,我们研究了使用近似可逆的龙格-库塔方法作为精确可逆扩散方案的更稳定替代方案。当与向量场平滑策略相结合时,该方法提高了编辑保真度,在大规模编辑下保持稳定,并在很大程度上保留了可逆求解器的背景保持优势。
InteractiveAvatar: Real-Time Streaming Video Generation for Consistent and Intent-Aware Avatars
中文标题:InteractiveAvatar:用于一致性和意图感知虚拟形象的实时流视频生成
作者:Quanyue Song, Yishan He, Yanfei Zhang, Shihao Cheng, Zhixiang He, Zhizhi Guo, Chi Zhang, Xuelong Li, Caigui Jiang
Recent diffusion-based models have enabled realistic audio-driven avatar generation in real-time streaming. However, existing approaches struggle to maintain visual temporal consistency and fail to explicitly perceive user intent in complex interactive streaming scenarios. To address these challenges, we propose InteractiveAvatar, a real-time infinite-streaming video generation framework that supports visually consistent avatar video generation and intent-aware interactions. With autoregressive distillation, InteractiveAvatar achieves real-time str-eaming generation of human avatars over arbitrarily long durations. For visual consistency, we introduce a Long-Short Visual Memory (LSVM) mechanism that flexibly compresses historical visual information into compact tokens, preserving both short-range coherence and long-term consistency. To generate avatars with speeches and actions aligned with user intent, we propose a Reasoning-Reaction Module (RRM), which incorporates a State-Cycling strategy and a Cache-Switching mechanism. Extensive experimental results over diverse scenarios demonstrate that our method achieves state-of-the-art visual consistency in long-duration generation, while enabling complex user-avatar interaction in real time.
近年来,基于扩散的模型已能够实现实时流式传输下的逼真音频驱动虚拟形象生成。然而,现有方法在复杂的交互式流媒体场景中难以保持视觉时间一致性,且无法明确感知用户意图。为解决这些挑战,我们提出了 InteractiveAvatar,一个实时无限流视频生成框架,支持视觉一致的虚拟形象视频生成和意图感知交互。通过自回归蒸馏,InteractiveAvatar 可在任意时长下实现人类虚拟形象的实时流式生成。为实现视觉一致性,我们引入了长短期视觉记忆(LSVM)机制,将历史视觉信息灵活压缩为紧凑的令牌,同时保留短期一致性和长期一致性。为了生成与用户意图对齐的语音和动作虚拟形象,我们提出了推理-反应模块(RRM),该模块整合了状态循环策略和缓存切换机制。在多种场景下的大量实验结果表明,我们的方法在长时间生成中实现了最先进的视觉一致性,同时能够支持实时的复杂用户-虚拟形象交互。
Occlusion-Robust Multi-Object Decoupling for Physics-Based Robotic Interaction
中文标题:基于物理的机器人交互的遮挡鲁棒多目标解耦
作者:Xin Dong, Lihan Zhang, Tianru Dai, Wenfeng Deng, Yansong Tang
We propose a mask-free method for lossless multi-object 3D reconstruction from sparse and occluded real-world views, enabling physically plausible robotic interaction via Material Point Method (MPM) simulation. Our key insight is that object coupling stems from occlusion and limited viewpoints, which we address by formulating multi-object decoupling as a sparse-view reconstruction problem. Using 3D Gaussian Splatting as base representation, we first obtain coarse instance partitions with a SAM2-trained segmentation field. Rather than relying on masks, we reconstruct fragmented geometries by leveraging a joint Score Distillation Sampling (SDS) process, which integrates reference-view supervision with novel-view synthesis guided by 2D and 3D diffusion priors to enforce both texture fidelity and 3D consistency. Furthermore, we incorporate geometry-aware priors such as intra-object and inter-object similarity to regularize geometric reasoning. Experimental results demonstrate that our method produces complete, simulation-ready 3D objects without requiring manual masks, enabling realistic dynamic interactions on both synthetic, robotic and real-world datasets.
本文提出一种无掩码方法,可从稀疏且遮挡的真实世界视图实现无损多目标3D重建,并通过物质点方法(Material Point Method, MPM)仿真实现物理合理的机器人交互。本文的核心理念是目标耦合源于遮挡和有限视角限制,因此将多目标解耦问题形式化为稀疏视图重建问题。以3D高斯溅射(3D Gaussian Splatting)作为基础表示,首先利用SAM2训练的分割场获得粗粒度实例划分。与依赖掩码不同,本文通过联合分数蒸馏采样(Score Distillation Sampling, SDS)过程重建碎片化几何,该过程将参考视图监督与由2D和3D扩散先验引导的新视图合成相结合,以同时保证纹理保真度和3D一致性。此外,本文还引入几何感知先验(如目标内和目标间相似性)来约束几何推理。实验结果表明,本方法无需人工掩码即可生成完整、可用于仿真的3D对象,在合成数据集、机器人数据集和真实世界数据集上均实现了逼真的动态交互效果。
Entropy-Controlled Flow Matching
中文标题:熵控制流匹配
作者:Chika Maduabuchi
Modern vision generators transport a base distribution to data through time-indexed measures, implemented as deterministic flows (ODEs) or stochastic diffusions (SDEs). Despite strong empirical performance, standard flow-matching objectives do not directly control the information geometry of the trajectory, allowing low-entropy bottlenecks that can transiently deplete semantic modes. We propose Entropy-Controlled Flow Matching (ECFM): a constrained variational principle over continuity-equation paths enforcing a global entropy-rate budget d/dt H(mu_t) >= -lambda. ECFM is a convex optimization in Wasserstein space with a KKT/Pontryagin system, and admits a stochastic-control representation equivalent to a Schrodinger bridge with an explicit entropy multiplier. In the pure transport regime, ECFM recovers entropic OT geodesics and Gamma-converges to classical OT as lambda -> 0. We further obtain certificate-style mode-coverage and density-floor guarantees with Lipschitz stability, and construct near-optimal collapse counterexamples for unconstrained flow matching.
现代视觉生成器通过时间索引测度将基础分布传输至数据,实现为确定性流(ODE)或随机扩散(SDE)。尽管标准流匹配目标具有强大的实证性能,但其不直接控制轨迹的信息几何,允许低熵瓶颈暂时耗尽语义模式。我们提出熵控制流匹配(ECFM):一种在连续性方程路径上的约束变分原理,强制执行全局熵率预算 d/dt H(μ_t) ≥ -λ。ECFM 是 Wasserstein 空间中的凸优化问题,具有 KKT/Pontryagin 系统,并可表示为具有显式熵乘子的随机控制,等价于 Schrödinger 桥。在纯传输机制下,ECFM 恢复熵最优传输测地线,并在 λ → 0 时 Γ 收敛至经典最优传输。我们进一步获得了具有 Lipschitz 稳定性的证书式模态覆盖和密度下界保证,并为无约束流匹配构建了近似最优的坍缩反例。
Image Compression 每日总览
今日Image Compression分类的论文数量较少,主要关注深度学习驱动的图像压缩与表示效率优化。REDI论文聚焦于视觉Transformer(DINOv3)中的Token Reduction技术,这属于模型层面的图像信息压缩而非传统图像编码压缩。当前趋势显示,研究正从单纯压缩向“压缩-理解-效率”三位一体方向发展,即在减少数据量的同时保持甚至提升模型的视觉理解能力。
重点论文推荐:
- REDI: Corpus Aware Patch Ranking for DINOv3 Token Reduction - 提出基于语料库的patch重要性排序方法,能有效识别并去除DINOv3中的冗余token,在降低30-50%计算量的同时保持模型性能,为视觉Transformer的高效部署提供了新思路。
REDI: Corpus Aware Patch Ranking for DINOv3 Token Reduction
中文标题:REDI:面向DINOv3令牌缩减的语料库感知块排序方法
作者:Chanjong Im, Sebastian Diem, Thomas Mandl
Most token reduction methods for Vision Transformers seek favorable tradeoffs between accuracy and efficiency by pruning, merging, or pooling patch tokens. REDI (Relevance for DINOv3 Token Reduction) studies this question through a controlled supervised reference: how should a fixed token budget be allocated across patches for image classification? REDI quantizes final block DINOv3 patch representations into a visual vocabulary and derives class conditioned corpus scores using supervised TF-IDF over visual words. For each validation image, the ground truth class selects a row of the TF-IDF table, and four transformed views produce a TF-IDF map aligned to a reference center crop. A separate dense pass on the same crop provides an attention map. After independent min max normalization, their elementwise product defines the REDI score. A fixed keep, merge, and compress operator then uses score rank to assign patch roles and score magnitude to weight merging and compression. With precomputed REDI scores, a frozen DINOv3 ViT-B/16 backbone, and the same linear classifier used for dense evaluation, the operator reduces the sequence length from 201 to 107 tokens, a 46.8% sequence reduction. The REDI variant based on incoming attention mass achieves 84.706% Top-1 accuracy on ImageNet-1K, compared with 83.514% for the dense baseline, 82.634% for incoming attention mass alone, and 81.796% for supervised TF-IDF alone. The same corpus term also improves reduced classification for three alternative attention formulations relative to their attention only counterparts. Together, these controlled comparisons indicate that class specific corpus statistics and image specific attention provide complementary signals for patch ranking in this setting.
大多数Vision Transformer的令牌缩减方法旨在通过剪枝、合并或池化图像块标记来寻求精度与效率之间的良好权衡。REDI(DINOv3令牌缩减相关性)通过受控的监督参考来研究这一问题:对于图像分类任务,应如何在各图像块之间分配固定令牌预算?REDI将DINOv3最终块的图像块表示量化为视觉词汇表,并利用监督TF-IDF方法在视觉词上导出类别条件化的语料库分数。对于每个验证图像,真实类别会选择TF-IDF表的一行,四个变换视图产生与参考中心裁剪对齐的TF-IDF图。同一裁剪上的独立密集传递提供注意力图。经过独立的最小最大归一化后,它们的逐元素乘积定义了REDI分数。固定的保留、合并和压缩算子随后使用分数排序来分配图像块角色,并利用分数幅度来加权合并和压缩操作。借助预计算的REDI分数、冻结的DINOv3 ViT-B/16骨干网络以及与密集评估相同的线性分类器,该算子将序列长度从201个令牌缩减至107个,实现46.8%的序列缩减。基于传入注意力质量的REDI变体在ImageNet-1K上达到84.706%的Top-1准确率,相比密集基线的83.514%、单独使用传入注意力质量的82.634%以及单独使用监督TF-IDF的81.796%均有提升。同样的语料库术语相对于其仅使用注意力的对应方法,还改进了三种替代注意力形式的缩减分类。综合来看,这些对照实验表明在此设置下,类别特定的语料库统计信息和图像特定的注意力为图像块排序提供了互补的信号。
今日未找到该分类的匹配论文。
今日未找到该分类的匹配论文。
今日未找到该分类的匹配论文。
今日未找到该分类的匹配论文。