每日 arXiv 论文简报
今日arXiv论文呈现以下趋势:
- 多模态统一与自回归建模:多篇论文聚焦于统一多模态模型架构,自回归方法与扩散模型形成互补。Pareto LoRA和统一多模态自回归建模等工作致力于解决模态不平衡和上下文-视觉 token 共享问题。
- 物理信息与仿真建模:Phys4D、Volterra 生成模型、MeiBRD 等工作将物理一致性引入生成模型,实现 4D 建模和术中变形预测。
- 实时与高效推理:DiFlow-TTS、MaineCoon、MimicIK 等工作关注低延迟、实时生成能力,涵盖语音合成、社交世界模型和逆运动学。
- 安全性与对齐:视频扩散模型的表征 steering、测试时训练等方法被用于安全对齐和鲁棒性提升。
今日最值得关注的论文:
- DiFlow-TTS:首个基于离散流匹配的紧凑低延迟零样本 TTS,突破传统自回归生成效率瓶颈。
- Pareto LoRA:创新性地将 Pareto 最优梯度整合用于解决统一多模态模型中的模态不平衡问题。
- Phys4D:从视频扩散模型中提取细粒度物理一致性的 4D 建模,对物理仿真和生成有重要意义。
- STAR:提出时空自适应奖励分配框架,为文本到图像模型的强化学习后训练提供新范式。
- Pulling The REINS:无需训练的表征 steering 方法实现视频扩散模型安全对齐,兼具创新性和实用性。
今日 Autoregressive 分类论文总览
今日 Autoregressive(自回归)相关论文整体呈现多模态融合与架构创新两大趋势。自回归模型不再局限于传统序列生成任务,而是向表格结构识别、多模态统一建模、零样本语音合成等场景深度渗透。值得注意的亮点包括:1)离散 token 化与流匹配技术的结合,为生成式模型提供更高效的离散表示方法;2)自回归框架在多任务表格识别中的结构依赖建模取得新进展;3)统一多模态模型通过共享上下文-视觉分词器实现跨模态协同。
- Discrete Autoregressive Transformer for Generative Mechanism Synthesis - 提出离散自回归变换器,为生成机制合成提供新思路,离散表示增强了可解释性与控制性。
- Revisiting Structural Dependency in Autoregressive Multi-Task Table Recognition - 创新性地引入顺序无关的单元格级表示,解决多任务表格识别中的结构依赖难题。
- Unified Multimodal Autoregressive Modeling with Shared Context-Visual Tokenizer - 通过共享分词器实现多模态统一,验证了统一架构在跨模态任务中的有效性。
- DiFlow-TTS: Compact and Low-Latency Zero-Shot Text-to-Speech - 离散流匹配结合自回归机制,实现轻量级零样本 TTS,低延迟特性具实用价值。
- Pareto LoRA: Mitigating Modality Imbalance in Unified Multimodal Models - Pareto 最优梯度整合方法缓解模态不平衡,为多模态自回归训练提供新策略。
Comprehensive pKa Data Augmentation from Limited Real Data through an Engineered Models-Quantum Framework
中文标题:基于工程化模型-量子框架的有限真实数据pKa数据综合增强
作者:Wang Rui, Liu Dinghao
Proton dissociation constants (pKa) are critical for functional molecule discovery and molecular modeling. Building on iBonD, the largest experimental pKa database established, we and other researchers have developed several methods including machine-learning-based empirical prediction and high-accuracy energy calculations. Despite this foundation, the rapid augmentation of high-quality pKa data remains fundamentally constrained. As part of this work, we performed large-scale regression-based pKa prediction on unlabeled molecular datasets using a collection of extensively optimized machine-learning models. The results indicate that, since the feature distributions of unlabeled molecular datasets, the pKa data distribution approximates normality, with extreme scarcity of tail-region samples. Although such augmentation is highly valuable for improving overall data availability and predictive modeling, it remains insufficient for efficiently discovering molecules with broad-spectrum pKa properties. To address this, we explore the targeted generation of molecules with sparse pKa properties from the vast chemical space. Given that traditional continuous latent space VAE-RNN methods for molecular generation suffer from insufficient stability and fail to demonstrate clear advantages in complementing sparse data, we design and implement a quantum-assisted sparse-pKa molecular generation. Feasibility is validated on a simulated quantum annealer, and superior extreme-value sampling is further achieved on physical coherent Ising machines (CIMs). (to be continued)
质子解离常数(pKa)在功能分子发现和分子建模中至关重要。在iBonD(最大的实验pKa数据库)的基础上,我们与其他研究人员开发了多种方法,包括基于机器学习的经验预测和高精度能量计算。尽管已有上述基础,高质量pKa数据的快速扩充仍受到根本性限制。作为本工作的一部分,我们使用大量经过充分优化的机器学习模型对未标记分子数据集进行了大规模回归pKa预测。结果表明,由于未标记分子数据集的特征分布,pKa数据分布近似正态分布,尾部区域样本极为稀缺。尽管这种数据增强对提高整体数据可用性和预测建模非常有价值,但对于高效发现具有广谱pKa特性的分子仍显不足。为此,我们探索从广阔化学空间中定向生成具有稀疏pKa特性的分子。由于传统的连续潜空间VAE-RNN分子生成方法稳定性不足,且在补充稀疏数据方面未能表现出明显优势,我们设计并实现了一种量子辅助的稀疏pKa分子生成方法。在模拟量子退火机上验证了可行性,并在物理相干伊辛机(CIMs)上进一步实现了优异的极值采样。(待续)
Discrete Autoregressive Transformer for Generative Mechanism Synthesis
中文标题:用于生成式机构综合的离散自回归Transformer
作者:Anar Nurizada, Anurag Purwar
Planar path synthesis requires mechanisms whose coupler curves match a prescribed trajectory; the mapping from curve to linkage is inherently one-to-many across four-, six-, and eight-bar topologies. We address this design problem with simulation-grounded evaluation on a curated corpus of over one million mechanisms, reporting Chamfer distance and dynamic time warping after forward kinematics and geometric alignment. We formulate synthesis as conditional autoregressive sequence modeling: joint coordinates are uniformly quantized to tokens and generated by a decoder-only transformer with a variational-autoencoder (VAE) latent of the target curve and an explicit mechanism-type token. Training combines token cross-entropy with a Gaussian-smoothed bin auxiliary loss that respects ordinal structure among bins. At inference, a bounded latent-noise schedule decodes all mechanism types at each noise level; we retain the top five candidates by geometric error, yielding diverse accurate families without dataset lookup. On held-out tests, aggregate mean Chamfer distance is $0.0132$ and mean dynamic time warping is $0.153$; a latent $k$-nearest-neighbor baseline that conditions on training-set neighbor latents in VAE space achieves matched-topology mean Chamfer distance $0.0071$ and mean dynamic time warping $0.117$ using the same decoder.
平面轨迹合成需要连杆曲线与给定轨迹相匹配的机构;从曲线到连杆的映射在四杆、六杆和八杆拓扑结构中本质上是一对多的关系。我们通过在一个包含超过一百万个机构的精选语料库上进行基于仿真的评估来解决这一设计问题,报告了正向运动学和几何对齐后的Chamfer距离和动态时间规整值。我们将综合问题形式化为条件自回归序列建模:联合坐标被均匀量化为token,并由一个仅解码器的Transformer生成,该Transformer以目标曲线的变分自编码器(VAE)潜在向量和显式的机构类型token为条件。训练结合了token交叉熵损失和高斯平滑分箱辅助损失,后者尊重分箱之间的序数结构。在推理时,有界潜在噪声调度在每个噪声水平下解码所有机构类型;我们通过几何误差保留前五个候选,无需数据集查找即可产生多样化的精确族。在保留测试集上,聚合平均Chamfer距离为0.0132,平均动态时间规整值为0.153;使用相同解码器的潜在k近邻基线在VAE空间中对训练集邻居潜在向量进行条件化,达到了匹配拓扑结构的平均Chamfer距离0.0071和平均动态时间规整值0.117。
Pareto LoRA: Mitigating Modality Imbalance in Unified Multimodal Models via Pareto-Optimal Gradient Integration
中文标题:Pareto LoRA:通过帕累托最优梯度整合缓解统一多模态模型中的模态不平衡问题
作者:Xiwen Wei, Mark Nutter, Madhusudhanan Srinivasan, Radu Marculescu
Unified multimodal models (UMMs) have recently emerged as a promising paradigm for integrating multimodal understanding and generation within a single autoregressive transformer. However, during multimodal instruction tuning, these models often exhibit pronounced modality imbalance: language gradients dominate optimization, thus leading to lower image generation quality, especially under parameter-efficient fine-tuning such as LoRA. In this work, we systematically analyze modality imbalance in LoRA-based fine-tuning of UMMs for interleaved text-image generation. We show that vision modality performance degrades substantially more than text modality performance when compared to unimodal counterparts, and that modality-specific gradients can differ by orders of magnitude across various tasks and layers. Motivated by this observation, we reformulate the multimodal instruction tuning as a bi-objective optimization problem and propose Pareto LoRA, a Pareto-optimal gradient integration strategy that balances the text and image objectives by modulating the gradient direction and strength. Experiments on the CoMM benchmark with Emu2 demonstrate that Pareto LoRA consistently improves multimodal generation balance, achieving up to 44.9% gains in perceptual image quality over vanilla LoRA while maintaining comparable text performance.
统一多模态模型(UMMs)作为一种新兴范式,能够将多模态理解与生成整合到单一自回归Transformer中。然而,在多模态指令微调过程中,这些模型常表现出明显的模态不平衡:语言梯度主导优化过程,导致图像生成质量下降,尤其是在使用LoRA等参数高效微调方法时。本工作系统分析了基于LoRA的UMMs微调在交错文本-图像生成任务中的模态不平衡问题。研究表明,与单模态模型相比,视觉模态性能下降程度远高于文本模态,且不同任务和层的模态特定梯度可能存在数量级差异。基于这一观察,我们将多模态指令微调重新表述为双目标优化问题,并提出Pareto LoRA——一种通过调节梯度方向和强度来平衡文本与图像目标的帕累托最优梯度整合策略。在CoMM基准测试上使用Emu2进行的实验表明,Pareto LoRA持续改善多模态生成平衡,在保持文本性能相当的前提下,图像感知质量较标准LoRA提升高达44.9%。
MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model
中文标题:MaineCoon:追求实时视听社交世界模型
作者:Lichen Bai, Tianhao Zhang, Shitong Shao, Dingwei Tan, Qiyu Zhong, Zhengpeng Xie, Haopeng Li, Qinghao Huang, Dandan Shen, Tengjiao Ji, Wei Wang, Peicheng Wu, Yuxuan Zhao, Xiangyu Zhu, Welly Luo, Shurui Yang, Zeke Xie
As an increasing majority of global video content is consumed on social platforms for interactive social purposes, video generation models built for social worlds are important but largely overlooked by previous studies. In this work, we define the position of social world models and build a prototype model as the first step towards this goal. While previous world models successfully simulate physical environments or gaming world exploration, they remain fundamentally detached from human-centric social dynamics. To bridge this gap as the first step to social world models, we present MaineCoon, the first real-time audio-visual autoregressive model that has 22B parameters and is capable of real-time streaming generation and sub-second interaction, with a record-breaking frame rate of up to 47.5 FPS, on a single GPU. To the best of our knowledge, MaineCoon is also the first real-time audio-visual generation model specifically optimized for social-interactive applications. To enable efficient and stable training, we introduce several novel techniques into MaineCoon, including self-resampling, cross-modal representation alignment, domain-aware preference optimization, and reinforced online-policy distillation (ROPD). We also design the first agentic streaming inference framework that supports thousand-second-scale or even longer generation while mitigating drift with agentic cache management and prompt planing. These innovations significantly accelerate training while optimizing real-time inference performance. We believe this work not only sets a new state-of-the-art (SOTA) performance benchmark for high-quality, low-latency, and long-horizon audio-visual autoregressive models, but also points out the paradigm shift desired for next-generation AI-native social platforms.
随着全球视频内容越来越多地在社交平台上被消费用于社交互动目的,为社交世界构建的视频生成模型非常重要,但此前的研究却基本忽略了这一领域。在本工作中,我们定义了社交世界模型的地位,并构建了一个原型模型作为实现这一目标的第一步。此前的气候模型成功模拟了物理环境或游戏世界的探索,但它们从根本上与以人为中心的社交动态脱节。作为迈向社交世界模型的第一步,我们提出了MaineCoon,这是首个实时视听自回归模型,拥有220亿参数,能够实现实时流式生成和亚秒级交互,在单块GPU上实现了高达47.5 FPS的创纪录帧率。据我们所知,MaineCoon也是首个专门针对社交互动应用优化的实时视听生成模型。为了实现高效稳定的训练,我们在MaineCoon中引入了多项新技术,包括自重采样、跨模态表示对齐、领域感知偏好优化和强化在线策略蒸馏(ROPD)。我们还设计了首个智能体流式推理框架,支持千秒级甚至更长时间的生成,同时通过智能体缓存管理和提示规划来缓解漂移。这些创新显著加速了训练过程,同时优化了实时推理性能。我们相信,本工作不仅为高质量、低延迟、长时序视听自回归模型设定了新的最先进(SOTA)性能基准,而且为下一代AI原生社交平台指明了范式转变的方向。
Revisiting Structural Dependency in Autoregressive Multi-Task Table Recognition via Order-Independent Cell-Level Representations
中文标题:通过顺序无关的单元格级表示重新审视自回归多任务表格识别中的结构依赖
作者:Takaya Kawakatsu
Multi-task table recognition jointly addresses table structure prediction, cell localization, and cell content recognition within a unified framework. Existing approaches often rely on autoregressive decoders to generate table structures and reuse their hidden states for cell localization and content recognition. This autoregressive generation process can make cell representations order-dependent, degrading global consistency across cells. This paper proposes a structural refinement module that produces order-independent cell features through non-causal attention. This design enables parallel inference of cell contents while conditioning each cell on global context encoded in the refined features. Experiments on two large datasets demonstrate consistent gains in cell localization and end-to-end recognition, while reducing overall inference time by around threefold.
多任务表格识别在一个统一框架内联合处理表格结构预测、单元格定位和单元格内容识别。现有方法通常依赖自回归解码器来生成表格结构,并重用其隐藏状态进行单元格定位和内容识别。这种自回归生成过程可能导致单元格表示具有顺序依赖性,从而削弱单元格间的全局一致性。本文提出了一种结构细化模块,通过非因果注意力生成顺序无关的单元格特征。该设计使得单元格内容能够并行推理,同时基于细化特征中编码的全局上下文对每个单元格进行条件化处理。在两个大型数据集上的实验表明,单元格定位和端到端识别性能均获得持续提升,同时整体推理时间减少约三分之一。
Adaptive Volumetric Mechanical Property Fields Invariant to Resolution
中文标题:自适应体积力学特性场(与分辨率无关)
作者:Rishit Dagli, Donglai Xiang, Vismay Modi, Xuning Yang, Gavriel State, David I. W. Levin, Maria Shugrina
Accurate mechanical properties (or materials) Young's modulus ($E$), Poisson&x27;s ratio ($\nu$) and density ($\rho$) are essential for reliable physics simulation of digital worlds, but most 3D assets lack this information. We propose AdaVoMP, a method for predicting accurate dense spatially-varying ($E$, $\nu$, $\rho$) for input 3D objects across representations, improving the resolution, accuracy, and memory efficiency over the state-of-the-art. The foundation of our technique is a sparse and adaptive voxel structure SAV that efficiently represents both the input 3D shape and the material field output. We replace the fixed-voxel model of the most accurate prior method, VoMP, with a novel sparse transformer encoder-decoder model that learns to generate a unique SAV autoregressively for every input shape to represent its materials, achieving a resolution $16^3\times$ higher than prior art. Experiments show that AdaVoMP estimates more accurate volumetric properties, even with lesser test-time compute than all prior art. This allows us to convert high-resolution complex 3D objects into simulation-ready assets, resulting in realistic deformable simulations.
准确的力学特性(即材料属性)杨氏模量($E$)、泊松比($ u$)和密度($ ho$)对于数字世界的可靠物理模拟至关重要,但大多数3D资产缺乏这些信息。我们提出了AdaVoMP方法,用于跨表示预测输入3D对象的精确密集空间变化($E$, $ u$, $ ho$),在分辨率、准确性和内存效率方面均优于现有最先进方法。我们技术的核心是稀疏自适应体素结构SAV,它能高效地表示输入3D形状和材料场输出。我们将最精确的先前方法VoMP的固定体素模型替换为一种新颖的稀疏Transformer编码器-解码器模型,该模型学习为每个输入形状自回归地生成独特的SAV来表示其材料特性,实现了比现有技术高$16^3$倍的分辨率。实验表明,AdaVoMP估计的体积特性更加精确,且测试时计算量比所有先前方法更少。这使我们能够将高分辨率复杂3D对象转换为仿真就绪资产,从而实现逼真的可变形模拟。
Unified Multimodal Autoregressive Modeling with Shared Context-Visual Tokenizer is Key to Unification
中文标题:使用共享上下文-视觉分词器的统一多模态自回归建模是实现统一的关键
作者:Wujian Peng, Lingchen Meng, Yuxuan Cai, Xianwei Zhuang, Yuhuan Yang, Rongyao Fang, Chenfei Wu, Junyang Lin, Zuxuan Wu, Shuai Bai
Unified Multimodal Modeling aims to integrate visual understanding and generation within a single system. However, existing approaches typically rely on two disparate visual tokenizers, which splits the representation space and hinders truly unified modeling. We propose UniAR, a unified autoregressive framework where a single discrete visual tokenizer serves as the key bridge between understanding and generation, enabling a shared context in which the model can directly interpret its own generated visual tokens without additional re-encoding. UniAR adapts a pretrained vision encoder with multi-level feature fusion and a lookup-free bitwise quantization scheme, preserving both high-level semantics and low-level details while scaling the effective visual vocabulary at minimal cost. Building on this, the unified autoregressive model adopts parallel-bitwise-prediction to jointly predict spatially grouped, multi-level visual codes, substantially reducing visual sequence length and accelerating generation. Finally, a diffusion-based visual decoder operates on discrete visual tokens to decode high-fidelity images. Through large-scale pre-training, followed by supervised fine-tuning and reinforcement learning, UniAR achieves state-of-the-art performance on image generation and image editing while remaining competitive on multimodal understanding benchmarks. The project page is available at https://sharelab-sii.github.io/uniar-web.
统一多模态建模旨在将视觉理解和生成集成在单一系统中。然而,现有方法通常依赖两个不同的视觉分词器,这导致表示空间分裂,阻碍了真正的统一建模。我们提出UniAR,这是一个统一的自回归框架,其中单一离散视觉分词器作为理解和生成之间的关键桥梁,使模型能够在共享上下文中直接解释自身生成的视觉标记,无需额外的重新编码。UniAR采用预训练视觉编码器配合多级特征融合和无查找表位量化方案,在以最低成本扩展有效视觉词汇的同时保留高级语义和低细节。在此基础上,统一自回归模型采用并行位预测来联合预测空间分组的多级视觉编码,大幅缩短视觉序列长度并加速生成。最后,基于扩散的视觉解码器对离散视觉标记进行操作以解码高保真图像。通过大规模预训练,随后进行监督微调和强化学习,UniAR在图像生成和图像编辑方面实现了最先进的性能,同时在多模态理解基准测试中保持竞争力。项目页面见 https://sharelab-sii.github.io/uniar-web。
DiFlow-TTS: Compact and Low-Latency Zero-Shot Text-to-Speech with Discrete Flow Matching
中文标题:DiFlow-TTS: 基于离散流匹配的紧凑低延迟零样本文本到语音合成
作者:Ngoc-Son Nguyen, Thanh V. T. Tran, Hieu-Nghia Huynh-Nguyen, Truong-Son Hy, Van Nguyen
Zero-shot text-to-speech (TTS) has made significant progress in replicating unseen voices, yet balancing generation quality and inference efficiency remains challenging. Autoregressive models suffer from high latency, while diffusion-based approaches are constrained by training-time configurations. Moreover, most flow-based methods operate in continuous space, which introduces optimization challenges because continuous token spaces are inherently more complex than discrete ones. To address these limitations, we propose DiFlow-TTS, a novel zero-shot TTS framework based on discrete flow matching. The model consists of a deterministic Phoneme-Content Mapper for linguistic modeling and a Factorized Discrete Flow Denoiser that simultaneously generates prosody and acoustic token streams. Experimental results demonstrate the effectiveness of our approach across multiple evaluation metrics.
零样本文本到语音合成(Zero-shot TTS)在复现未见声音方面取得了显著进展,然而在生成质量与推理效率之间取得平衡仍具挑战性。自回归模型存在高延迟问题,而基于扩散的方法受限于训练时的配置约束。此外,大多数基于流的方法在连续空间中进行操作,这带来了优化难题,因为连续token空间本质上比离散空间更为复杂。为解决这些局限性,我们提出了DiFlow-TTS,一个基于离散流匹配的新型零样本TTS框架。该模型由一个确定性的音素内容映射器(Phoneme-Content Mapper)用于语言建模,以及一个因子化离散流去噪器(Factorized Discrete Flow Denoiser)同时生成韵律和声学token流组成。实验结果验证了我们的方法在多个评估指标上的有效性。
Diffusion 分类论文概述(2025年6月)
今日 Diffusion 相关论文呈现出多模态扩展与细化的趋势。既有针对 Text-to-Image 后训练的对齐优化(STAR),也有面向视频扩散模型的安全防护(REINS),同时3D/4D生成与编辑继续是热点方向(Edit3DGS、TextMesh4D、Phys4D)。值得注意的是,零样本与训练-free 技术持续受到关注,细节增强与测试时训练成为提升效率的关键路径。此外,离散扩散(DVD)与流匹配(DiFlow-TTS)等新架构也在探索更高效的生成范式。
- STAR:提出时空自适应奖励分配机制,解决 Text-to-Image RL 后训练中偏好学习效率低下的问题,为生成质量优化提供了新思路。
- Pulling The REINS:首个训练-free 的视频扩散模型安全对齐框架,通过表征 steering 实现无需微调的伤害内容防护,具有重要实用价值。
- Detail++:提出训练-free 的细节增强器,无需额外训练即可提升扩散模型的细节质量,兼顾效率与效果。
- Phys4D:从视频扩散模型中提取物理一致的4D建模,实现细粒度动作与物理规律的结合,推动生成内容向真实世界模拟迈进。
- Edit3DGS:统一框架结合2D指令引导扩散与3D Gaussian Splatting,实现动态头部的高质量编辑,弥合了2D与3D编辑的鸿沟。
STAR: SpatioTemporal Adaptive Reward Allocation for Text-to-Image RL Post-Training
中文标题:STAR:用于文本到图像强化学习后训练的时空自适应奖励分配
作者:Jinjie Shen, Wei Deng, Xian Hu, Daiguo Zhou, Jian Luan
Existing RL post-training methods for text-to-image generation usually convert the final-image reward into a single scalar advantage and apply it with the same strength to the entire generative trajectory. However, text-to-image generation naturally has temporal and spatial structure: different denoising steps are responsible for different generation stages, and the content that truly determines text alignment often appears only in part of the image. This granularity mismatch makes it difficult for policy updates to focus on the generative components that actually affect the reward. To address this issue, we propose \textbf{SpatioTemporal Adaptive Reward (STAR) Allocation} for RL post-training of text-to-image diffusion and flow models. STAR uses text-image attention inside the generative model and starts from the core content that the user truly cares about in the prompt. It constructs spatial allocation maps that dynamically vary across denoising steps and rollouts, and allocates the same group-relative advantage to more relevant latent regions with almost no additional computational overhead. STAR then applies stronger policy updates to these regions through a spatially resolved policy objective. We use Stable Diffusion 3.5 Medium as the base model and evaluate on three tasks: GenEval, OCR text rendering, and PickScore. Experimental results show that STAR improves compositional semantic alignment, text rendering, and preference optimization without changing the external reward source, achieving $\mathbf{0.9759}$, $\mathbf{0.9757}$, and $\mathbf{23.60}$ on GenEval, OCR, and PickScore, respectively.
现有文本到图像生成的强化学习后训练方法通常将最终图像奖励转换为单个标量优势值,并以相同强度应用于整个生成轨迹。然而,文本到图像生成自然具有时间和空间结构:不同的去噪步骤负责不同的生成阶段,而真正决定文本对齐的内容通常仅出现在图像的局部区域。这种粒度不匹配使得策略更新难以聚焦于实际影响奖励的生成组件。为解决这一问题,我们为文本到图像扩散模型和流模型的强化学习后训练提出了一种时空自适应奖励分配(STAR)方法。STAR利用生成模型内部的文本-图像注意力机制,从用户提示中真正关心的核心内容出发,构建了跨去噪步骤和采样动态变化的空间分配图,并以几乎可忽略的计算开销为更相关的潜空间区域分配同组相对优势。STAR随后通过空间解析的策略目标对这些区域实施更强的策略更新。我们以Stable Diffusion 3.5 Medium为基座模型,在三个任务上进行评估:GenEval、OCR文本渲染和PickScore。实验结果表明,STAR在不改变外部奖励源的情况下提升了组合语义对齐、文本渲染和偏好优化性能,分别在GenEval、OCR和PickScore上达到0.9759、0.9757和23.60的分数。
ANEForge: Python for direct computation on the Apple Neural Engine
中文标题:ANEForge:用于直接在Apple Neural Engine上进行计算的Python工具
作者:Spencer H. Bryngelson
ANEForge is a Python package that programs the Apple Neural Engine (ANE), the fixed-function neural accelerator on every recent Apple device, directly and without CoreML. In production the engine is reachable only through CoreML, which treats it as a scheduling option: no configuration requires the ANE, and a model can silently run on the CPU or GPU instead. ANEForge compiles a lazy tensor graph, built from 58 fused operators and 19 native bridge operators, into a single ANE program. The program is dispatched through the same ANE daemon and kernel-driver stack as Apple's internal framework. Beyond inference, the package reaches the engine&x27;s native fused attention, streams int8, int4, and sparse weights, keeps decoder and optimizer state resident across steps, and runs the forward pass, backward pass, and optimizer update of training on the engine. A small fused program completes a call in about 90us, near the engine's 70us per-program dispatch floor, and a pretrained ResNet-18 forward runs end-to-end in 0.33ms. ResNet-18, a sentence encoder, and a Vision Transformer run end-to-end against framework references, and a Stable Diffusion U-Net validates its forward pass. ANEForge targets Apple Silicon under macOS 14 and later. Each release is verified against a recorded macOS and ANE-compiler version.
ANEForge是一个Python包,可直接对Apple Neural Engine(ANE)进行编程,这是苹果近期设备上的固定功能神经网络加速器,无需使用CoreML。在生产环境中,该引擎只能通过CoreML访问,CoreML将其作为一个调度选项:没有任何配置要求必须使用ANE,模型可能会静默地在CPU或GPU上运行。ANEForge将一个由58个融合算子和19个原生桥接算子构成的惰性张量图,编译成单个ANE程序。该程序通过与苹果内部框架相同的ANE守护进程和内核驱动栈进行调度。除了推理之外,该包还能访问引擎的原生融合注意力机制,支持int8、int4和稀疏权重的流式传输,保持解码器和优化器状态在多个步骤间驻留,并在引擎上运行训练的前向传播、反向传播和优化器更新。一个小型融合程序完成一次调用约需90微秒,接近引擎每程序70微秒的调度基线,预训练的ResNet-18前向传播端到端运行仅需0.33毫秒。ResNet-18、句子编码器和Vision Transformer与框架参考进行端到端对比验证,Stable Diffusion U-Net验证了其前向传播。ANEForge面向macOS 14及更高版本的Apple Silicon。每次发布都经过针对记录的macOS和ANE编译器版本进行验证。
Pulling The REINS: Training-Free Safety Alignment of Video Diffusion Models via Representation Steering
中文标题:拉起REINS:通过表征空间推理时 steering 实现视频扩散模型的免训练安全对齐
作者:Rohit Kundu, Arindam Dutta, Sarosij Bose, Athula Balachandran, Amit K. Roy-Chowdhury
Open-weight video diffusion models can generate photorealistic unsafe content, from violence to misinformation, yet existing defenses either require expensive safety fine-tuning that degrades general capability, or apply external filters that are trivially bypassed by adversarial prompts. We present REINS (REpresentation-space INference-time Safety steering), a training-free method that aligns video diffusion models at inference time by steering their internal representations toward safe generation. Our key finding is that safety-relevant structure is linearly encoded in the hidden-state activations of video diffusion transformers, and a single direction, discovered via Supervised PCA on binary safety labels, suffices to separate safe from unsafe generation trajectories. At inference, adding this direction to hidden states at an intermediate transformer layer redirects generation from harmful content to semantically related safe alternatives, with no weight updates, no concept enumeration, and negligible computational overhead. Through mechanistic analysis, we reveal that while safety information accumulates monotonically with transformer depth, steering effectiveness peaks at intermediate layers (~50% depth), exposing a fundamental tradeoff between information availability and downstream propagation capacity. We evaluate REINS across 9 video diffusion models, multiple parameter scales (1.3B-5B), and both text-to-video and image-to-video generation, to our knowledge, the broadest safety evaluation suite in the video generation literature.
开源权重的视频扩散模型能够生成包含暴力到虚假信息在内的真实感不安全内容,而现有防御措施要么需要昂贵的会对通用能力造成损害的安全微调,要么应用可被对抗性提示轻易绕过的外部过滤器。我们提出了REINS(表征空间推理时安全 steering),这是一种在推理时通过对模型内部表征进行 steering 来实现视频扩散模型安全对齐的免训练方法。我们的关键发现是,安全相关结构在视频扩散 transformer 的隐藏状态激活中呈线性编码,通过在二元安全标签上进行监督主成分分析发现的单个方向足以分离安全与不安全的生成轨迹。在推理时,在中间 transformer 层的隐藏状态上添加该方向即可将生成从有害内容重定向到语义相关的安全替代方案,无需权重更新、无需概念枚举,且计算开销可忽略不计。通过机制分析我们揭示了安全信息虽随 transformer 深度单调积累,但 steering 有效性在中间层(约 50% 深度)达到峰值,这揭示了信息可用性与下游传播能力之间的根本权衡。我们对 REINS 进行了跨 9 个视频扩散模型、多种参数规模(1.3B-5B)的评估,涵盖文本到视频和图像到视频生成,据我们所知,这是视频生成文献中最为广泛的安全评估套件。
MeiBRD: Meta-Learning Intraoperative Biomechanical Residual Deformation
中文标题:MeiBRD: 元学习术中生物力学残余变形
作者:Casey Meisenzahl, Jon Heiselman, Michael Holtz, Yubo Ye, Michael Miga, Linwei Wang
Accurate intraoperative liver registration is challenging due to substantial soft-tissue deformation yet sparse intraoperative measurements. Biomechanical models regularize this ill-posedness with prior knowledge but exhibit persistent prediction bias due to simplifying assumptions, while data-driven learning solutions struggle with data efficiency, generalization, and physical plausibility. We propose a hybrid registration framework that adapts a biomechanical prior using sparse intraoperative correspondences. Rather than learning a full deformation field, we learn a residual deformation function that corrects linear biomechanical predictions, modeled as a graph neural diffusion function with geometry-aware attention over the 3D liver mesh. To enable long-range information transfer of sparse observations, we take a novel perspective of sparse intraoperative measurements as \textit{context} samples where input-output pairs of the residual deformation function are fully observed, casting the problem into learning-to-learn this residual function from intraoperative context samples with feedforward meta-learners. Experiments on a deformable liver phantom dataset demonstrate improved registration accuracy and generalization compared to rigid, biomechanical, and data-driven baselines, particularly for out-of-distribution geometries and deformations.
由于显著的软组织变形和稀疏的术中测量值,精确的术中肝脏配准仍具挑战性。生物力学模型利用先验知识对这一欠定问题进行正则化,但由于简化假设而存在持续的预测偏差,而数据驱动的学习方案在数据效率、泛化性和物理合理性方面存在困难。我们提出了一种混合配准框架,利用稀疏的术中对应关系来适配生物力学先验。与学习完整变形场不同,我们学习的是一个残余变形函数,用于校正线性生物力学预测,该函数被建模为在3D肝脏网格上具有几何感知注意力的图神经扩散函数。为了实现稀疏观测的长程信息传递,我们将稀疏术中测量值作为一种上下文样本的全新视角来看待,在该视角下残余变形函数的输入-输出对被完全观测到,从而将问题转化为通过前馈元学习器从术中上下文样本中学习该残余函数。在可变形肝脏仿体数据集上的实验表明,与刚性、生物力学和数据驱动的基线方法相比,配准精度和泛化性能得到改善,特别是在分布外几何形状和变形情况下。
Volterra Generative Models
中文标题:沃尔泰拉生成模型
作者:Yusen Jia, Bingyan Han
Score-based diffusion models typically use Brownian perturbations, which provide tractable reverse-time dynamics but impose memoryless noising. We introduce Volterra generative models, a continuous-time score-based framework whose forward process injects path-dependent noise through fractional kernels. To handle the non-Markovian and non-semimartingale dynamics, we construct finite-dimensional Markovian lifts using Gaussian quadrature in both regimes and a hybrid finite-difference exponential approximation in the smooth regime. We prove squared error bounds, derive an augmented linear-Gaussian forward process, and show that the learning can remain data-dimensional by considering residual states and analytic auxiliary Gaussian scores. We also identify covariance and reverse-time degeneracies caused by shared Brownian factors and signed smooth-regime weights. The degeneracy motivates stabilized conditioning and, for stiff larger lifts, a Gaussian-bridge reconstruction sampler. Experiments on MNIST and CIFAR-10 show that persistent fractional perturbations with small Markovian lifts can improve score-based generation on MNIST and provide a promising extension to natural images, while the bridge sampler provides a stability mechanism for larger lifts.
分数基础扩散模型通常使用布朗扰动,这种扰动虽可提供易于处理的时间反向动力学,但会产生无记忆的噪声化过程。本文提出沃尔泰拉生成模型,这是一种连续时间分数基础框架,其前向过程通过分数核注入路径依赖噪声。为处理非马尔可夫和非半鞅动力学,我们采用两种机制下的高斯求积以及平滑机制下的混合有限差分指数近似来构建有限维马尔可夫提升。我们证明了均方误差界,导出了增广线性高斯前向过程,并表明通过考虑残差状态和解析辅助高斯分数,学习过程可以保持数据维度。我们还识别了由共享布朗因子和带符号平滑区域权重引起的协方差和时间反向退化问题。这种退化促使我们进行稳定化条件设计,对于刚性较大的提升情形,提出了一种高斯桥重构采样器。在MNIST和CIFAR-10上的实验表明,具有小马尔可夫提升的持续分数扰动可以改善MNIST上的分数基础生成,并为自然图像提供了有前景的扩展,同时桥采样器为较大提升情形提供了稳定性机制。
ReAge3D: Re-Aging 3D Faces with View Consistency
中文标题:ReAge3D:具有视图一致性的3D人脸年龄重塑
作者:Libing Zeng, Li Ma, Mingming He, Ning Yu, Paul Debevec, Nima Khademi Kalantari
We present a novel framework for realistic and controllable 3D face re-aging which produces highly detailed, identity-preserving results. Existing 3D editing methods, while effective for coarse semantic changes, are not well suited for re-aging, as even small inconsistencies across re-aged 2D views can lead to over-smoothing of subtle but perceptually important age-related details. To address this challenge, we first introduce a 2D diffusion-based re-aging model, DiffReaging, trained on synthetically generated image pairs. We further propose a center-out editing propagation strategy that leverages this re-aging model to reconstruct multi-view-consistent re-aged images. Specifically, starting from a re-aged frontal pivot view, we reconstruct the remaining views through warping and our proposed Masked-DiffReaging process. By injecting existing content at every step of the diffusion process, Masked-DiffReaging ensures that the reconstructed regions remain coherent with existing pixels. The resulting consistent set of re-aged views supervises the optimization of the re-aged 3D representation. Our method outperforms existing 3D editing techniques both visually and quantitatively, enabling smooth, fine-grained control over age transformations in 3D face models.
我们提出了一个用于真实且可控的3D人脸年龄重塑的新颖框架,该框架能够产生高度详细且保持身份特征的结果。现有的3D编辑方法虽然对粗粒度语义变化有效,但在年龄重塑方面表现不佳,因为即使重塑后的2D视图之间存在微小不一致,也可能导致微妙但感知上重要的年龄相关细节被过度平滑化。为解决这一挑战,我们首先引入了一个基于2D扩散的年龄重塑模型DiffReaging,该模型在合成生成的图像对上进行训练。我们进一步提出了一种中心向外的编辑传播策略,利用该年龄重塑模型来重建具有多视图一致性的重塑图像。具体而言,我们从一个重塑后的正面枢轴视图开始,通过变形和我们提出的Masked-DiffReaging过程来重建其余视图。通过在扩散过程的每一步注入现有内容,Masked-DiffReaging确保重建区域与现有像素保持一致。由此产生的重塑视图一致集合监督重塑3D表示的优化。我们的方法在视觉和定量方面均优于现有的3D编辑技术,能够对3D人脸模型中的年龄变换进行平滑、细粒度的控制。
Detail++: Training-Free Detail Enhancer for Text-to-Image Diffusion Models
中文标题:Detail++:用于文本到图像扩散模型的免训练细节增强器
作者:Lifeng Chen, Jiner Wang, Zihao Pan, Beier Zhu, Xiaofeng Yang, Chi Zhang
Recent advances in text-to-image (T2I) generation have led to impressive visual results. However, these models still face significant challenges when handling complex prompt, particularly those involving multiple subjects with distinct attributes. Inspired by the human drawing process, which first outlines the composition and then incrementally adds details, we propose Detail++, a training-free framework that introduces a novel Progressive Detail Injection (PDI) strategy to address this limitation. Specifically, we decompose a complex prompt into a sequence of simplified sub-prompts, guiding the generation process in stages. This staged generation leverages the inherent layout-controlling capacity of self-attention to first ensure global composition, followed by precise refinement. To achieve accurate binding between attributes and corresponding subjects, we exploit cross-attention mechanisms and further introduce a Centroid Alignment Loss at test time to reduce binding noise and enhance attribute consistency. Extensive experiments on T2I-CompBench and a newly constructed style composition benchmark demonstrate that Detail++ significantly outperforms existing methods, particularly in scenarios involving multiple objects and complex stylistic conditions.
文本到图像(T2I)生成的最新进展已取得了令人印象深刻的视觉效果。然而,这些模型在处理复杂提示词时仍面临重大挑战,特别是涉及具有不同属性的多个主体的情况。受到人类绘画过程的启发——首先勾勒构图,然后逐步添加细节——我们提出了Detail++,一个免训练框架,引入了新型的渐进式细节注入(PDI)策略来应对这一局限性。具体而言,我们将复杂提示词分解为一系列简化的子提示词,分阶段引导生成过程。这种分阶段生成利用了自注意力机制的内在布局控制能力,首先确保全局构图,随后进行精确细化。为了实现属性与对应主体之间的精确绑定,我们利用交叉注意力机制,并在测试时进一步引入质心对齐损失,以减少绑定噪声并增强属性一致性。在T2I-CompBench和一个新构建的风格组合基准上进行的广泛实验表明,Detail++显著优于现有方法,特别是在涉及多个物体和复杂风格条件的场景中。
From Noise to Order: Learning to Rank via Denoising Diffusion
中文标题:从噪声到有序:基于去噪扩散的学习排序
作者:Sajad Ebrahimi, Bhaskar Mitra, Negar Arabzadeh, Ye Yuan, Haolun Wu, Fattane Zarrinkalam, Ebrahim Bagheri
Learning-to-rank (LTR) methods have traditionally been limited to discriminative machine learning approaches that model the probability of the document being relevant to the query given some feature representation of the query-document pair. We propose an alternative denoising diffusion-based generative approach to LTR that instead models the full joint distribution over features and relevance labels. While in discriminative LTR, an over-parameterized ranking model may find different ways to fit the training data, we posit that candidate solutions that can explain the full data distribution under the generative setting maybe better at estimating relevance. Thus, we propose DiffusionRank that extends TabDiff, an existing diffusion model for tabular datasets, to create generative alternatives to classical discriminative pointwise and pairwise LTR objectives. Our work demonstrates improvements from DiffusionRank over discriminative counterparts on four standard LTR datasets and points to a rich space for future exploration to leverage ongoing advancements in deep generative models for LTR. Our code is publicly available at https://github.com/sadjadeb/DiffusionRank.
学习排序方法传统上局限于判别式机器学习方法,这些方法对给定查询-文档对特征表示的文档与查询相关的概率进行建模。我们提出了一种基于去噪扩散的生成式学习排序替代方法,该方法对特征和相关性标签的完整联合分布进行建模。在判别式学习排序中,过度参数化的排序模型可能通过不同方式拟合训练数据,我们假设能够在生成式设置下解释完整数据分布的候选解决方案可能更好地估计相关性。因此,我们提出了DiffusionRank,该方法扩展了TabDiff(一种针对表格数据集的现有扩散模型),为经典的判别式点级和成对学习排序目标创建了生成式替代方案。我们的工作展示了DiffusionRank在四个标准学习排序数据集上相比判别式方法的改进,并指出了一个丰富的未来探索空间,以利用深度生成模型在学习排序领域的最新进展。我们的代码已在https://github.com/sadjadeb/DiffusionRank上公开。
Phys4D: Fine-Grained Physics-Consistent 4D Modeling from Video Diffusion
中文标题:Phys4D: 从视频扩散模型进行细粒度物理一致的4D建模
作者:Haoran Lu, Shang Wu, Songling Liu, Jianshu Zhang, Maojiang Su, Guo Ye, Chenwei Xu, Lie Lu, Pranav Maneriker, Fan Du, Manling Li, Zhaoran Wang, Han Liu
Recent video diffusion models have achieved impressive capabilities as large-scale generative world models. However, these models often struggle with fine-grained physical consistency, exhibiting physically implausible dynamics over time. In this work, we present \textbf{Phys4D}, a pipeline for learning physics-consistent 4D world representations from video diffusion models. Phys4D adopts \textbf{a three-stage training paradigm} that progressively lifts appearance-driven video diffusion models into physics-consistent 4D world representations. We first bootstrap robust geometry and motion representations through large-scale pseudo-supervised pretraining, establishing a foundation for 4D scene modeling. We then perform physics-grounded supervised fine-tuning using simulation-generated data, enforcing temporally consistent 4D dynamics. Finally, we apply simulation-grounded reinforcement learning to correct residual physical violations that are difficult to capture through explicit supervision. To evaluate fine-grained physical consistency beyond appearance-based metrics, we introduce a set of \textbf{4D world consistency evaluation} that probe geometric coherence, motion stability, and long-horizon physical plausibility. Experimental results demonstrate that Phys4D substantially improves fine-grained spatiotemporal and physical consistency compared to appearance-driven baselines, while maintaining strong generative performance. Our project page is available at https://sensational-brioche-7657e7.netlify.app/
近期视频扩散模型作为大规模生成式世界模型已展现出令人瞩目的能力。然而,这些模型在细粒度物理一致性方面常常存在不足,表现出随时间推移在物理上不可信的动力学行为。在本工作中,我们提出了Phys4D,一个从视频扩散模型学习物理一致的4D世界表示的pipeline。Phys4D采用三阶段训练范式,逐步将外观驱动的视频扩散模型提升为物理一致的4D世界表示。我们首先通过大规模伪监督预训练来引导稳健的几何和运动表示,为4D场景建模奠定基础。然后,我们使用模拟生成的数据进行基于物理的有监督微调,以强制执行时间一致的4D动力学。最后,我们应用基于模拟的强化学习来纠正难以通过显式监督捕捉的残留物理违规问题。为了评估超越外观指标的细粒度物理一致性,我们引入了一套4D世界一致性评估方法,用于探测几何一致性、运动稳定性和长时域物理可信度。实验结果表明,与外观驱动的基线方法相比,Phys4D显著提升了细粒度时空和物理一致性,同时保持了强大的生成性能。
SP-GCRL: Influence Maximization on Incomplete Social Graphs
中文标题:SP-GCRL:面向不完整社交图的影响力最大化
作者:Haohua Niu, Yuxuan Yang, Lingfeng Zhang, Hao Li, Jiao Liang, Zongfu Luo, Luca Rossi
Influence maximization (IM) in real platforms is challenged by incomplete, noisy social graphs and non-stationary diffusion dynamics. We propose SP-GCRL, a social-propagation-aware graph contrastive reinforcement learning framework that learns end-to-end seed selection under partial observability.We first introduce a social-propagation-aware nonlinear diffusion function to model reinforcement/diminishing effects and probability drift under repeated exposure; we then construct dual structural views and perform contrastive learning to obtain node representations robust to missing edges and weak ties, while replacing expensive strategy metrics with a GAT-based regression surrogate to improve efficiency and scalability; finally, we use DDQN to learn an end-to-end seed selection policy on top of these representations. Experiments on multiple real-world networks show that SP-GCRL achieves significant gains over heuristic and learning-based baselines across budgets and topologies, while maintaining strong large-scale scalability.
现实平台中的影响力最大化(IM)面临不完整、有噪声的社交图以及非平稳扩散动态的挑战。我们提出了SP-GCRL,这是一种社会传播感知的图对比强化学习框架,能够在部分可观测条件下学习端到端的种子选择。我们首先引入了一种社会传播感知的非线性扩散函数,用于建模重复暴露下的强化/衰减效应和概率漂移;随后构建双重结构视图并进行对比学习,以获得对缺失边和弱关系具有鲁棒性的节点表示,同时用基于GAT的回归替代模型取代昂贵的策略指标,以提高效率和可扩展性;最后,我们在这些表示的基础上使用DDQN学习端到端的种子选择策略。在多个真实网络上的实验表明,SP-GCRL在不同预算和拓扑条件下相比启发式和基于学习的基线方法均取得了显著的性能提升,同时保持了较强的大规模可扩展性。
MimicIK: Real-Time Generative Inverse Kinematics from Teleoperation with FK Consistency
中文标题:MimicIK: 基于遥操作的实时生成式逆运动学及FK一致性
作者:Jiahao Yang, Shenhao Yan, Fan Feng, Chengsi Yao, Ge Wang, Zhixin Mai, Yiming Zhao, Yatong Han
Inverse kinematics (IK) remains a critical bottleneck for real-time robot manipulation. Classical numerical solvers achieve high geometric precision but often suffer from discontinuous branch switching and unstable behavior near kinematic singularities during closed-loop deployment. Meanwhile, learned IK approaches frequently struggle to balance spatial accuracy, motion smoothness, and real-time efficiency, particularly when trained on noisy human teleoperation data. We present \textbf{MimicIK}, a real-time generative inverse kinematics framework that learns smooth and robust joint-space motion priors from teleoperation demonstrations through conditional flow matching. Given the current joint configuration and a target end-effector pose, MimicIK predicts continuous delta-joint commands using an efficient two-step iterative refinement process based on a Minimal Iterative Policy (MIP) backbone. To enforce physical consistency, we further introduce an FK consistency loss, a differentiable forward-kinematics regularization that penalizes task-space deviations from the target pose during training. We evaluate MimicIK on a real-world 6-DOF robot dataset containing 8,848 teleoperation demonstrations. MimicIK achieves a mean position error of 4.65 mm, a 10 mm success rate of 92.01\%, and a trajectory spike rate of only 7.99\%. Compared with a UNet diffusion baseline, our method improves both spatial accuracy and motion smoothness while reducing inference latency from 21.66 ms to 6.74 ms. Furthermore, unlike deterministic MLP baselines that catastrophically diverge under out-of-distribution deployment, MimicIK remains stable near singular configurations and enables robust 20 Hz real-time control on deployment hardware.
逆运动学(IK)仍是实时机器人操作的关键瓶颈。经典数值求解器虽能达到较高的几何精度,但在闭环部署中常面临分支切换不连续及运动学奇异点附近行为不稳定的问题。同时,学习型IK方法在平衡空间精度、运动平滑性和实时效率方面存在困难,尤其是在基于噪声较大的遥操作人类数据进行训练时。我们提出了MimicIK,这是一个实时生成式逆运动学框架,通过条件流匹配从遥操作演示中学习平滑且稳健的关节空间运动先验。给定当前关节构型和目标末端执行器姿态,MimicIK使用基于最小迭代策略(MIP)主干的高效两阶段迭代细化过程预测连续的关节增量命令。为确保物理一致性,我们进一步引入了FK一致性损失,这是一种可微分的正向运动学正则化方法,在训练过程中惩罚任务空间与目标姿态的偏差。我们在包含8,848个遥操作演示的真实世界6-DOF机器人数据集上对MimicIK进行了评估。MimicIK实现了4.65 mm的平均位置误差、92.01%的10 mm成功率,以及仅7.99%的轨迹尖峰率。与UNet扩散基线相比,我们的方法在空间精度和运动平滑性方面均有提升,同时将推理延迟从21.66 ms降低至6.74 ms。此外,与确定性MLP基线在分布外部署时会出现灾难性发散不同,MimicIK在奇异构型附近保持稳定,并能在部署硬件上实现稳健的20 Hz实时控制。
SierpinskiCam: Camera-Controlled Video Retaking with Sierpinski Triangle Pattern Cues
中文标题:SierpinskiCam:基于谢尔宾斯基三角形模式线索的相机控制视频重制
作者:Suttisak Wizadwongsa, Hyelin Nam, Supasorn Suwajanakorn, Jeong Joon Park
Generating novel renderings of a scene along user-defined camera trajectories from a single monocular video, dubbed video retaking, is a compelling but difficult problem in content creation and visual effects. Existing geometry-guided approaches reconstruct a 4D representation from the source video and render it along the target trajectory to condition video diffusion models. However, this guidance degrades as the target camera departs from the source trajectory, leaving newly revealed regions sparse or entirely missing. We propose SierpinskiCam, which addresses this limitation by augmenting geometry-based guidance with Sierpinski dome texture cues that contains rich trackable features even under large viewpoint changes. We further introduce a reference video conditioning mechanism that appends source-video tokens to the target-token sequence and separates the two streams with negative RoPE indices, enabling appearance grounding without architectural modification or per-video adaptation. Extensive experiments show that SierpinskiCam achieves significant gains in camera controllability, geometric consistency, and video quality across diverse and challenging retaking scenarios. Project page: https://hyelinnam.github.io/SierpinskiCam/.
从单目视频沿用户定义的相机轨迹生成场景的新渲染(称为视频重制)是内容创作和视觉特效领域中一项引人注目但极具挑战性的问题。现有的几何引导方法从源视频重建4D表示,并沿目标轨迹进行渲染以条件化视频扩散模型。然而,随着目标相机偏离源轨迹,这种引导效果会下降,导致新显现的区域稀疏或完全缺失。我们提出SierpinskiCam,通过将基于几何的引导与谢尔宾斯基圆顶纹理线索相结合来解决这一限制,这些线索即使在大幅视角变化下仍包含丰富的可追踪特征。我们进一步引入了一种参考视频条件机制,将源视频标记追加到目标标记序列中,并使用负RoPE索引分隔这两个流,从而无需修改架构或进行逐视频适配即可实现外观锚定。大量实验表明,SierpinskiCam在多样化和具有挑战性的重制场景中在相机可控性、几何一致性和视频质量方面取得了显著提升。项目主页:https://hyelinnam.github.io/SierpinskiCam/.
Learning a Maximum Entropy Model for Visual Textures using Diffusion
中文标题:基于扩散学习的视觉纹理最大熵模型
作者:Xinyuan Zhao, Eero P. Simoncelli
Visual textures -- spatially homogeneous image regions containing repeated elements (e.g. a field of grass, the bark of a tree) -- are ubiquitous in visual scenes and provide important cues for recognizing and analyzing materials and objects. A number of existing texture models extract essential statistics from a single texture image, and can then generate high-quality samples that are visually similar to the original by matching these statistics. However, their statistics are either hand-designed or based on a network pretrained for another purpose (e.g., object recognition). Here, we develop the first principled method for unsupervised learning of a set of statistics that are used to constrain a maximum entropy probability model. We leverage methods developed for generative diffusion models to derive training and sampling procedures, and compare these to the traditional method of sampling via matching the statistics. Despite the compactness of our trained model (512 statistics), it generates texture images whose quality is as good as or better than the current state-of-the-art model (~177k statistics). A more direct comparison of the two models, obtained by synthesizing images that are indistinguishable for one model but maximally different for the other, reveals their relative strengths and weaknesses. Finally, we show that unlike previous statistical texture models, a straight trajectory in the representation space of our model generates homogeneous texture samples that interpolate smoothly between the features of the two end points.
视觉纹理——包含重复元素的空间同质图像区域(如草地、树皮)——在视觉场景中无处不在,为识别和分析材质与物体提供了重要线索。许多现有的纹理模型从单幅纹理图像中提取关键统计量,随后通过匹配这些统计量来生成与原图视觉相似的高质量样本。然而,这些统计量要么是人工设计的,要么是基于为其他目的(如目标识别)预训练的网络。本文提出了首个基于原则性的无监督学习方法,用于学习一组用于约束最大熵概率模型的统计量。我们利用生成式扩散模型中开发的方法推导出训练和采样程序,并将这些方法与传统的通过匹配统计量进行采样的方法进行比较。尽管我们训练得到的模型更加紧凑(512个统计量),但它生成的纹理图像质量与当前最先进模型(约177k个统计量)相当甚至更好。通过合成对一种模型无法区分但对另一种模型差异最大的图像,我们对两种模型进行了更直接的比较,揭示了它们的相对优势和劣势。最后,我们的研究表明,与以往统计纹理模型不同,我们模型表示空间中的直线路径可以生成同质纹理样本,并在两个端点的特征之间实现平滑插值。
Visual Retrieval-Augmented Generation for Silhouette-Guided Animal Art
中文标题:用于剪影引导动物艺术的视觉检索增强生成
作者:Quoc-Duy Tran, Anh-Tuan Vo, Trung-Nghia Le
Generative AI has advanced the ability to render photorealistic or artistic images, yet it remains limited in a key aspect of human creativity: interpreting ambiguous shapes. This phenomenon, rooted in pareidolia, allows humans to perceive meaningful forms in random patterns such as clouds, stones, or leaves. To computationally replicate this imaginative process, we introduce Visual Retrieval-Augmented Generation (Visual-RAG), a framework that generates animal art directly from natural silhouettes. Our method retrieves structurally similar animal shapes from a curated corpus of 28,586 high-quality silhouettes and uses them as reference exemplars to guide diffusion-based generation with ControlNet and IP-Adapter. Ablation studies confirm that shape Context with RANSAC provides the most accurate alignment, while removing shape standardization reduces the inlier ratio to just 13.4\%, underscoring the importance of structural fidelity in Visual-RAG. A user study with 12 participants evaluated the outputs in terms of aesthetics, silhouette fidelity, and overall impression. Results reveal that while Visual-RAG provides plausible interpretations, challenges remain in achieving high perceptual impact. This work lays the foundation for computational pareidolia, showing how machines can contribute to the early stages of imaginative discovery.
生成式AI在渲染逼真或艺术图像方面已取得显著进展,但在人类创造力的一个关键方面仍存在局限:解读模糊形状的能力。这一现象源于错视现象(pareidolia),即人类能够在云朵、石头或树叶等随机图案中感知到有意义的形态。为了在计算层面复现这一想象过程,我们提出了视觉检索增强生成(Visual Retrieval-Augmented Generation,Visual-RAG)框架,能够直接从自然剪影生成动物艺术。我们的方法从包含28,586个高质量剪影的精选语料库中检索结构相似的动物形状,并将其作为参考示例,通过ControlNet和IP-Adapter引导扩散模型生成。消融实验证实,带有RANSAC的形状上下文提供了最准确的对齐效果,而去除形状标准化将内点比率降低至仅13.4%,凸显了结构保真度在视觉检索增强生成中的重要性。一项包含12名参与者的用户研究从美学、剪影保真度和整体印象三个维度对输出进行了评估。结果表明,虽然视觉检索增强生成能够提供合理的解读,但在实现高感知冲击力方面仍面临挑战。本工作为计算错视现象奠定了基础,展示了机器如何为想象发现的早期阶段做出贡献。
Test-Time Training for Robust Text-Guided Open-Vocabulary Object Counting
中文标题:用于鲁棒文本引导开放词汇目标计数的测试时训练
作者:Hao-Yuan Ma, Yuda Zou, Li Zhang, Yongchao Xu
Text-guided Open-vocabulary Object Counting (TOOC) enables counting arbitrary object categories specified by text prompts, offering substantially greater flexibility than conventional closed-set counting. However, existing TOOC methods are developed and evaluated primarily on ideal images, while real-world scenes often suffer from adverse conditions such as rain, fog, darkness, and sensor noise, which severely degrade visual quality and impair vision-language alignment. To bridge this gap, we introduce Robust-TOOC, the first benchmark for evaluating TOOC under diverse corruption conditions, which covers six representative degradation types: rain, fog, darkness, Gaussian noise, salt-and-pepper noise, and mixed corruption. To improve robustness while preserving the original counting architecture, we propose Dual-TTT, a dual-architecture test-time training framework for TOOC. Specifically, during test-time training, Dual-TTT updates only the Text-guided Lightweight Denoising module (TL-Denoiser), while keeping the original counting network frozen. Inspired by diffusion models, the TL-Denoiser is optimized to remove corruption-aware noise from image representations under degraded conditions. Since only the TL-Denoiser is trained at test time, Dual-TTT is annotation-free and can be seamlessly integrated into existing TOOC models without modifying their original architecture. Extensive experiments on multiple recent TOOC baselines demonstrate the effectiveness of our method.
文本引导的开放词汇目标计数(TOOC)能够对文本提示指定的任意目标类别进行计数,相比传统封闭集计数提供了更大的灵活性。然而,现有的TOOC方法主要在理想图像上开发和评估,而现实场景经常遭受雨、雾、黑暗和传感器噪声等恶劣条件,这些会严重降低视觉质量并损害视觉-语言对齐。为弥补这一差距,我们引入了Robust-TOOC,这是首个在多种退化条件下评估TOOC的基准,涵盖六种代表性退化类型:雨、雾、黑暗、高斯噪声、椒盐噪声和混合退化。为在保持原始计数架构的同时提高鲁棒性,我们提出了Dual-TTT,这是一种用于TOOC的双架构测试时训练框架。具体而言,在测试时训练期间,Dual-TTT仅更新文本引导轻量级去噪模块(TL-Denoiser),同时冻结原始计数网络。受扩散模型启发,TL-Denoiser经过优化,可在退化条件下从图像表示中去除污染感知噪声。由于仅在测试时训练TL-Denoiser,Dual-TTT无需标注,可无缝集成到现有TOOC模型中,无需修改其原始架构。在多个最新TOOC基线模型上的广泛实验证明了该方法的有效性。
Flux-Guard: Facial Identity Protection using diffusion models
中文标题:Flux-Guard:基于扩散模型的人脸身份保护方法
作者:Jie Wang, Tao Wang, Ru Zhang, Jianyi Liu
The widespread deployment of face recognition (FR) systems exposes personal images shared on social media and public platforms to identity linkage and privacy risks. Existing adversarial privacy protection methods can degrade unauthorized FR performance but are not compatible with generative face editing. Artificial intelligence-driven face editing tools are gaining popularity, which has significantly increased user demand for personalized portrait generation and social sharing. However, current editing methods often preserve identity features, making the edited images still susceptible to tracking by malicious FR systems. Thus, this paper proposes Flux-Guard, a privacy-preserving face editing framework based on adversarial attacks, which integrates face editing and privacy protection within a unified generative process. Specifically, we design a flow trajectory control method to align semantic manipulations with the generative process and introduce latent-space adversarial optimization with an adaptive perceptual-loss-driven weighting strategy, dynamically adjusting adversarial strength to maximize attack effectiveness while preserving visual quality. Extensive experiments demonstrate that Flux-Guard supports face editing while significantly improving attack success rates against cross-domain face recognition models on the CelebA-HQ and LADN datasets. Furthermore, evaluation results for commercial APIs have confirmed its effectiveness in real-world applications. The code is released at https://github.com/JLMWang/Flux-Guard.
人脸识别系统的广泛应用使得社交媒体和公共平台上分享的个人图像面临身份关联和隐私泄露风险。现有的对抗性隐私保护方法虽能降低未授权人脸识别的性能,但与生成式人脸编辑不兼容。由人工智能驱动的人脸编辑工具正日益普及,显著增加了用户对个性化肖像生成和社交分享的需求。然而,当前的编辑方法往往保留了身份特征,使得编辑后的图像仍易受到恶意人脸识别系统的追踪。因此,本文提出Flux-Guard,一种基于对抗攻击的隐私保护人脸编辑框架,将人脸编辑和隐私保护集成到统一的生成过程中。具体而言,我们设计了一种流轨迹控制方法,使语义操作与生成过程对齐,并引入基于自适应感知损失加权策略的潜空间对抗优化方法,动态调整对抗强度以在保持视觉质量的同时最大化攻击效果。大量实验表明,Flux-Guard在支持人脸编辑的同时,显著提高了针对CelebA-HQ和LADN数据集上跨域人脸识别模型的攻击成功率。此外,对商业API的评估结果也证实了其在实际应用中的有效性。代码已发布于https://github.com/JLMWang/Flux-Guard。
MOCHI: Motion Enhancement of Collaborative Human-object Interactions
中文标题:MOCHI:协作式人-物体交互的运动增强
作者:Jiye Lee, Yonghun Choi, Jungdam Won
Collaborative human-object interaction shows dynamic and complex movements that require mutual anticipation and continuous adjustment between participants and the shared object. Modeling such collaborative multi-human object interaction (MHOI) scenarios requires high-quality data acquisition as a foundational step; however, this is challenging due to the inherent complexity of MHOI where human-human and human-object interactions occur simultaneously. Such complexity leads to noisy MHOI captures characterized by several artifacts: contact misalignment between hands and objects, motion jitter and temporal inconsistencies in the captured sequences, and missing or incomplete finger-level articulation details. To address these challenges, we present MOCHI (MOtion Enhancement of Collaborative Human-object Interactions), a two-stage framework for enhancing noisy MHOI data. Our approach first generates physically plausible hand grasps through optimization from noisy body input, producing grasps that are both physically plausible and semantically consistent with the body pose, where these optimized grasps are extended into complete hand-object interaction sequences. Consequently, the full-body motion for all participants are refined through a diffusion-based noise optimization framework that uses single-person motion priors. During the optimization process, we introduce optimization objectives to encode human-object and human-human interaction information within these single-person priors. Experimental results demonstrate the effectiveness of our pipeline across diverse MHOI data, either acquired by existing capture methods or synthesized by generative models. We further show robustness of our system across varying numbers of participants and types of interactions, and demonstrate various applications including keyframe-based MHOI creation and data augmentation through varying object geometries.
协作式人-物体交互表现出动态且复杂的运动,需要参与者与共享物体之间的相互预测和持续调整。对此类协作式多人人-物体交互(MHOI)场景进行建模需要以高质量数据采集作为基础步骤;然而,由于MHOI inherently复杂——人-人交互和人-物体交互同时发生——数据采集面临巨大挑战。这种复杂性导致获取的MHOI数据存在噪声,表现为多种伪影:手部与物体之间的接触错位、捕获序列中的运动抖动和时间不一致性,以及手指级关节细节的缺失或不完整。为解决这些挑战,我们提出了MOCHI(MOtion Enhancement of Collaborative Human-object Interactions,协作式人-物体交互的运动增强),一个用于增强噪声MHOI数据的双阶段框架。我们的方法首先通过从噪声身体输入进行优化来生成物理上可信的手部抓取,产生既物理上可信又与身体姿态语义一致的抓取,并将这些优化后的抓取扩展为完整的手-物体交互序列。随后,通过使用单人运动先验的基于扩散的噪声优化框架来优化所有参与者的全身运动。在优化过程中,我们引入优化目标以将这些单人先验中编码的人-物体和人-人交互信息融合进去。实验结果表明,我们的方法在多样的MHOI数据上具有有效性,这些数据既可以通过现有捕获方法获取,也可以通过生成模型合成。我们进一步展示了系统在不同数量的参与者和交互类型下的鲁棒性,并展示了多种应用场景,包括基于关键帧的MHOI创建和通过改变物体几何形状进行的数据增强。
Unified Multimodal Autoregressive Modeling with Shared Context-Visual Tokenizer is Key to Unification
中文标题:使用共享上下文-视觉分词器的统一多模态自回归建模是实现统一的关键
作者:Wujian Peng, Lingchen Meng, Yuxuan Cai, Xianwei Zhuang, Yuhuan Yang, Rongyao Fang, Chenfei Wu, Junyang Lin, Zuxuan Wu, Shuai Bai
Unified Multimodal Modeling aims to integrate visual understanding and generation within a single system. However, existing approaches typically rely on two disparate visual tokenizers, which splits the representation space and hinders truly unified modeling. We propose UniAR, a unified autoregressive framework where a single discrete visual tokenizer serves as the key bridge between understanding and generation, enabling a shared context in which the model can directly interpret its own generated visual tokens without additional re-encoding. UniAR adapts a pretrained vision encoder with multi-level feature fusion and a lookup-free bitwise quantization scheme, preserving both high-level semantics and low-level details while scaling the effective visual vocabulary at minimal cost. Building on this, the unified autoregressive model adopts parallel-bitwise-prediction to jointly predict spatially grouped, multi-level visual codes, substantially reducing visual sequence length and accelerating generation. Finally, a diffusion-based visual decoder operates on discrete visual tokens to decode high-fidelity images. Through large-scale pre-training, followed by supervised fine-tuning and reinforcement learning, UniAR achieves state-of-the-art performance on image generation and image editing while remaining competitive on multimodal understanding benchmarks. The project page is available at https://sharelab-sii.github.io/uniar-web.
统一多模态建模旨在将视觉理解和生成集成在单一系统中。然而,现有方法通常依赖两个不同的视觉分词器,这导致表示空间分裂,阻碍了真正的统一建模。我们提出UniAR,这是一个统一的自回归框架,其中单一离散视觉分词器作为理解和生成之间的关键桥梁,使模型能够在共享上下文中直接解释自身生成的视觉标记,无需额外的重新编码。UniAR采用预训练视觉编码器配合多级特征融合和无查找表位量化方案,在以最低成本扩展有效视觉词汇的同时保留高级语义和低细节。在此基础上,统一自回归模型采用并行位预测来联合预测空间分组的多级视觉编码,大幅缩短视觉序列长度并加速生成。最后,基于扩散的视觉解码器对离散视觉标记进行操作以解码高保真图像。通过大规模预训练,随后进行监督微调和强化学习,UniAR在图像生成和图像编辑方面实现了最先进的性能,同时在多模态理解基准测试中保持竞争力。项目页面见 https://sharelab-sii.github.io/uniar-web。
Where Should Action Generation Begin? A Learnable Source Prior for Generative Robot Policies
中文标题:动作生成应从何始?——一种用于生成式机器人策略的可学习源先验方法
作者:Meipo Dai, Qiyuan Zhuang, He-Yang Xu, Ying-Jie Shuai, Yijun Wang, Qi Dou, Xiu-Shen Wei
Generative robot policies typically begin action generation from an observation-independent standard Gaussian distribution, leaving the choice of source distribution underexplored. This work asks a simple question: where should action generation begin? We propose LeaP, a Learnable source Prior that replaces the standard Gaussian with a proprioception-conditioned diagonal Gaussian over action chunks. Parameterized by a lightweight MLP, LeaP jointly predicts the mean and state-adaptive variance of the source distribution, while keeping the downstream generator architecture and inference solver unchanged. This design provides an observation-informed yet stochastic initialization, allowing the generator to focus on precise action refinement rather than transporting samples from an uninformed noise source. On 15 RoboTwin manipulation tasks, LeaP achieves an average success rate of 81.6%, outperforming four representative baselines -- including deterministic-source methods, a no-prior counterpart, and a diffusion-bridge policy -- by 6.5 to 25.5 percentage points. The same prior consistently improves both flow-matching and diffusion-bridge generators, while using fewer parameters and converging faster. The advantage carries over to real-world deployment, where LeaP attains the best performance. These results suggest that the source distribution is an independent and reusable design axis for generative robot policies, complementary to the choice of generative dynamics.
生成式机器人策略通常从与观测无关的标准高斯分布开始动作生成,而源分布的选择却未被充分探索。本工作提出一个简单的问题:动作生成应该从哪里开始?本研究提出LeaP(可学习源先验),用一个基于本体感知调节的动作块对角高斯分布取代标准高斯分布。LeaP由一个轻量级MLP参数化,可联合预测源分布的均值和状态自适应方差,同时保持下游生成器架构和推理求解器不变。这种设计提供了一种观测知情但随机的初始化方式,使生成器能够专注于精确的动作细化,而非从无信息噪声源传输样本。在15个RoboTwin操作任务上,LeaP实现了81.6%的平均成功率,在四项代表性基线方法上(包括确定性源方法、无先验对应方法以及扩散桥策略)提升了6.5至25.5个百分点。同样的先验方法持续改善了流匹配和扩散桥生成器的性能,同时使用更少的参数并收敛更快。该优势可迁移至真实世界部署,LeaP取得了最佳表现。这些结果表明,源分布是生成式机器人策略的一个独立且可复用的设计轴,可与生成式动力学模型的选择形成互补。
Edit3DGS: Unified Framework for Dynamic Head Editing via 2D Instruction-Guided Diffusion and 3D Gaussian Splatting
中文标题:Edit3DGS:基于2D指令引导扩散和3D高斯溅射的动态头部编辑统一框架
作者:Duy-Dat Tran, Trung-Nghia Le
We present Edit3DGS, a unified framework for dynamic 3D head editing that integrates 2D instruction-guided diffusion with 3D Gaussian splatting. Unlike prior approaches that separately address frame-based edits or static 3D reconstruction, our method couples semantic controllability in the image domain with photorealistic, temporally consistent 3D representations. Given an input video, editable facial regions are masked and modified using a text-conditioned diffusion model to support fine-grained operations such as expression transformation, attribute modification, and appearance refinement. The edited frames are then aggregated through 3D Gaussian splatting to produce a coherent, high-fidelity avatar that preserves both identity and motion dynamics. To enforce consistency, Edit3DGS incorporates multi-view batch editing and lightweight inpainting strategies that recover lost expressions across timesteps. Experimental results demonstrate that our framework enables controllable, artifact-free head editing with smooth temporal transitions, offering practical applications in virtual avatars, immersive communication, film production, and interactive media.
我们提出了Edit3DGS,一个将2D指令引导扩散与3D高斯溅射相融合的动态3D头部编辑统一框架。与先前单独处理基于帧的编辑或静态3D重建的方法不同,本方法将图像域中的语义可控性与逼真且时间一致的3D表示相结合。给定输入视频,可编辑的面部区域被蒙版并使用文本条件扩散模型进行修改,以支持表情变换、属性修改和外观细化等细粒度操作。随后,通过3D高斯溅射聚合编辑后的帧,生成一致、高保真的虚拟化身,同时保留身份信息和运动动态。为确保一致性,Edit3DGS引入了多视角批量编辑和轻量级修复策略,以跨时间步恢复丢失的表情。实验结果表明,本框架实现了可控、无伪影的头部编辑和平滑的时间过渡,在虚拟化身、沉浸式通信、电影制作和交互媒体等领域具有实际应用价值。
SceneCompleter: Dense 3D Scene Completion for Generative Novel View Synthesis
中文标题:SceneCompleter:用于生成式新视角合成的密集3D场景补全
作者:Weiliang Chen, Jiayi Bi, Yuanhui Huang, Wenzhao Zheng, Yueqi Duan
Generative models have shown great promise for novel view synthesis (NVS) by leveraging strong image generation priors. However, existing approaches typically follow a 2D inpainting paradigm, first completing missing image regions and then performing 3D reconstruction. This strategy often causes geometry distortion and appearance drift, as 2D inpainting models cannot reliably infer the underlying 3D structure required for cross-view consistent generation. In this paper, we propose \textbf{SceneCompleter}, a geometry-aware framework that reformulates generative NVS as dense 3D scene completion. Instead of hallucinating isolated 2D views, SceneCompleter jointly completes geometry and appearance through a geometry-appearance dual-stream diffusion model in a spatially aligned RGBD latent space. To provide holistic scene context, we further introduce a Scene Embedder that conditions generation on global semantic and stylistic information from reference images. The completed RGBD predictions are then aligned and integrated into an expandable 3D scene representation, enabling iterative and coherent scene completion. Extensive experiments on in-domain and out-of-distribution datasets demonstrate that SceneCompleter produces visually plausible and geometrically consistent novel views across diverse scenarios. Project Page: https://chen-wl20.github.io/SceneCompleter
生成式模型通过利用强大的图像生成先验在新视角合成(NVS)中展现出巨大潜力。然而,现有方法通常遵循2D修复范式,先完成缺失图像区域的修复,再进行3D重建。这种策略往往导致几何失真和外观漂移,因为2D修复模型无法可靠地推断跨视角一致生成所需的底层3D结构。本文提出SceneCompleter,一个几何感知的框架,将生成式NVS重新定义为密集3D场景补全。SceneCompleter不再幻想孤立的2D视图,而是通过几何-外观双流扩散模型在空间对齐的RGBD潜在空间中联合完成几何和外观的补全。为提供全面的场景上下文,我们进一步引入Scene Embedder,利用参考图像的全局语义和风格信息来调节生成过程。完成的RGBD预测随后被对齐并整合到一个可扩展的3D场景表示中,实现迭代且一致的场景补全。在域内和分布外数据集上的大量实验表明,SceneCompleter能够在各种场景中产生视觉上合理且几何一致的新视角。项目主页:https://chen-wl20.github.io/SceneCompleter
TextMesh4D: Zero-shot Text-to-4D Mesh Generation
中文标题:TextMesh4D:零样本文本到4D网格生成
作者:Sisi Dai, Xinxin Su, Kai Xu
Large-scale, high-quality dynamic 3D (4D) assets are essential for learning physically grounded representations, but remain costly to capture and annotate at scale. This limits the viability of supervised 4D learning and motivates zero-shot text-to-4D generation leveraging pretrained diffusion priors. To model complex dynamics, prior methods typically adopt implicit 3D representations (e.g., NeRFs or 3DGS) for their deformation capacity. However, their implicit nature provides limited control over surface topology, which hinders high-fidelity geometry and makes temporally coherent surface reconstruction challenging. To address these limitations, we explore zero-shot text-to-4D mesh generation. However, a structural mismatch arises when combining diffusion-based guidance with topology-constrained meshes: the guidance is noisy and spatially inconsistent, while meshes impose severe topological constraints, making direct vertex-level deformation unstable. In this paper, we introduce TextMesh4D, the first zero-shot framework for text-to-4D that directly generates dynamic meshes by addressing the above challenge at two complementary levels. Geometrically, we shift deformation modeling from vertices to faces via a Jacobian Deformation Field (JDF), enabling topology-aware surface reconstruction through an integrability-enforcing integration formulation. Semantically, we propose a Local-Global Semantic Regularizer (LGSR) that preserves identity over time by jointly constraining local deformation plausibility and global shape consistency. Extensive experiments demonstrate state-of-the-art temporal consistency, structural fidelity, and visual quality, while remaining efficient on a single 24GB GPU.
大规模、高质量的动态3D(4D)资产对于学习物理基础表示至关重要,但大规模捕获和标注这些资产成本高昂。这限制了监督式4D学习的可行性,并促使研究转向利用预训练扩散先验的零样本文本到4D生成方法。早期方法通常采用隐式3D表示(如NeRF或3DGS)来建模复杂动态,因为它们具有较强的变形能力。然而,隐式表示对表面拓扑的控制有限,这阻碍了高保真几何的实现,并使得时间一致性表面重建面临挑战。为解决这些局限性,我们探索零样本文本到4D网格生成。然而,当将基于扩散的指导与拓扑约束网格相结合时会出现结构不匹配:扩散指导存在噪声且空间不一致,而网格施加了严格的拓扑约束,使得直接进行顶点级变形变得不稳定。本文提出TextMesh4D,这是首个零样本文本到4D框架,通过在两个互补层面解决上述挑战来直接生成动态网格。在几何层面,我们将变形建模从顶点转移到面,提出了雅可比变形场(JDF),通过可积性强制积分公式实现拓扑感知的表面重建。在语义层面,我们提出了局部-全局语义正则化器(LGSR),通过联合约束局部变形合理性和全局形状一致性来保持时间身份一致性。大量实验表明,本方法在时间一致性、结构保真度和视觉质量方面达到了最先进水平,同时在单张24GB GPU上保持高效运行。
4DSloMo: 4D Reconstruction for High Speed Scene with Asynchronous Capture
中文标题:4DSloMo:基于异步采集的高速场景4D重建
作者:Yutian Chen, Shi Guo, Tianshuo Yang, Lihe Ding, Xiuyuan Yu, Jinwei Gu, Tianfan Xue
Reconstructing fast-dynamic scenes from multi-view videos is crucial for high-speed motion analysis and realistic 4D reconstruction. However, the majority of 4D capture systems are limited to frame rates below 30 FPS (frames per second), and a direct 4D reconstruction of high-speed motion from low FPS input may lead to undesirable results. In this work, we propose a high-speed 4D capturing system only using low FPS cameras, through novel capturing and processing modules. On the capturing side, we propose an asynchronous capture scheme that increases the effective frame rate by staggering the start times of cameras. By grouping cameras and leveraging a base frame rate of 25 FPS, our method achieves an equivalent frame rate of 100-200 FPS without requiring specialized high-speed cameras. On processing side, we also propose a novel generative model to fix artifacts caused by 4D sparse-view reconstruction, as asynchrony reduces the number of viewpoints at each timestamp. Specifically, we propose to train a video-diffusion-based artifact-fix model for sparse 4D reconstruction, which refines missing details, maintains temporal consistency, and improves overall reconstruction quality. Experimental results demonstrate that our method significantly enhances high-speed 4D reconstruction compared to synchronous capture.
从多视角视频重建快速动态场景对于高速运动分析和真实感4D重建至关重要。然而,大多数4D采集系统的帧率限制在30 FPS以下,直接从低帧率输入进行高速运动的4D重建可能导致不理想的结果。本工作提出了一种仅使用低帧率相机的高速4D采集系统,通过创新的采集和处理模块实现。在采集方面,我们提出了一种异步采集方案,通过错开相机的启动时间来有效提高帧率。通过相机分组并利用25 FPS的基础帧率,我们的方法实现了100-200 FPS的等效帧率,无需使用专业高速相机。在处理方面,我们还提出了一种新型生成模型来修复由4D稀疏视角重建产生的伪影,因为异步性减少了每个时间点的视角数量。具体而言,我们提出训练一个基于视频扩散的伪影修复模型用于稀疏4D重建,该模型能够完善缺失细节、保持时间一致性并提升整体重建质量。实验结果表明,我们的方法相较于同步采集显著提升了高速4D重建效果。
FUSER: Feed-Forward MUltiview 3D Registration Transformer and SE(3)$^N$ Diffusion Refinement
中文标题:FUSER:前馈多视图三维配准Transformer与SE(3)^N扩散细化
作者:Haobo Jiang, Jin Xie, Jian Yang, Liang Yu, Jianmin Zheng
Registration of multiview point clouds conventionally relies on extensive pairwise matching to build a pose graph for global synchronization, which is computationally expensive and inherently ill-posed without holistic geometric constraints. This paper proposes FUSER, the first feed-forward multiview registration transformer that jointly processes all scans in a unified, compact latent space to directly predict global poses without any pairwise estimation. To maintain tractability, FUSER encodes each scan into low-resolution superpoint features via a sparse 3D CNN that preserves absolute translation cues, and performs efficient intra- and inter-scan reasoning through a Geometric Alternating Attention module. Particularly, we transfer 2D attention priors from off-the-shelf foundation models to enhance 3D feature interaction and geometric consistency. Building upon FUSER, we further introduce FUSER-DF, an SE(3)$^N$ diffusion refinement framework to correct FUSER's estimates via denoising in the joint SE(3)$^N$ space. FUSER acts as a surrogate multiview registration model to construct the denoiser, and a prior-conditioned SE(3)$^N$ variational lower bound is derived for denoising supervision. Extensive experiments on 3DMatch, ScanNet and ArkitScenes demonstrate that our approach achieves the superior registration accuracy and outstanding computational efficiency.
多视图点云配准传统上依赖于广泛的成对匹配来构建姿态图以进行全局同步,这在计算上非常昂贵,且在没有整体几何约束的情况下本质上是不适定的问题。本文提出FUSER,这是首个前馈多视图配准变换器,它在统一、紧凑的潜在空间中联合处理所有扫描,直接预测全局姿态,无需任何成对估计。为保持可处理性,FUSER将每个扫描编码为低分辨率超点特征,通过稀疏3D卷积神经网络保留绝对平移线索,并通过几何交替注意力模块执行高效的扫描内和扫描间推理。特别地,我们从现成的基础模型中迁移2D注意力先验,以增强3D特征交互和几何一致性。在此基础上,我们进一步引入FUSER-DF,这是一个SE(3)^N扩散细化框架,通过在联合SE(3)^N空间中的去噪来纠正FUSER的估计值。FUSER作为代理多视图配准模型来构建去噪器,并推导了基于先验条件的SE(3)^N变分下界用于去噪监督。在3DMatch、ScanNet和ArkitScenes上的大量实验表明,我们的方法实现了卓越的配准精度和出色的计算效率。
RAIGen: Rare Attribute Identification in Text-to-Image Generative Models
中文标题:RAIGen:文本到图像生成模型中的稀有属性识别
作者:Silpa Vadakkeeveetil Sreelatha, Dan Wang, Serge Belongie, Muhammad Awais, Anjan Dutta
Text-to-image diffusion models achieve impressive generation quality but inherit and amplify training-data biases, skewing coverage of semantic attributes. Prior work addresses this in two ways. Closed-set approaches mitigate biases in predefined fairness categories (e.g., gender, race), assuming socially salient minority attributes are known a priori. Open-set approaches frame the task as bias identification, highlighting majority attributes that dominate outputs. Both overlook a complementary task: uncovering rare or minority features underrepresented in the data distribution (social, cultural, or stylistic) yet still encoded in model representations. We introduce RAIGen, the first framework, to our knowledge, for label-free rare-attribute discovery in diffusion models, requiring no predefined minority categories. RAIGen leverages Matryoshka Sparse Autoencoders and a novel minority metric combining neuron activation frequency with semantic distinctiveness to identify interpretable neurons whose top-activating images reveal underrepresented attributes. Experiments show RAIGen discovers attributes beyond fixed fairness categories in Stable Diffusion, scales to larger models such as SDXL, supports systematic auditing across architectures, and enables targeted amplification of rare attributes during generation. The project page is available at https://vssilpa.github.io/RAIGen_webpage/ .
文本到图像扩散模型取得了令人印象深刻的生成质量,但继承并放大了训练数据偏差,使语义属性的覆盖范围出现偏斜。现有工作通过两种方式解决这一问题。封闭集方法缓解预定义公平性类别(如性别、种族)中的偏见,假设社会显著的少数属性是先验已知的。开放集方法将任务框架化为偏见识别,突出主导输出的多数属性。两者都忽略了一个互补任务:揭示在数据分布中代表不足但在模型表示中仍然被编码的稀有或少数特征(社会的、文化的或风格的)。我们提出RAIGen,据我们所知,这是扩散模型中首个无需标签的稀有属性发现框架,无需预定义的少数类别。RAIGen利用套娃稀疏自编码器(Matryoshka Sparse Autoencoders)和一种新的少数指标,结合神经元激活频率与语义独特性,以识别可解释的神经元,其最高激活图像揭示了代表不足的属性。实验表明,RAIGen在Stable Diffusion中发现了超出固定公平性类别的属性,可扩展至更大的模型如SDXL,支持跨架构的系统审计,并能在生成过程中对稀有属性进行有针对性的增强。项目主页见https://vssilpa.github.io/RAIGen_webpage/。
CASR: A Robust Cyclic Framework for Arbitrary Large-Scale Super-Resolution with Distribution Alignment and Self-Similarity Awareness
中文标题:CASR:一种面向任意大尺度超分辨率的分布对齐与自相似性感知鲁棒循环框架
作者:Wenhao Guo, Zhaoran Zhao, Peng Lu, Sheng Li, Qian Qiao, DeRui Li
Arbitrary-Scale SR (ASISR) remains fundamentally limited by cross-scale distribution shift: once the inference scale leaves the training range, noise, blur, and artifacts accumulate sharply. We revisit this challenge from a cross-scale distribution transition perspective and propose CASR, a simple yet highly efficient cyclic SR framework that reformulates ultra-magnification as a sequence of in-distribution scale transitions. This design ensures stable inference at arbitrary scales while requiring only a single model. CASR tackles two major bottlenecks: distribution drift across iterations and patch-wise diffusion inconsistencies. The proposed SSAM module aligns structural distributions via superpixel aggregation, preventing error accumulation, while SARM module restores high-frequency textures by enforcing correlation-guided consistency and preserving self-similarity structure through correlation alignment. Despite using only a single model, our approach significantly reduces distribution drift, preserves long-range texture consistency, and achieves superior generalization even at extreme magnification.
任意尺度超分辨率(ASISR)本质上受限于跨尺度分布偏移问题:一旦推理尺度超出训练范围,噪声、模糊和伪影会急剧累积。我们从跨尺度分布转换的视角重新审视这一挑战,并提出CASR——一种简洁而高效的循环超分辨率框架,将超大规模放大重新表述为一系列分布内尺度转换。该设计确保了在任意尺度下的稳定推理,同时仅需单一模型即可实现。CASR解决了两个主要瓶颈:迭代过程中的分布漂移和块状扩散不一致性问题。提出的SSAM模块通过超像素聚合对齐结构分布,防止误差累积;而SARM模块则通过相关性引导的一致性约束和相关性对齐保留自相似性结构,从而恢复高频纹理。尽管仅使用单一模型,本方法显著降低了分布漂移,保持了远距离纹理一致性,并在极端放大倍数下实现了卓越的泛化能力。
DVD: Discrete Voxel Diffusion for 3D Generation and Editing
中文标题:DVD:用于3D生成与编辑的离散体素扩散
作者:Zhengrui Xiang, Jiaqi Wu, Fupeng Sun, Heliang Zheng, Yingzhen Li
We introduce Discrete Voxel Diffusion (DVD), a discrete diffusion framework to generate, assess, and edit sparse voxels for SLat (Structured LATent) based 3D generative pipelines. Although discrete diffusion has not generally displaced continuous diffusion in image-like generation, we show that it can be an effective first-stage prior for sparse voxel scaffolds. By treating voxel occupancy as a native discrete variable, DVD avoids continuous-to-discrete thresholding and provides a simple framework for voxel generation, uncertainty estimation, and editing. Beyond quality gains, DVD provides more interpretable generation dynamics through explicit categorical modeling. Furthermore, we leverage the predictive entropy as a robust uncertainty metric to identify ambiguous voxel regions and complicated samples, facilitating tasks such as data filtering and quality assessment. Finally, we propose a lightweight fine-tuning strategy using block-structured perturbation patterns. This approach empowers the model to inpaint and edit voxels within a single sampling round, requiring negligible auxiliary computation and no additional model evaluations. Code is available at https://github.com/TeCai/DVD.
我们提出了离散体素扩散(Discrete Voxel Diffusion,DVD),这是一个离散扩散框架,用于为基于SLat(Structured LATent,结构化潜空间)的3D生成管道生成、评估和编辑稀疏体素。尽管离散扩散通常未能取代图像类生成中的连续扩散,但我们表明,它可以作为稀疏体素脚手架的有效先验。通过将体素占用作为原生离散变量,DVD避免了连续到离散的阈值化问题,并提供了一个简洁的体素生成、不确定性估计和编辑框架。除了质量提升外,DVD通过显式分类建模提供了更易解释的生成动态。此外,我们利用预测熵作为鲁棒的不确定性度量来识别模糊体素区域和复杂样本,从而支持数据过滤和质量评估等任务。最后,我们提出了一种轻量级的微调策略,采用块结构扰动模式。该方法使模型能够在单次采样轮次内完成体素的修补和编辑,且仅需极少的辅助计算,无需额外的模型评估。代码见 https://github.com/TeCai/DVD。
Flash-GRPO: Efficient Alignment for Video Diffusion via One-Step Policy Optimization
中文标题:Flash-GRPO: 通过单步策略优化实现视频扩散模型的高效对齐
作者:Xiaoxuan He, Siming Fu, Zeyue Xue, Weijie Wang, Ruizhe He, Yuming Li, Dacheng Yin, Shuai Dong, Haoyang Huang, Hongfa Wang, Nan Duan, Bohan Zhuang
Group Relative Policy Optimization has emerged as essential for aligning video diffusion models with human preferences, but faces a critical computational bottleneck: training a 14B parametered model typically demands hundreds of GPU days per experiment. Existing efficiency methods reduce costs through sliding window subsampling training timesteps, but fundamentally compromise optimization, exhibiting severe instability and failing to reach full trajectory performance. We present Flash-GRPO, a single-step training framework that outperforms full trajectory training in alignment quality under low computational budgets while substantially improving training efficiency. Flash-GRPO addresses two critical challenges: iso-temporal grouping eliminates timestep-confounded variance by enforcing prompt-wise temporal consistency, decoupling policy performance from timestep difficulty; temporal gradient rectification neutralizes the time-dependent scaling factor that causes vastly inconsistent gradient magnitudes across timesteps. Experiments on 1.3B to 14B parameter models validate Flash-GRPO's effectiveness, demonstrating substantial training acceleration with consistent stability and state-of-the-art alignment quality.
组相对策略优化(GRPO)已成为视频扩散模型与人类偏好对齐的关键技术,但面临严重的计算瓶颈:训练一个140亿参数的模型通常每个实验需要数百个GPU天。现有的效率方法通过滑动窗口子采样训练时间步来降低成本,但从根本上损害了优化过程,表现出严重的不稳定性,无法达到完整轨迹的性能。我们提出了Flash-GRPO,这是一种单步训练框架,在低计算预算下,对齐质量优于完整轨迹训练,同时显著提高了训练效率。Flash-GRPO解决了两个关键挑战:等时间分组通过强制执行提示级时间一致性消除了时间步混淆的方差,将策略性能与时间步难度解耦;时间梯度校正中和了导致不同时间步梯度幅度严重不一致的时间依赖缩放因子。在13亿到140亿参数模型上的实验验证了Flash-GRPO的有效性,展示了显著的训练加速效果,同时保持了一致的稳定性和最先进的对齐质量。
Show the Signal, Hide the Noise: Spectral Forcing for Pixel-Space Diffusion
中文标题:显示信号,隐藏噪声:用于像素空间扩散的谱强制方法
作者:Weichen Fan, Haiwen Diao, Penghao Wu, Ziwei Liu
Pixel-space diffusion models are trained on full-bandwidth noisy images, yet the useful signal available to the denoiser is strongly frequency dependent. Under rectified-flow diffusion and natural-image power-law spectra, the per-band data-to-noise contour $k^{*}(t) = (1-t)^{-2/\alpha}$ separates a signal-bearing low-frequency region from a noise-dominated high-frequency region at each time $t$. We show that this implicit coarse-to-fine structure is not merely descriptive: it induces a capacity-allocation problem. A standard pixel-space denoiser must discover the moving bandwidth boundary internally and can spend computation on frequency-time regions where the optimal prediction collapses to deterministic baselines rather than data-distribution modeling. To make this boundary explicit, we introduce Spectral Forcing, a parameter-free, time-conditional 2D-DCT low-pass operator applied to the noisy input before the patch embedder. Its cutoff expands monotonically with the diffusion time and becomes the identity at the data endpoint. Through controlled synthetic experiments, we identify the regime in which the operator is beneficial: coarse patch tokenization and data whose high-frequency content is predominantly noise rather than essential signal. On ImageNet-256 with JiT-700M/32, Spectral Forcing consistently improves both FID and Inception Score across different training epochs, demonstrating robust gains throughout training; at finer tokenization, the spectral forcing is still competitive. We further insert the unchanged operator into SenseNova-U1, a unified text-to-image model, where it improves DPG-Bench and GenEval, showing that the input-side spectral prior transfers beyond class-conditional generation. These results suggest a route to capacity-efficient pixel-space diffusion by showing the signal and hiding the noise.
像素空间扩散模型在全带宽噪声图像上进行训练,然而去噪器可利用的有效信号与频率密切相关。在整流流扩散和自然图像幂律谱条件下,每个时间步t的每频带数据-噪声轮廓k*(t) = (1-t)^(-2/α)将携带信号的低频区域与噪声主导的高频区域分隔开来。我们表明这种隐含的粗到细结构不仅是描述性的:它还会引发容量分配问题。标准像素空间去噪器必须在内部发现移动的带宽边界,并可能在频域-时间区域上浪费计算资源,其中最优预测退化为确定性基线而非数据分布建模。为了使这一边界显式化,我们引入谱强制(Spectral Forcing),这是一种无参数、时间条件的二维离散余弦变换低通算子,应用于patch嵌入器之前的噪声输入。其截止频率随扩散时间单调扩展,在数据端点处变为恒等算子。通过受控的合成实验,我们确定了该算子有益的 regime:粗粒度patch标记化且数据的高频内容主要为噪声而非本质信号。在ImageNet-256上使用JiT-700M/32,谱强制在不同训练轮次下持续改善FID和Inception Score,展示了训练过程中的稳健增益;在更细粒度的标记化下,谱强制仍具竞争力。我们进一步将不变算子插入SenseNova-U1(一种统一文生图模型),其中它改善了DPG-Bench和GenEval,表明输入端频谱先验可迁移至类别条件生成之外。这些结果通过显示信号和隐藏噪声,为容量高效的像素空间扩散指明了一条路径。
Variational Test-time Optimization for Diffusion Synchronization
中文标题:扩散同步的变分测试时优化
作者:Hyunsoo Lee, Farrin Marouf Sofian, Kushagra Pandey, Stephan Mandt
Collaborative generation, which coordinates multiple diffusion trajectories to extend the capabilities of pretrained priors, has emerged as a powerful paradigm for extending the applicability of diffusion models. Among existing approaches, diffusion synchronization provides a scenario-agnostic solution by introducing general guidance mechanisms. However, current synchronization approaches rely heavily on heuristics and still require task-specific tailoring, which limits their generalizability and performance. In this work, we mathematically derive a synchronization framework based on optimal control, providing a principled explanation of diffusion synchronization. During sampling, we optimize control variables to guide multiple trajectories toward coherent solutions while remaining close to the underlying diffusion prior. Our method operates entirely at test-time without additional training, thereby enabling broad applicability across diverse generation scenarios when combined with strong pretrained priors. We demonstrate consistent improvements over baselines on three representative collaborative generation tasks, covering a wide range of modalities and applications. Beyond performance gains, our work establishes a novel foundation for collaborative generation, opening a principled path toward extending pretrained generative models to new collaborative generation settings.
协同生成通过协调多条扩散轨迹来扩展预训练先验的能力,已成为扩展扩散模型适用性的强大范式。在现有方法中,扩散同步通过引入通用引导机制提供了一种场景无关的解决方案。然而,当前的同步方法严重依赖启发式设计,仍然需要针对具体任务进行调整,这限制了它们的通用性和性能表现。本研究从数学上推导出一个基于最优控制的同步框架,为扩散同步提供了原则性的解释。在采样过程中,我们优化控制变量以引导多条轨迹趋向一致解,同时保持接近底层扩散先验。我们的方法完全在测试时运行,无需额外训练,因此与强大的预训练先验相结合时能够在各种生成场景中实现广泛适用性。我们在三个具有代表性的协同生成任务上展示了相较于基线方法的一致性改进,涵盖了广泛的应用领域和场景。除了性能提升之外,本研究还为协同生成奠定了新的基础,为将预训练生成模型扩展到新的协同生成设置开辟了一条原则性的路径。
DiFlow-TTS: Compact and Low-Latency Zero-Shot Text-to-Speech with Discrete Flow Matching
中文标题:DiFlow-TTS: 基于离散流匹配的紧凑低延迟零样本文本到语音合成
作者:Ngoc-Son Nguyen, Thanh V. T. Tran, Hieu-Nghia Huynh-Nguyen, Truong-Son Hy, Van Nguyen
Zero-shot text-to-speech (TTS) has made significant progress in replicating unseen voices, yet balancing generation quality and inference efficiency remains challenging. Autoregressive models suffer from high latency, while diffusion-based approaches are constrained by training-time configurations. Moreover, most flow-based methods operate in continuous space, which introduces optimization challenges because continuous token spaces are inherently more complex than discrete ones. To address these limitations, we propose DiFlow-TTS, a novel zero-shot TTS framework based on discrete flow matching. The model consists of a deterministic Phoneme-Content Mapper for linguistic modeling and a Factorized Discrete Flow Denoiser that simultaneously generates prosody and acoustic token streams. Experimental results demonstrate the effectiveness of our approach across multiple evaluation metrics.
零样本文本到语音合成(Zero-shot TTS)在复现未见声音方面取得了显著进展,然而在生成质量与推理效率之间取得平衡仍具挑战性。自回归模型存在高延迟问题,而基于扩散的方法受限于训练时的配置约束。此外,大多数基于流的方法在连续空间中进行操作,这带来了优化难题,因为连续token空间本质上比离散空间更为复杂。为解决这些局限性,我们提出了DiFlow-TTS,一个基于离散流匹配的新型零样本TTS框架。该模型由一个确定性的音素内容映射器(Phoneme-Content Mapper)用于语言建模,以及一个因子化离散流去噪器(Factorized Discrete Flow Denoiser)同时生成韵律和声学token流组成。实验结果验证了我们的方法在多个评估指标上的有效性。
概述:
今日该分类(Image Compression)下的论文数量较少,仅有2篇论文发布。从论文标题来看,这两篇论文严格来说并不属于传统的图像压缩领域。
第一篇论文《Adversarial Attacks Leverage Interference Between Features in Superposition》研究对抗性攻击利用叠加特征间的干扰现象,属于AI安全与可解释性领域,与图像压缩无直接关联。
第二篇论文《Learning a Maximum Entropy Model for Visual Textures using Diffusion》研究基于扩散模型的视觉纹理最大熵建模,涉及图像生成与建模技术。虽然扩散模型在图像压缩领域也有应用(如基于扩散的图像重建),但该论文主要聚焦于纹理建模,可能与图像压缩中间表示学习有一定相关性。
推荐关注:
- Learning a Maximum Entropy Model for Visual Textures using Diffusion - 该论文探索了扩散模型在视觉纹理建模中的应用,其最大熵建模方法可能为图像压缩中的特征表示学习提供新思路,具有一定的参考价值。
注:今日图像压缩领域更新较少,建议关注其他相关分类如Image Generation、Computer Vision等获取更多内容。
Adversarial Attacks Leverage Interference Between Features in Superposition
中文标题:对抗攻击利用叠加中特征间的干扰
作者:Edward Stevinson, Lucas Prieto, Melih Barsbey, Tolga Birdal
Why do adversarial examples exist, and why do they transfer between models? Existing explanations appeal to high-dimensional geometry, non-robust patterns in the input, and decision boundary structure, but none provides a representation-level mechanism that explains why specific perturbations succeed and why attacks transfer between models. In this paper, we show that adversarial vulnerability can stem from efficient information encoding in neural networks. Specifically, vulnerability can arise from superposition - the phenomenon where networks represent more concepts than they have dimensions, forcing non-orthogonal representation and thus interference. This interference causes perturbations targeting one representation to affect others, creating vulnerabilities determined by interference patterns. In synthetic settings with precisely controlled superposition, we establish that superposition suffices to create adversarial vulnerability. The resulting attacks are predictable: PGD-discovered perturbations align with theoretically optimal perturbations derived from the interference geometry. Models trained on similar data develop similar interference patterns, explaining attack transferability. We then show that successful attacks on image classifiers exhibit the structure predicted by our proposed mechanism. These findings reveal that adversarial vulnerability can be a byproduct of networks' representational compression, complementing existing explanations based on data properties or architectural factors.
为什么对抗样本存在?为什么对抗样本可以在不同模型之间迁移?现有的解释诉诸于高维几何、输入中的非鲁棒模式以及决策边界结构,但没有提供表示层面的机制来解释特定扰动为何能够成功,以及攻击为何能在模型之间迁移。在本文中,我们表明对抗脆弱性可能源于神经网络中的高效信息编码。具体而言,脆弱性可能源于叠加现象——即网络所表示的概念数量超过其维度数量的现象,这迫使网络采用非正交表示,从而产生干扰。这种干扰导致针对某一表示的扰动会影响其他表示,形成由干扰模式决定的脆弱性。在精确控制叠加的合成设置中,我们证明了叠加足以产生对抗脆弱性。攻击结果是可预测的:PGD发现的扰动与从干扰几何推导出的理论最优扰动相一致。在相似数据上训练的模型会发展出相似的干扰模式,这解释了攻击的可迁移性。随后,我们表明对图像分类器的成功攻击呈现出我们提出的机制所预测的结构。这些发现揭示了对抗脆弱性可能是网络表示压缩的副产品,补充了基于数据特性或架构因素的现有解释。
Learning a Maximum Entropy Model for Visual Textures using Diffusion
中文标题:基于扩散学习的视觉纹理最大熵模型
作者:Xinyuan Zhao, Eero P. Simoncelli
Visual textures -- spatially homogeneous image regions containing repeated elements (e.g. a field of grass, the bark of a tree) -- are ubiquitous in visual scenes and provide important cues for recognizing and analyzing materials and objects. A number of existing texture models extract essential statistics from a single texture image, and can then generate high-quality samples that are visually similar to the original by matching these statistics. However, their statistics are either hand-designed or based on a network pretrained for another purpose (e.g., object recognition). Here, we develop the first principled method for unsupervised learning of a set of statistics that are used to constrain a maximum entropy probability model. We leverage methods developed for generative diffusion models to derive training and sampling procedures, and compare these to the traditional method of sampling via matching the statistics. Despite the compactness of our trained model (512 statistics), it generates texture images whose quality is as good as or better than the current state-of-the-art model (~177k statistics). A more direct comparison of the two models, obtained by synthesizing images that are indistinguishable for one model but maximally different for the other, reveals their relative strengths and weaknesses. Finally, we show that unlike previous statistical texture models, a straight trajectory in the representation space of our model generates homogeneous texture samples that interpolate smoothly between the features of the two end points.
视觉纹理——包含重复元素的空间同质图像区域(如草地、树皮)——在视觉场景中无处不在,为识别和分析材质与物体提供了重要线索。许多现有的纹理模型从单幅纹理图像中提取关键统计量,随后通过匹配这些统计量来生成与原图视觉相似的高质量样本。然而,这些统计量要么是人工设计的,要么是基于为其他目的(如目标识别)预训练的网络。本文提出了首个基于原则性的无监督学习方法,用于学习一组用于约束最大熵概率模型的统计量。我们利用生成式扩散模型中开发的方法推导出训练和采样程序,并将这些方法与传统的通过匹配统计量进行采样的方法进行比较。尽管我们训练得到的模型更加紧凑(512个统计量),但它生成的纹理图像质量与当前最先进模型(约177k个统计量)相当甚至更好。通过合成对一种模型无法区分但对另一种模型差异最大的图像,我们对两种模型进行了更直接的比较,揭示了它们的相对优势和劣势。最后,我们的研究表明,与以往统计纹理模型不同,我们模型表示空间中的直线路径可以生成同质纹理样本,并在两个端点的特征之间实现平滑插值。
今日未找到该分类的匹配论文。
今日未找到该分类的匹配论文。
今日未找到该分类的匹配论文。
今日未找到该分类的匹配论文。