每日 arXiv 论文简报
今日arXiv论文呈现以下趋势:扩散模型仍占据主导地位(26篇),从图像生成延伸到视频生成、图结构建模、机器人控制等多领域;自回归与扩散的边界正在模糊,如ARDY、OPSD-V等工作尝试将两者混合;视频生成成为最大热点,多篇论文从不同角度解决效率与质量问题;图神经网络与扩散的结合是新兴方向,DiPhon、Hypergraph Neural Stochastic Diffusion等将扩散应用于图结构数据;此外,国产AI模型(JuZhou 1.0)开始出现在国际舞台。整体上,高效化(少步蒸馏、采样优化)、多模态融合、机器人通用化是核心主题。
- 最值得关注:
- GeoProp:将机器人状态锚定视觉实现通用操作,是多模态具身智能的重要进展
- SAGA(两篇皆有):同时出现在Autoregressive和Diffusion中,提供视频生成的稳定加速指导,跨方向影响力强
- D2PO:通过动态偏好优化扩散采样器,直接提升生成质量与效率
- JuZhou 1.0:首个完全基于国产AI加速器训练的图像基础模型,填补国产化空白
- ARDY:自回归扩散与混合表示的结合,代表生成模型架构演进新思路
Autoregressive 类别今日概览:
今日Autoregressive相关论文聚焦于视频生成与视觉理解两大方向。视频生成领域出现多个创新工作,包括稳定加速指导、自蒸馏策略优化,以及将自回归与扩散模型结合的混合架构;视觉理解方面则关注深度模型的纹理表示与人类感知的对比。整体趋势显示,自回归模型正逐步融合扩散技术以提升生成质量,同时在无线通信、具身智能(RL策略、运动生成)等应用场景持续拓展。
- SAGA: Stable Acceleration Guidance for Autoregressive Video Generation — 提出稳定加速指导机制,有效解决自回归视频生成中的误差累积问题,提升长视频生成质量。
- OPSD-V: On-Policy Self-Distillation for Post-Training Few-Step Autoregressive Video Generators — 引入策略自蒸馏方法,实现少步自回归视频生成器的后训练优化,平衡生成效率与质量。
- LongE2V: Long-Horizon Event-based Video Reconstruction, Prediction, and Frame Interpolation — 首次将视频扩散模型应用于长程事件相机视频任务,兼顾重建、预测与插值三大功能。
- ARDY: Autoregressive Diffusion with Hybrid Representation for Interactive Human Motion Generation — 创新结合自回归与扩散机制,采用混合表示提升交互式人体运动生成的自然度与可控性。
- Texture Representations in Deep Vision Models — 系统比较CNN、ViT与人类视觉的纹理表征差异,为理解深度模型感知机制提供重要参考。
ORCAID: Oblique Rule-Based Continuous-Action Interpretation for Deep RL Policies
中文标题:ORCAID:面向深度强化学习策略的倾斜规则式连续动作解释方法
作者:Ignacio D. Lopez-Miguel, Ezio Bartocci, Thomas Eiter, Martin Tappler
Explainability remains a key issue in reinforcement learning (RL). Distilling an interpretable policy from an agent trained in a complex environment is particularly challenging when the action space is continuous. We introduce ORCAID, a novel method for extracting interpretable rule-based policies from RL agents operating in mixed continuous-discrete environments with continuous action spaces. Our main contribution is an efficient oblique decision tree training algorithm that partitions the state space by hyperplanes and fits local linear models. The key idea lies in a three-stage split search: efficient random initialization, local refinement, and backward elimination. Finally, adjacent leaves are merged to yield a concise set of interpretable rules describing a given deep RL policy. We evaluate ORCAID across multiple RL environments, demonstrating that the extracted rule-based policies maintain strong performance with a low number of parameters and can even be used to improve the performance of the original deep RL policy.
可解释性仍是强化学习(RL)中的关键问题。当动作空间为连续空间时,从复杂环境中训练的智能体中提取可解释策略尤为困难。我们提出了ORCAID,这是一种从混合连续-离散环境(具有连续动作空间)中运行的强化学习智能体提取可解释的基于规则的策略的新方法。我们的主要贡献在于提出了一种高效的倾斜决策树训练算法,该算法通过超平面划分状态空间并拟合局部线性模型。其核心思想在于三阶段分割搜索:高效随机初始化、局部精炼和后向消除。最后,合并相邻叶子节点以产生一组简洁的可解释规则,用于描述给定的深度强化学习策略。我们在多个强化学习环境中对ORCAID进行了评估,结果表明提取的基于规则的策略能够以较少的参数维持较强的性能,甚至可用于提升原始深度强化学习策略的性能。
Contrastive Predictive Coding with Compression for Enhanced Channel State Feedback in Wireless Networks
中文标题:用于增强无线网络信道状态反馈的压缩对比预测编码
作者:Ahmed Y. Radwan, Fahad Syed Muhammad, Matthew Baker, Hina Tabassum
Accurate and timely channel state information (CSI) is essential for next-generation wireless systems, yet existing works treat CSI compression and CSI prediction as separate problems, both in academia and in current 3GPP studies. Consequently, channel aging remains insufficiently addressed within standardized CSI feedback pipelines. In this article, we propose a unified compression-prediction framework that integrates Contrastive Predictive Coding (CPC) directly into the 3GPP-compliant CSI compression architecture. Instead of predicting high-dimensional CSI matrices, our approach forecasts future latent representations and jointly optimizes reconstruction fidelity and temporal predictive coherence via a combined 1-SGCS and InfoNCE objective. This design enables temporal representation learning without increasing feedback overhead. We present two variants: CPC-before-Compression, which performs autoregressive modeling on encoded features prior to quantization, and CPC-after-Compression, which shifts temporal modeling to the base-station to reduce the complexity of users' devices. Evaluations on 3GPP-compliant datasets from Nokia, Oppo, and CATT show that CPC-before-Compression achieves over 90% reconstruction accuracy with 32x lower decoder GFLOPs than the 3GPP baseline, while CPC-after-Compression preserves an identical encoder footprint and the same 64-bit feedback overhead. By unifying compression and prediction within a standardized pipeline, the proposed framework provides an age-aware, computationally efficient CSI feedback solution. The source code is publicly available at: https://github.com/AhmedRadwan02/cpc-3gpp
准确且及时的信道状态信息(CSI)对下一代无线系统至关重要,然而现有研究无论是学术界还是当前的3GPP研究都将CSI压缩和CSI预测视为两个独立的问题。因此,信道老化问题在标准化的CSI反馈流程中仍未得到充分解决。本文提出了一种统一的压缩-预测框架,将对比预测编码(CPC)直接集成到符合3GPP标准的CSI压缩架构中。我们的方法不是直接预测高维CSI矩阵,而是对未来潜在表示进行预测,并通过联合优化1-SGCS和InfoNCE目标函数来实现重建保真度与时间预测一致性的双重优化。该设计能够在不增加反馈开销的情况下进行时间表征学习。我们提出了两种变体:压缩前CPC(在量化之前对编码特征进行自回归建模)和压缩后CPC(将时间建模移至基站端以降低用户设备的复杂度)。在来自诺基亚、OPPO和中国信科的3GPP兼容数据集上的评估表明,压缩前CPC在保持与3GPP基线相同的64位反馈开销下,实现了超过90%的重建精度,且解码器GFLOPs降低了32倍;而压缩后CPC则保持了相同的编码器结构。通过在标准化流程中统一压缩与预测,所提框架提供了一种时效感知、计算高效的CSI反馈解决方案。源代码已公开于:https://github.com/AhmedRadwan02/cpc-3gpp
SAGA: Stable Acceleration Guidance for Autoregressive Video Generation
中文标题:SAGA:自回归视频生成的稳定加速引导方法
作者:Thanh-Nhan Vo, Trong-Thuan Nguyen, Trung-Hoang Le, Tam V. Nguyen, Minh-Triet Tran
Autoregressive video diffusion enables efficient streaming and long-horizon video generation, but repeatedly reusing generated latents as causal context can amplify temporal errors, resulting in flickering, motion jitter, and structural drift. In this paper, we investigate this failure mode from a spectral kinematic perspective and identify discrete latent acceleration as an effective signal for revealing unstable high-frequency temporal perturbations. To this end, we propose SAGA, a training-free \textbf{\textit{s}}table \textbf{\textit{a}}cceleration \textbf{\textit{g}}uidance approach for \textbf{\textit{a}}utoregressive video generation. SAGA integrates an acceleration domain spectral guidance objective based on finite-window Slepian projections with a structured autoregressive noise initialization strategy that suppresses short-range temporal correlations while preserving long-range motion structure. Without retraining or modifying the backbone, SAGA can be directly applied to existing chunk-wise autoregressive diffusion models, which is the prevalent setting for high-quality generation. Extensive experiments show that SAGA consistently improves temporal quality across multiple autoregressive diffusion models. On Self-Forcing, SAGA improves Temporal Quality from 97.30 to 97.91 and Image Quality from 69.60 to 70.51. Moreover, spectral analysis and human preference studies demonstrate that SAGA reduces temporal instability while maintaining visual fidelity.
自回归视频扩散能够实现高效的流式生成和长时序视频生成,但重复使用已生成的潜在表示作为因果上下文可能会放大时间误差,导致闪烁、运动抖动和结构漂移等问题。本文从光谱运动学视角研究这一失效模式,并识别出离散潜在加速度作为揭示不稳定高频时间扰动的有效信号。基于此,本文提出了SAGA,一种无需训练的自回归视频生成稳定加速引导方法。SAGA将基于有限窗口Slepian投影的加速度域谱引导目标与结构化自回归噪声初始化策略相结合,在保留长程运动结构的同时抑制短程时间相关性。SAGA无需重新训练或修改主干网络,可直接应用于现有的分块自回归扩散模型,这也是高质量视频生成的主流设置。大量实验表明,SAGA在多个自回归扩散模型上持续提升时间质量。在Self-Forcing数据集上,SAGA将时间质量从97.30提升至97.91,将图像质量从69.60提升至70.51。此外,光谱分析和人类偏好研究表明,SAGA在保持视觉保真度的同时减少了时间不稳定性。
Texture Representations in Deep Vision Models: Comparing CNNs, Vision Transformers, and Human Perception
中文标题:深度视觉模型中的纹理表征:比较CNN、视觉Transformer与人类感知
作者:Ludovica de Paolis, Marco Baroni, Alessandro Laio, Eugenio Piasini
In computational vision science, Convolutional Neural Networks (CNNs) have emerged as a popular model of biological vision because of the alignment they can exhibit with neural and behavioral data in humans and animals. However, it remains unclear to what extent this alignment persists for visual tasks that extend beyond the canonical object recognition paradigm based on well defined semantic content. In this study, we diverge from the common object-centric view by focusing on another aspect of vision: texture perception. We consider textures of different complexity generated with three different algorithms from the same source images. Using a rank-based statistic, we quantify the information encoded in the internal representations of a CNN and three Vision Transformers (ViTs), and we compare the similarity of these representations to those inferred from human psychophysics data. We find that the representation of textures is aligned in different ViTs, but not between the ViTs and the CNN; that ViTs form similar representations for textures of different complexity; that human performance in recognizing textures can be better predicted from ViTs representations rather than CNN representations. Taken together, these results suggest that ViTs may capture more faithfully than CNNs how texture patterns are visually processed by humans, and that the representations of texture stimuli in computational models may be driven by the network architecture.
在计算视觉科学中,卷积神经网络(CNN)已成为生物视觉的流行模型,因为它们能够与人类和动物的神经和行为数据保持一致。然而,这种一致性在超出基于明确语义内容的典型目标识别范式的视觉任务中能否持续仍不清楚。本研究从常见的以目标为中心的视角转向视觉的另一个方面:纹理感知。我们考虑了使用三种不同算法从相同源图像生成的不同复杂度的纹理。通过基于排名的统计量,我们量化了CNN和三个视觉Transformer(ViT)内部表征中编码的信息,并将这些表征与从人类心理物理学数据推断出的表征进行比较。我们发现,不同ViT之间的纹理表征是一致的,但ViT与CNN之间则不一致;ViT对不同复杂度的纹理形成相似的表征;人类识别纹理的表现可以更好地从ViT表征而非CNN表征来预测。综合来看,这些结果表明,ViT可能比CNN更真实地捕捉人类视觉处理纹理模式的方式,纹理刺激在计算模型中的表征可能由网络架构驱动。
OPSD-V: On-Policy Self-Distillation for Post-Training Few-Step Autoregressive Video Generators
中文标题:OPSD-V: 面向后训练少步自回归视频生成器的策略自蒸馏方法
作者:Hongyu Liu, Chun Wang, Feng Gao, Xuanhua He, Yue Ma, Ziyu Wan, Yong Zhang, Xiaoming Wei, Qifeng Chen
We propose OPSD-V, an on-policy self-distillation paradigm for post-training few-step autoregressive (AR) video diffusion models. Existing few-step AR video generators can produce long videos with low latency, but still suffer from error accumulation and weakened motion dynamics during long autoregressive rollout. OPSD-V reduces long-horizon degradation while preserving the original few-step inference path. The key idea is to introduce real long-video data as temporal context during training and use it to provide dense trajectory-level supervision. Specifically, the student follows the exact inference-time rollout, generating each chunk conditioned on its own previously generated KV cache. In parallel, the teacher is evaluated at the same student-visited denoising states, but uses a cleaner AR-consistent temporal cache in which older history can be replaced by real-video context. This provides dense denoising-level corrective targets under on-policy AR cache dynamics, without changing the sampler, number of denoising steps, or inference-time cache mechanism. We apply OPSD-V to representative few-step AR video models, including Self-Forcing and LongLive. Experiments show consistent improvements in visual quality, motion dynamics, and VBenchLong scores. A user study with 10 participants comparing 20 video pairs shows that OPSD-V is preferred over the base models in 66.0% of overall-preference judgments (82.5% excluding ties).
本文提出OPSD-V,一种面向后训练少步自回归视频扩散模型的策略自蒸馏范式。现有的少步自回归视频生成器虽然能够以低延迟生成较长的视频,但在长时自回归展开过程中仍面临误差累积和运动动态减弱的问题。OPSD-V在保留原有少步推理路径的同时,降低了长程退化问题。其核心思想是在训练过程中引入真实长视频数据作为时序上下文,并利用其提供密集的轨迹级监督。具体而言,学生模型遵循推理时的展开方式,基于自身生成的KV缓存来生成每个片段;与此同时,教师模型在学生访问的相同去噪状态下进行评估,但使用更清晰的自回归一致时序缓存,其中较旧的历史可以被真实视频上下文所替代。这在保持采样器、去噪步数及推理时缓存机制不变的情况下,基于策略自回归缓存动态提供了密集的去噪级纠正目标。本文将OPSD-V应用于代表性的少步自回归视频模型,包括Self-Forcing和LongLive。实验表明,OPSD-V在视觉质量、运动动态和VBenchLong分数方面均取得一致的提升。针对20对视频进行的10人用户研究表明,在总体偏好判断中,OPSD-V在66.0%的案例中更受青睐(排除平局后为82.5%)。
LongE2V: Long-Horizon Event-based Video Reconstruction, Prediction, and Frame Interpolation with Video Diffusion Models
中文标题:LongE2V:基于视频扩散模型的长时序事件相机视频重建、预测与帧插值
作者:Cheng-De Fan, Chun-Wei Tuan Mu, Chen-Wei Chang, Chin-Yang Lin, Kun-Ru Wu, Yu-Chee Tseng, Yu-Lun Liu
Recovering high-quality video from sparse event streams is a challenging task. Regression methods often blur textures, while existing generative models struggle with long-term stability. We propose LongE2V, a novel approach that leverages pre-trained video diffusion priors to jointly handle event-based video reconstruction, prediction, and frame interpolation. By fine-tuning a foundational video model, our approach achieves high data efficiency and superior perceptual quality. We introduce Autoregressive Unrolling and Adaptive Context Switching to mitigate temporal drift in extremely long sequences. We also propose Reencoding Alignment with Cross Residual Correction to ensure precise bidirectional consistency during frame interpolation. Furthermore, Event Voxel Density Augmentation ensures robustness across varying sensor resolutions. Extensive experiments on real-world benchmarks demonstrate that LongE2V outperforms state-of-the-art methods across all three tasks, exhibiting exceptional temporal coherence and zero-shot generalization. Project page: https://cdfan0627.github.io/LongE2V-page/
从稀疏事件流中恢复高质量视频是一项具有挑战性的任务。回归方法往往会导致纹理模糊,而现有生成模型在长期稳定性方面存在不足。我们提出LongE2V,这是一种利用预训练视频扩散先验的新方法,可联合处理基于事件的视频重建、预测和帧插值。通过微调基础视频模型,我们的方法实现了较高的数据效率和卓越的感知质量。我们引入自回归展开和自适应上下文切换来缓解极长序列中的时间漂移问题。此外,我们还提出重新编码对齐与交叉残差校正,以确保帧插值过程中的精确双向一致性。事件体素密度增强则确保了模型在不同传感器分辨率下的鲁棒性。在真实世界基准数据集上的广泛实验表明,LongE2V在所有三项任务上均优于最先进的方法,表现出优异的时间一致性和零样本泛化能力。项目主页:https://cdfan0627.github.io/LongE2V-page/
ARDY: Autoregressive Diffusion with Hybrid Representation for Interactive Human Motion Generation
中文标题:ARDY:用于交互式人体运动生成的自回归扩散混合表示框架
作者:Kaifeng Zhao, Mathis Petrovich, Haotian Zhang, Tingwu Wang, Siyu Tang, Davis Rempe
Generating realistic 3D human motions in real-time within interactive applications is key for animation, simulation, and humanoid robotics. While recent offline motion generation approaches offer precise control via text and kinematic constraints, they lack the inference speed required for interactive settings. Conversely, existing online methods enable real-time synthesis but often sacrifice controllability or struggle with complex text semantics and long-horizon goals due to limited context windows. In this work, we introduce ARDY, a streaming generation framework that bridges this gap by enabling high-fidelity motion generation controllable via online text prompts and flexible kinematic constraints. ARDY employs a hybrid representation that combines explicit root features with a latent body embedding, balancing precise trajectory control with efficient generative learning. We propose a two-stage autoregressive transformer denoiser that features variable history context and supports conditioning on flexible, long-horizon kinematic constraints. By training on a large-scale motion capture dataset and being directly conditioned on text labels and kinematic constraints sampled from ground truth poses, ARDY natively learns controllable generation that supports online prompting and flexible long-horizon goals. Extensive evaluations on the HumanML3D benchmark and the large-scale, high-fidelity Bones Rigplay dataset demonstrate ARDY's high motion quality and constraint adherence, validating the efficacy of our key architectural decisions. Finally, we demonstrate the method&x27;s practical versatility through an interactive demo featuring dynamic text control, diverse keyframe pose constraints, path following, and interactive locomotion control via mouse and keyboard. Supplementary video results, code, and model releases can be found at https://research.nvidia.com/labs/sil/projects/ardy/.
在交互式应用中实时生成逼真的3D人体运动对于动画、仿真和类人机器人学至关重要。尽管近期离线运动生成方法可通过文本和运动学约束实现精确控制,但缺乏交互场景所需的推理速度。相反,现有在线方法虽能实现实时合成,但往往牺牲了可控性,或因有限的时间上下文窗口而难以处理复杂文本语义和长时序目标。本工作提出ARDY,一个流式生成框架,通过支持高保真运动生成并可通过在线文本提示和灵活的运动学约束进行控制来弥合这一差距。ARDY采用结合显式根节点特征与潜在身体嵌入的混合表示,在精确轨迹控制与高效生成学习之间取得平衡。我们提出了一种具有可变历史上下文并支持灵活长时序运动学约束条件的两阶段自回归Transformer去噪器。通过在大规模动作捕捉数据集上训练,并直接以文本标签和从真实姿态采样的运动学约束为条件,ARDY原生学习可控生成,支持在线提示和灵活的长时序目标。在HumanML3D基准和大规模高保真Bones Rigplay数据集上的广泛评估表明,ARDY具有高运动质量和约束依从性,验证了我们关键架构决策的有效性。最后,我们通过一个交互式演示展示了该方法的实际通用性,包括动态文本控制、多样化关键帧姿态约束、路径跟随以及通过鼠标和键盘实现的交互式运动控制。补充视频结果、代码和模型发布于https://research.nvidia.com/labs/sil/projects/ardy/。
今日Diffusion领域论文整体趋势呈现多模态应用深化与采样训练双线并进的态势。视频生成依然是核心热点,涉及长时序重建、自回归加速、3D可控转换等多个子方向;同时,图结构与扩散模型的结合成为新晋亮点,在图匹配、图生成、神经表征等领域均有探索。采样效率优化和可控生成依然是关键技术突破点,涵盖动态偏好采样、时间步加权、对比引导等方法。此外,扩散模型向机器人控制、医疗影像、深度伪造检测等垂直场景延伸的趋势明显。
- D2PO: Optimizing Diffusion Samplers via Dynamic Preference - 提出动态偏好优化扩散采样器,为采样效率提升提供了新范式,值得关注。
- ContrastiveCFG: Guiding Diffusion Sampling by Contrasting Positive and Negative Concepts - 创新性地通过对比正负概念引导扩散采样,为可控生成提供了简洁有效的思路。
- SAGA: Stable Acceleration Guidance for Autoregressive Video Generation - 解决自回归视频生成中的加速与稳定性矛盾,是视频生成领域的重要进展。
- ardy: Autoregressive Diffusion with Hybrid Representation for Interactive Human Motion Generation - 混合表示的自回归扩散方法,在交互式人体动作生成上展现潜力。
- LongE2V: Long-Horizon Event-based Video Reconstruction with Video Diffusion Models - 聚焦长时序事件相机视频的重建与插帧,拓展了扩散模型在事件相机领域的应用。
D2PO: Optimizing Diffusion Samplers via Dynamic Preference
中文标题:D2PO:基于动态偏好的扩散采样器优化
作者:Jinkyu Kim, Jinyoung Choi, Bohyung Han
We propose D2PO (Dynamic Direct Preference Optimization), a principled framework for optimizing diffusion sampling policies with respect to timestep schedules and classifier-free guidance (CFG) weights. Our work is motivated by a fundamental limitation of existing student-teacher regression frameworks; low-NFE student samplers are trained to mimic high-NFEteachers, often sacrificing high-frequency texture fidelity while preserving coarse global structures, thereby misaligning the sampler with perceptual quality. D2PO addresses this challenge by reformulating sampler optimization as a preference-based alignment problem, leveraging the Direct Preference Optimization (DPO) framework. To make DPO applicable to diffusion samplers, we model the sampling policy as an energy-based model (EBM), transforming preference comparisons into tractable energy differences. We further introduce a novel energy formulation derived directly from the pretrained score network, enabling preference evaluation in perturbed spaces that jointly capture structural consistency and fine-grained details. Moreover, we introduce dynamic preferences, where the preferred samples used for alignment progressively improve as the sampling policies are learned. This self-improving mechanism replaces rigid static teacher supervision with an iterative, preference-guided refinement process, providing progressively stronger alignment signals. Extensive experiments demonstrate that D2PO aligns diffusion samplers with perceptual quality more faithfully, unlocking the full potential of high-quality teachers and consistently outperforming conventional regression-based schedulers under low-NFE constraints.
我们提出了D2PO(动态直接偏好优化),这是一个优化扩散采样器时间步调度器和无分类器引导(CFG)权重的原则性框架。我们的工作源于现有学生-教师回归框架的一个根本局限性:低NFE学生采样器被训练以模仿高NFE教师采样器,通常在保留粗粒度全局结构的同时牺牲了高频纹理保真度,从而导致采样器与感知质量出现偏差。D2PO通过将采样器优化重新表述为基于偏好的对齐问题来解决这一挑战,并利用直接偏好优化(DPO)框架。为了使DPO适用于扩散采样器,我们将采样策略建模为能量模型(EBM),将偏好比较转化为可处理的能量差。我们进一步引入了一种直接从预训练得分网络推导出的新型能量公式,能够在扰动空间中进行偏好评估,同时捕获结构一致性和细粒度细节。此外,我们引入了动态偏好,即用于对齐的偏好样本随着采样策略的学习而逐步改进。这种自我改进机制用迭代的、偏好引导的细化过程取代了僵化的静态教师监督,提供了逐步增强的对齐信号。大量实验表明,D2PO能够更忠实地将扩散采样器与感知质量对齐,释放高质量教师的全部潜力,并在低NFE约束下始终优于传统的基于回归的调度器。
Dynamic-in-Few-Step: Unifying Dynamic Computation and Few-Step Distillation for Efficient Video Generation
中文标题:Dynamic-in-Few-Step:统一动态计算与少步蒸馏的高效视频生成方法
作者:Yu Cheng, Siyue Yao, Zhongang Qi, Shanyan Guan, Wei Li, Fajie Yuan
Video Diffusion Models (VDMs) have demonstrated superior generation quality but suffer from prohibitive computational costs. While recent few-step distillation techniques significantly accelerate inference, they typically enforce a static model architecture across all denoising stages, ignoring the varying computational demands inherent to different noise levels. In this work, we propose a novel post-training acceleration framework that exploits this redundancy by integrating dynamic structural sparsification directly into the distillation process. Unlike conventional post-hoc compression applied to a fixed diffusion pipeline, our approach jointly optimizes the denoising steps and structured model sparsity, transforming a pre-trained VDM into a compact, step-specific Mixture-of-Models (MoM). To address the training instability arising from this joint optimization, we introduce a Progressive Training Strategy coupled with an Output Rollout Mechanism, which ensures the coherent learning of structural decisions across timesteps. Furthermore, we develop a specialized inference engine to deploy the resulting MoM efficiently. Our method is orthogonal to existing acceleration techniques and highly effective: On Wan-14B, it removes 24% of the per-step FLOPs on top of 4-step distillation, adding a 1.2x wall-clock gain and reaching a 30x speedup over the 50-step teacher while preserving competitive generation quality.
视频扩散模型(Video Diffusion Models, VDMs)虽然展示了优异的生成质量,但存在计算成本过高的问题。尽管近期提出的少步蒸馏技术显著加速了推理过程,但这些方法通常在所有去噪阶段强制使用静态模型架构,忽略了不同噪声水平所固有的计算需求差异。本工作提出了一种新型的后训练加速框架,通过将动态结构稀疏化直接集成到蒸馏过程中来利用这种冗余。与传统的针对固定扩散管道进行的后处理压缩不同,本方法联合优化去噪步数和结构化模型稀疏性,将预训练的VDM转化为紧凑的步态特定混合模型(Mixture-of-Models, MoM)。为解决该联合优化带来的训练不稳定问题,本研究引入了渐进训练策略(Progressive Training Strategy)结合输出Rollout机制(Output Rollout Mechanism),以确保结构决策在时间步长上的一致性学习。此外,本研究还开发了专门的推理引擎以高效部署所得到的混合模型。本方法与现有加速技术正交且效果显著:在Wan-14B模型上,本方法在4步蒸馏的基础上进一步消除了每步24%的FLOPs计算量,实现了1.2倍的实际时间收益,并在保持具有竞争力的生成质量前提下,达到了相比50步教师模型30倍的加速效果。
Diffusion enabled Optimal Transport distances for graph matching
中文标题:用于图匹配的扩散增强最优传输距离
作者:Iman Seyedi, Francesco Archetti
This paper introduces Diffusion Semi-Relaxed Fused Gromov-Wasserstein (DsrFGW), a novel method for graph comparison that unifies node features and structural connectivity through optimal transport. While traditional Gromov-Wasserstein and semi-relaxed variants (srGW, srFGW) capture graph structure, they often struggle with sparse, noisy, or partially observed graphs. Inspired by Graph Diffusion Distance, which posits graphs are similar if they enable similar information transmission patterns, DsrFGW incorporates diffusion processes allowing information propagation across nodes, capturing local and global structural patterns while reducing sensitivity to noise or missing edges. An extensive evaluation on 36 synthetic pairwise graph matching tasks (easy, medium, hard) demonstrates consistent superiority over srFGW, achieving accuracy improvements of 0-20 percentage points and dramatic Adjusted Rand Index (ARI) gains: in medium-difficulty scenarios, srFGW often achieves negative ARI (worse than random) while DsrFGW offers better performance in terms of both internal and external clustering quality measures (i.e., Adjusted Rank Index and Accuracy with respect to the true underlying clusters, respectively). Even under severe noise, DsrFGW improves clustering quality in 92% of the synthetic tasks with optimal diffusion scales adapting to problem difficulty, establishing DsrFGW as a robust framework for graph comparison under structural uncertainty.
本文提出了扩散半松弛融合Gromov-Wasserstein(DsrFGW)方法,这是一种通过最优传输统一节点特征和结构连通性的图比较新方法。传统的Gromov-Wasserstein及其半松弛变体(srGW、srFGW)虽然能够捕捉图结构,但在处理稀疏、有噪声或部分观测的图时往往表现不佳。受图扩散距离的启发——该观点认为如果图具有相似的信息传输模式则图是相似的——DsrFGW引入扩散过程,允许信息在节点间传播,从而捕捉局部和全局结构模式,同时降低对噪声或缺失边的敏感性。在36个合成成对图匹配任务(简单、中等、困难)上的广泛评估表明,DsrFGW始终优于srFGW,准确率提升0-20个百分点,并在调整兰德指数(ARI)方面取得显著增益:在中等难度场景中,srFGW经常获得负ARI(比随机更差),而DsrFGW在内部和外部聚类质量度量方面表现更优(即分别相对于真实底层聚类的调整秩指数和准确率)。即使在严重噪声条件下,DsrFGW仍能在92%的合成任务中改善聚类质量,并具有适应问题难度的最优扩散尺度,确立了DsrFGW作为结构不确定性下图比较的稳健框架的地位。
Latent graph encoding of multimodal neuroimaging features with generative AI architectures
中文标题:基于生成式AI架构的多模态神经影像特征潜在图编码方法
作者:Ishaan Batta, Meenu Ajith, Vince Calhoun
While generative models enable encoding of complex neuroimaging data for feature generation and reconstruction, developing optimal architectural frameworks with appropriate encoding and latent space processes is crucial for studying structural and functional properties of the brain. We design a multimodal generative framework for structural and functional magnetic resonance imaging (MRI) features through systematic evaluation of encoding strategies, latent multimodal fusion, and generative model selection. Using structural gray matter volume (GMV) and static functional network connectivity (sFNC) features from a large neuroimaging dataset, we analyze generative frameworks involving variational autoencoders (VAEs), transformers, generative adversarial networks (GANs), and diffusion models. Architectures that employ modality-aware graph encoding of functional connectivity into a lower-dimensional latent space outperform vectorized encoders or direct data space approaches. The proposed multimodal graph VAE (gMMVAE) surpasses alternative generative variants across multiple metrics for generation fidelity, reconstruction quality, efficiency, and latent space discriminability, highlighting its potential for robust multimodal neuroimaging analysis.
生成模型能够对复杂神经影像数据进行编码以实现特征生成和重构,而开发具有适当编码和潜在空间处理机制的优化架构框架对于研究大脑的结构与功能特性至关重要。我们设计了一个用于结构和功能磁共振成像(MRI)特征的多模态生成框架,对编码策略、潜在多模态融合和生成模型选择进行了系统评估。利用大型神经影像数据集的结构灰质体积(GMV)和静态功能网络连接(sFNC)特征,我们分析了涉及变分自编码器(VAE)、Transformer、生成对抗网络(GAN)和扩散模型的生成框架。将功能连接以模态感知图编码方式映射到低维潜在空间的架构表现优于向量化编码器或直接数据空间方法。所提出的多模态图变分自编码器(gMMVAE)在生成保真度、重构质量、效率和潜在空间可区分性等多个指标上均优于其他生成变体,突显出其在稳健多模态神经影像分析中的潜力。
GeoProp: Grounding Robot State in Vision for Generalist Manipulation
中文标题:GeoProp:将机器人状态扎根于视觉的通用操作方法
作者:Guoyang Zhao, Quanhao Qian, Gongjie Zhang, Wenhao Li, Jiuniu Wang, Xiaowei Lu, Deli Zhao, Ran Xu
Proprioception is fundamental to robotic manipulation, yet standard fusion methods often treat it as an isolated vector lacking explicit alignment with visual tokens. Without a direct correspondence between 3D kinematics and 2D feature maps, manipulation policies struggle to ground the robot's state within the scene, frequently underperforming even vision-only baselines. To address this, we introduce GeoProp, a lightweight, plug-and-play adapter that aligns proprioception with vision through explicit geometric grounding and spatial feature sampling. GeoProp projects the robot state onto the image plane to sample localized visual features, constructing a grounded state token. It then injects state-derived spatial priors into the corresponding visual features via FiLM modulation. To capture motion intent, GeoProp further samples features at a short-horizon predicted coordinate derived from recent kinematics, providing look-ahead visual context. Across 67 tasks, GeoProp improves Diffusion Policy by 8.7% on 63 simulation tasks and pi_0 by 4.0% on the RoboTwin subset, and yields a 10.6% average gain across both policy families in the real world, while adding only 2-3% to the parameter count. These results demonstrate that GeoProp is a simple yet high-impact inductive bias for generalist embodied policies. Project page: https://alibaba-damo-academy.github.io/GeoProp/.
本体感觉是机器人操作的基础,然而标准融合方法往往将其视为缺乏与视觉token显式对齐的孤立向量。由于缺少3D运动学与2D特征图之间的直接对应关系,操作策略难以将机器人状态扎根于场景中,经常甚至不如纯视觉基线方法。为此,我们提出了GeoProp,这是一种轻量级即插即用适配器,通过显式几何扎根和空间特征采样将本体感觉与视觉对齐。GeoProp将机器人状态投影到图像平面上以采样局部视觉特征,从而构建扎根状态token。随后,它通过FiLM调制将状态衍生的空间先验注入相应的视觉特征中。为了捕捉运动意图,GeoProp还会对基于最近运动学推导出的短视界预测坐标处的特征进行采样,从而提供前瞻视觉上下文。在67项任务中,GeoProp在仿真任务上将Diffusion Policy提升了8.7%,在RoboTwin子集上将pi_0提升了4.0%,并在真实世界中为两种策略家族带来了10.6%的平均增益,同时仅增加了2-3%的参数量。这些结果表明GeoProp是一种简洁而高影响力的归纳偏置,适用于通用具身策略。项目页面:https://alibaba-damo-academy.github.io/GeoProp/。
DiPhon: Diffusion on Graphons for Scalable Graph Generation
中文标题:DiPhon:基于图on的可扩展图生成扩散模型
作者:Sergio Rozada, Yiming Qin, Manuel Madeira, Pascal Frossard, Alejandro Ribeiro
Diffusion models represent a leading paradigm for graph generation, with notable impact in domains such as molecular design. Yet, scaling these models to large graphs remains an open problem. We approach this question in the dense-graph setting through the lens of graphons, the size-agnostic limit objects of dense graph sequences, to study how structural graph statistics behave across node-size scales. This perspective leads to DiPhon, a diffusion framework for size-scalable graph generation. Specifically, we formulate a continuous diffusion process on the graphon space via a Jacobi stochastic differential equation (SDE), and propose DiPhon, a discretized graph-level process that mimics these dynamics on finite graphs. We further derive the corresponding reverse-time process, which requires access to the marginal score. For the Jacobi process, this score interestingly admits a tractable form, which we estimate from data via graph denoising and plug into the reverse process to generate graph samples. We prove that DiPhon matches exactly the first moment of the marginal distributions induced by the continuous graphon process, and approximates the second moment up to a closed-form discrepancy. Thus, DiPhon inherits key size-agnostic statistical properties of the graphon dynamics, providing a principled route toward scalable graph generation. Empirically, we demonstrate this scalability by training on small graphs and generating progressively larger graphs at inference time, without retraining, while preserving their core topological properties.
扩散模型是图生成领域的领先范式,在分子设计等领域产生了显著影响。然而,将这些模型扩展至大规模图仍然是一个开放性问题。本文从图on的视角来研究密集图设置,图on是密集图序列的无尺寸极限对象,用于探究结构图统计量在不同节点规模下的行为规律。基于这一视角,本文提出了DiPhon,一个用于尺寸可扩展图生成的扩散框架。具体而言,本文通过Jacobi随机微分方程(SDE)在图on空间上构建连续扩散过程,并提出DiPhon,一种在有限图上模拟该动力学的离散图级过程。本文进一步推导了对应的时间反向过程,该过程需要获取边缘分数。有趣的是,对于Jacobi过程,该分数具有可处理的形式,本文通过图去噪从数据中估计该分数,并将其代入反向过程以生成图样本。本文证明了DiPhon精确匹配由连续图on过程诱导的边缘分布的一阶矩,并对二阶矩进行近似,误差为闭合形式。因此,DiPhon继承了图on动力学的关键无尺寸统计特性,为可扩展图生成提供了原则性的实现路径。实验上,本文通过在小图上训练并在推理时生成逐步增大的图来验证这种可扩展性,无需重新训练,同时保持图的核心拓扑特性。
Hypergraph Neural Stochastic Diffusion: An SDE Framework for Uncertainty Estimation
中文标题:超图神经随机扩散:一种用于不确定性估计的SDE框架
作者:Zhiheng Zhou, Mengyao Zhou, Dengyi Zhao, Xingqin Qi, Guiying Yan
Hypergraph neural networks have shown powerful capability in modeling higher-order relations, yet their predictive uncertainty remains underexplored. Unlike pairwise graphs, uncertainty in hypergraphs arises not only from noisy attributes and ambiguous labels, but also from variations in node-hyperedge incidence structures and complex higher-order dependencies. Existing approaches mainly estimate uncertainty from final predictions or rely on computationally expensive ensembles and Bayesian inference, limiting their ability to capture uncertainty evolution during representation learning. In this paper, we propose Hypergraph Neural Stochastic Diffusion(HyperNSD), a stochastic differential equation framework for uncertainty estimation on hypergraphs. HyperNSD models hypergraph representations as stochastic processes evolving over node-hyperedge incidence structures. A learnable drift function captures deterministic higher-order diffusion dynamics, while a learnable stochastic forcing function characterizes structural ambiguity and representation noise. Predictive uncertainty is directly quantified through the variability of stochastic representation trajectories, providing an intrinsic uncertainty measure beyond post-hoc confidence scores. We formulate HyperNSD with neural drift and diffusion networks, enabling joint learning of prediction and uncertainty propagation. Theoretical analyses establish well posedness, perturbation stability,permutation equivariance, and numerical convergence of the proposed stochastic dynamics. Experiments on multiple hypergraph benchmarks demonstrate that HyperNSD achieves reliable uncertainty estimation for out-of-distribution and misclassification detection while preserving competitive prediction accuracy. These results provide a principled stochastic-dynamical framework for trustworthy higher-order representation learning.
超图神经网络在建模高阶关系方面展现出强大的能力,但其预测不确定性仍未得到充分探索。与成对图不同,超图中的不确定性不仅来源于噪声属性和模糊标签,还源于节点-超边关联结构的变化以及复杂的高阶依赖关系。现有的方法主要从最终预测结果估计不确定性,或依赖于计算开销较大的集成方法和贝叶斯推理,这限制了它们在表示学习过程中捕捉不确定性演化的能力。本文提出了超图神经随机扩散(Hypergraph Neural Stochastic Diffusion,HyperNSD),一种用于超图不确定性估计的随机微分方程框架。HyperNSD将超图表示建模为在节点-超边关联结构上演化的随机过程。可学习的漂移函数捕捉确定性高阶扩散动力学,而可学习的随机驱动力函数则表征结构歧义和表示噪声。预测不确定性通过随机表示轨迹的变异性直接量化,提供了一种超越事后置信度分数的内在不确定性度量。我们使用神经漂移和扩散网络构建HyperNSD,实现了预测与不确定性传播的联合学习。理论分析建立了所提随机动力学框架的适定性、扰动稳定性、置换等变性和数值收敛性。在多个超图基准数据集上的实验表明,HyperNSD在保持竞争力的预测精度的同时,能够对分布外检测和误分类检测进行可靠的不确定性估计。这些结果为可信赖的高阶表示学习提供了一个原则性的随机动力学框架。
Quantum simulation of real-world nonlinear dynamics via Koopman method
中文标题:基于Koopman方法的真实世界非线性动力学量子模拟
作者:Baoyang Zhang, Dong An, Zhaoyuan Meng, Yefei Yu, Xiaoxiao Xiao, Zhen Lu, Yue Yang
Nonlinear dynamics is ubiquitous in nature, ranging from chemical pattern formation to ocean circulation, yet its simulation on quantum computers is fundamentally limited by the unitary nature of quantum evolution. We propose the quantum Koopman method, a data-driven framework that embeds nonlinear dynamics into a learned linear representation and implements the resulting evolution using shallow quantum circuits. This method learns Koopman observables from trajectory data, projects the lifted dynamics onto a finite-dimensional subspace, and decomposes the corresponding non-unitary propagator into parallel spectral channels. We utilize the Koopman method on a superconducting processor to simulate three distinct nonlinear systems, comprising reaction-diffusion dynamics, fluid motion on a sphere, and satellite-derived observations of Gulf Stream currents, employing up to 32 parallel circuits of 10 qubits. These quantum simulations capture the dominant multiscale patterns and statistical signatures of the underlying dynamics, and reveal a transition from performance limited by hardware noise in weakly nonlinear systems to performance limited by finite-dimensional Koopman representations as nonlinear scale interactions increase. This transition identifies a practical boundary for quantum-amenable nonlinear dynamics, establishing a hardware-validated route for simulating moderately nonlinear dynamics on near-term quantum hardware.
非线性动力学在自然界中无处不在,从化学图案形成到海洋环流,然而其在量子计算机上的模拟从根本上受限于量子演化的酉特性。我们提出量子Koopman方法,这是一种数据驱动框架,将非线性动力学嵌入到学习的线性表示中,并利用浅层量子电路实现所得到的演化。该方法从轨迹数据中学习Koopman可观测量,将提升的动力学投影到有限维子空间,并将其对应的非酉传播子分解为并行的谱通道。我们在超导处理器上利用Koopman方法模拟了三个不同的非线性系统,包括反应扩散动力学、球面流体运动以及来自卫星观测的墨西哥湾流数据,采用多达32个并行的10比特电路。这些量子模拟捕获了底层动力学的主导多尺度模式和统计特征,并揭示了从弱非线性系统中硬件噪声限制性能到随着非线性尺度相互作用增加而有限维Koopman表示限制性能的转变。这一转变确定了量子可行非线性动力学的实用边界,为在近term量子硬件上模拟中等非线性动力学建立了硬件验证路径。
Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF
中文标题:面向样本高效扩散模型RLHF的选择性时间步加权与基于优势的重放策略
作者:Eric Zhu, Abhinav Shrivastava, Soumik Mukhopadhyay
Reinforcement learning from human feedback (RLHF) has emerged as a powerful paradigm for aligning generative models with human preferences. However, applying RLHF to diffusion models remains highly feedback inefficient, as existing approaches typically require large amounts of human or reward model evaluations. This limitation reduces the practicality of diffusion RLHF in realworld settings where feedback is the primary bottleneck. In this paper, we propose two complementary strategies that substantially improve the feedback efficiency of diffusion RLHF while preserving generalization to unseen prompts. Our key observation is that reward information in diffusion trajectories is unevenly distributed: not all denoising timesteps or trajectories contribute equally to learning from a reward signal. By emphasizing informative timesteps and trajectories during optimization, we obtain more effective gradient updates. First, we introduce a per-timestep weighting scheme that reweights denoising steps during policy optimization. We theoretically connect this weighting to the optimal convergence properties of proximal policy optimization (PPO) and approximate the resulting weighting trend empirically. Second, we introduce a replay mechanism that prioritizes informative trajectories, enabling the model to reuse past samples instead of repeatedly querying new rewards. Together, these strategies significantly improve the feedback efficiency of diffusion RLHF. Under identical hyperparameter settings, our approach achieves up to a 6$\times$ improvement in sample efficiency compared to widely used diffusion RLHF baselines.
强化学习从人类反馈(RLHF)已成为将生成模型与人类偏好对齐的强大范式。然而,将RLHF应用于扩散模型仍然存在反馈效率低下的问题,因为现有方法通常需要大量的人类或奖励模型评估。这种局限性降低了扩散模型RLHF在反馈为主要瓶颈的实际应用场景中的实用性。本文提出了两种互补策略,在保持对未见提示泛化能力的同时显著提高扩散模型RLHF的反馈效率。我们的关键观察是:扩散轨迹中的奖励信息分布不均,并非所有去噪时间步或轨迹都对从奖励信号中学习具有同等贡献。通过在优化过程中强调信息丰富的时间步和轨迹,我们获得了更有效的梯度更新。首先,我们引入了一种每时间步加权方案,在策略优化过程中对去噪步骤进行重新加权。我们从理论上将这种加权与近端策略优化(PPO)的最优收敛性质联系起来,并经验性地近似了加权趋势。其次,我们引入了一种优先考虑信息丰富轨迹的重放机制,使模型能够重用过去的样本而非重复查询新的奖励。这些策略共同显著提高了扩散模型RLHF的反馈效率。在相同的超参数设置下,我们的方法与广泛使用的扩散模型RLHF基线相比,样本效率提升高达6倍。
ContrastiveCFG: Guiding Diffusion Sampling by Contrasting Positive and Negative Concepts
中文标题:ContrastiveCFG:通过对比正负概念引导扩散采样
作者:Jinho Chang, Changsun Lee, Hyungjin Chung, Jong Chul Ye
As Classifier-Free Guidance (CFG) has proven effective in conditional diffusion model sampling for improved condition alignment, many applications use a negated CFG term as a Negative Prompting (NP) to filter out unwanted features from samples. However, simply negating CFG guidance creates an inverted probability distribution, often distorting samples away from the marginal distribution. Inspired by recent advances in conditional diffusion models for inverse problems, here we present a novel method to achieve guidance toward the given condition using contrastive loss. Specifically, our guidance term aligns or repels the denoising direction based on the given condition through contrastive loss, achieving a similar guiding effect to traditional CFG for positive conditions while overcoming the limitations of existing negative guidance methods. Experimental results demonstrate that our approach effectively injects or removes the given concepts while maintaining sample quality across diverse scenarios, from simple class conditions to complex and overlapping text prompts.
鉴于无分类器引导(CFG)在条件扩散模型采样中已被证明能有效提升条件对齐效果,许多应用采用取反的CFG项作为负向提示(NP)来从样本中过滤掉不需要的特征。然而,简单地对CFG引导取反会导致概率分布倒置,往往会使样本偏离边缘分布。受条件扩散模型在逆问题处理方面最新进展的启发,本文提出了一种利用对比损失实现条件引导的新方法。具体而言,我们的引导项通过对比损失根据给定条件对齐或排斥去噪方向,在实现与传统CFG对正向条件相似引导效果的同时,克服了现有负向引导方法的局限性。实验结果表明,我们的方法在从简单类别条件到复杂重叠文本提示的多种场景中,能够有效注入或移除给定概念,同时保持样本质量。
JuZhou 1.0 Technical Report: The First Edge-Native Text-to-Image Foundation Model Trained Entirely on China-Developed AI Accelerators
中文标题:JuZhou 1.0 技术报告:首个完全基于国产AI加速器训练的边缘原生文本到图像基础模型
作者:Ce Chen, Congrui Wang, Yonglin Li, Zhenchen Wan, Mingyang Geng, Junhao Xiao, Zhengpeng Xing, Yaqing Hu, Yao Wu, Zhaoyang Qu, Long Lan, Xinwang Liu, Yingqi Peng, Shijia Li, Zufeng Zhang, Chen Ma, Jingjing Zhou, Xingyu Wang, Qilin Lu, Bin Jiang, Qilin Sun, Shanzhi Gu, Yaoguang Jin, Tongliang Liu, Kede Ma, Yifan Peng
Text-to-image (T2I) diffusion models typically require substantial computational resources and cloud infrastructure, posing significant challenges for edge deployment in terms of latency, cost, and user privacy. We present JuZhou 1.0, an ultra-lightweight T2I foundation model designed for fully offline, on-device execution. JuZhou 1.0 achieves its efficiency through four key designs: (1) a compact image-generation backbone consisting of a 0.385B-parameter denoising U-Net and a 1.90M-parameter distilled decoder, totaling approximately 0.387B parameters; (2) Rectified Flow training combined with DMD2 distillation, reducing inference to 4 sampling steps; (3) Chinese semantic alignment trained on 9M curated image-text pairs, enabling direct Chinese prompting without external translation at inference time; and (4) a training and distillation pipeline completed on domestically developed Sugon K100 AI accelerators without relying on NVIDIA GPUs for training or distillation. Despite its compact scale, the 28-step base model of JuZhou 1.0 achieves an overall GenEval score of 0.69, outperforming published baselines including SDXL (2.6B, 0.55), SD3-Medium (2B, 0.62), and IF-XL (4.3B, 0.61). We further validate the full poetry-to-image pipeline on Android and the core CLIP-U-Net-VAE generation branch on iOS. On a smartphone powered by the Snapdragon 8 Elite Gen 5 Mobile Platform, the 4-step U-Net denoising branch runs in approximately 1.6 seconds, while the full Android poetry-to-image pipeline takes 4.5 seconds with on-device prompt refinement on Xiaomi 17 Pro Max. These results position JuZhou 1.0 as a practical approach to mobile text-to-image generation and provide a concrete reference for Chinese-native generation, domestic-compute training, and fully offline on-device deployment after one-time installation.
文本到图像(T2I)扩散模型通常需要大量计算资源和云基础设施,在延迟、成本和用户隐私方面对边缘部署构成重大挑战。我们提出JuZhou 1.0,一款专为完全离线、设备端运行设计的超轻量级T2I基础模型。JuZhou 1.0通过四项关键设计实现高效能:(1)紧凑的图像生成骨干网络,包含0.385B参数的去噪U-Net和1.90M参数的蒸馏解码器,总计约0.387B参数;(2)结合DMD2蒸馏的修正流训练,将推理降至4步采样;(3)基于900万精选图文对训练的中文语义对齐,使推理时无需外部翻译即可直接使用中文提示词;(4)完全在国产曙光K100 AI加速器上完成的训练和蒸馏流程,不依赖NVIDIA GPU进行训练或蒸馏。尽管体积紧凑,JuZhou 1.0的28步基础模型在GenEval评测中达到0.69的整体分数,优于已发布的基线模型,包括SDXL(2.6B,0.55)、SD3-Medium(2B,0.62)和IF-XL(4.3B,0.61)。我们进一步在Android上验证了完整的诗歌到图像流程,并在iOS上验证了核心CLIP-U-Net-VAE生成分支。在搭载骁龙8 Elite Gen 5移动平台手机上,4步U-Net去噪分支运行时间约1.6秒,而完整的Android诗歌到图像流程在小米17 Pro Max上通过设备端提示词优化需4.5秒。这些成果使JuZhou 1.0成为移动端文本到图像生成的实用方案,并为中文原生生成、国产计算训练以及一次性安装后的完全离线设备端部署提供了具体参考。
GIRAF: Towards Generalizable Human Interactions with Articulated Objects
中文标题:GIRAF:面向可泛化人体与关节对象的交互合成
作者:Xiaohan Zhang, Sebastian Starke, Alexander Winkler, Federica Bogo, Samir Aroudj, Yuting Ye
Synthesizing realistic full-body human interactions with articulated objects is a fundamental challenge for embodied AI and graphics, with applications in robotics training and virtual agents. Existing models remain limited: some focus on simple activities with static objects, while others restrict attention to hand-only manipulation. This leaves open the problem of generating coordinated full-body motion that approaches, manipulates, and moves articulated objects in a realistic and generalizable way. The key difficulty lies in reasoning jointly about locomotion, fine-grained contact, and object articulation. Models must capture subtle hand-object correspondences that transfer across object geometries, while also producing seamless transitions from navigation to manipulation. At the same time, the scarcity of large-scale paired motion-scene data makes it difficult to generalize across diverse object positions and shapes. We introduce a text-conditioned diffusion model that addresses these challenges through three core ideas: an object-centric representation that unifies hand-object contact with object surfaces, a mixed-domain training strategy that balances locomotion and interaction, and a contact-based augmentation scheme that expands training diversity. Through experiments, our method demonstrated strong generalization to unseen object configurations, surpassing current state-of-the-art methods.
合成逼真的全身人体与关节对象交互是具身人工智能和计算机图形学领域的一项基础挑战,在机器人训练和虚拟智能体等场景中具有广泛应用。现有模型仍存在局限性:部分模型仅关注静态对象的简单活动,另一部分则将研究范围限制在仅手部操作上。这使得如何生成协调的全身运动以逼真且可泛化的方式接近、操作和移动关节对象这一问题仍未得到有效解决。核心难点在于联合推理运动能力、精细接触和对象关节运动。模型必须捕捉可跨对象几何形状迁移的细微手-对象对应关系,同时实现从导航到操作的平滑过渡。与此同时,大规模配对运动-场景数据的稀缺使得跨不同对象位置和形状的泛化变得困难。我们提出了一种文本条件扩散模型,通过三个核心思路应对这些挑战:以对象为中心的表示方法,统一手-对象接触与对象表面;混合域训练策略,平衡运动能力与交互任务;基于接触的数据增强方案,扩展训练多样性。实验表明,我们的方法在未见过的对象配置上展现出强大的泛化能力,超越了当前最先进的方法。
LightCrafter: PBR-Conditioned Video Diffusion Refinement for Controllable and Consistent Relighting
中文标题:LightCrafter:基于PBR条件的视频扩散细化实现可控一致重光照
作者:Zixin Guo, Yehonathan Litman, Yifeng He, John Miller, Chuhan Chen, Deva Ramanan
Video relighting requires balancing long-form temporal consistency with a physically grounded understanding of light transport, which depends on accurate estimation of intrinsic scene properties such as materials, geometry, and illumination. Existing methods follow two paradigms: (1) reconstruct a video's photometric properties via inverse rendering and relight them to a target illumination via forward rendering, using physically-based rendering (PBR) or a neural renderer; these suffer from noisy reconstructions and struggle with hard-to-model effects such as global illumination. (2) Frame the task as generative video-to-video translation conditioned on relighting targets (a target environment map or text); this limits relighting control and temporal stability, since diffusion models struggle to translate long-form videos, and is constrained by the availability of input/relit training pairs. We propose LightCrafter, a hybrid pipeline that reformulates video relighting as video translation of a proxy video: rather than translating the input video directly to the target, we translate a PBR rendering of the input under the target illumination to the final target. This bakes illumination targets into the PBR proxy, removing the need to teach the diffusion model illumination concepts like environment maps, and enables more intricate lighting control while naturally providing long-form temporal consistency. We show PBR renders alone already outperform some prior art but struggle with effects like global illumination; to capture these, we leverage photometric priors in video generation models by post-training CogVideoX on synthetic video pairs and real-world unpaired videos. We outperform prior state-of-the-art on existing real-world relighting benchmarks and contribute a synthetic benchmark for further analysis. We will release our dataset, benchmark, metrics, and code.
视频重光照需要在长视频时间一致性与基于物理的光传输理解之间取得平衡,而这依赖于对材质、几何和光照等内在场景属性的准确估计。现有方法遵循两种范式:(1)通过逆向渲染重建视频的光度属性,然后通过前向渲染利用基于物理的渲染(PBR)或神经渲染器将视频重光照至目标光照;这些方法受到重建噪声的困扰,难以处理全局光照等难以建模的效应。(2)将任务框定为以重光照目标(目标环境贴图或文本)为条件的生成式视频到视频转换;由于扩散模型难以转换长视频,且受限于输入/重光照训练对的可用性,这种方式限制了重光照控制和时序稳定性。我们提出LightCrafter,这是一种混合pipeline,将视频重光照重新表述为代理视频的视频转换:不是将输入视频直接转换到目标,而是将输入在目标光照下的PBR渲染转换到最终目标。这将光照目标烘焙到PBR代理中,无需教导扩散模型理解环境贴图等光照概念,同时实现了更精细的光照控制,并自然地提供了长视频时间一致性。我们证明,仅PBR渲染已经优于一些现有方法,但在全局光照等效应上存在困难;为捕捉这些效果,我们通过在合成视频对和现实世界无配对视频上对CogVideoX进行后训练来利用视频生成模型的光度先验。我们在现有的现实世界重光照基准上超越了先前最先进的方法,并贡献了一个合成基准以供进一步分析。我们将发布数据集、基准、指标和代码。
SAGA: Stable Acceleration Guidance for Autoregressive Video Generation
中文标题:SAGA:自回归视频生成的稳定加速引导方法
作者:Thanh-Nhan Vo, Trong-Thuan Nguyen, Trung-Hoang Le, Tam V. Nguyen, Minh-Triet Tran
Autoregressive video diffusion enables efficient streaming and long-horizon video generation, but repeatedly reusing generated latents as causal context can amplify temporal errors, resulting in flickering, motion jitter, and structural drift. In this paper, we investigate this failure mode from a spectral kinematic perspective and identify discrete latent acceleration as an effective signal for revealing unstable high-frequency temporal perturbations. To this end, we propose SAGA, a training-free \textbf{\textit{s}}table \textbf{\textit{a}}cceleration \textbf{\textit{g}}uidance approach for \textbf{\textit{a}}utoregressive video generation. SAGA integrates an acceleration domain spectral guidance objective based on finite-window Slepian projections with a structured autoregressive noise initialization strategy that suppresses short-range temporal correlations while preserving long-range motion structure. Without retraining or modifying the backbone, SAGA can be directly applied to existing chunk-wise autoregressive diffusion models, which is the prevalent setting for high-quality generation. Extensive experiments show that SAGA consistently improves temporal quality across multiple autoregressive diffusion models. On Self-Forcing, SAGA improves Temporal Quality from 97.30 to 97.91 and Image Quality from 69.60 to 70.51. Moreover, spectral analysis and human preference studies demonstrate that SAGA reduces temporal instability while maintaining visual fidelity.
自回归视频扩散能够实现高效的流式生成和长时序视频生成,但重复使用已生成的潜在表示作为因果上下文可能会放大时间误差,导致闪烁、运动抖动和结构漂移等问题。本文从光谱运动学视角研究这一失效模式,并识别出离散潜在加速度作为揭示不稳定高频时间扰动的有效信号。基于此,本文提出了SAGA,一种无需训练的自回归视频生成稳定加速引导方法。SAGA将基于有限窗口Slepian投影的加速度域谱引导目标与结构化自回归噪声初始化策略相结合,在保留长程运动结构的同时抑制短程时间相关性。SAGA无需重新训练或修改主干网络,可直接应用于现有的分块自回归扩散模型,这也是高质量视频生成的主流设置。大量实验表明,SAGA在多个自回归扩散模型上持续提升时间质量。在Self-Forcing数据集上,SAGA将时间质量从97.30提升至97.91,将图像质量从69.60提升至70.51。此外,光谱分析和人类偏好研究表明,SAGA在保持视觉保真度的同时减少了时间不稳定性。
Multimodal 3D LUT Generation via StatLUT with Statistical Features for Photorealistic Style Transfer
中文标题:基于StatLUT与统计特征的多模态3D LUT生成方法用于逼真风格迁移
作者:Yifan Wang, Zhixiang Hao, Yu Wang, Congchao Zhu
Photorealistic Style Transfer (PST) aims to transfer the color and tonal style of a reference to a content image while strictly preserving its structural integrity. However, existing deep learning-based methods inherently suffer from semantic entanglement caused by pre-trained image encoders, leading to unnatural spatial distortions. Moreover, current pixel-level mapping paradigms often ignore color gamut topology, resulting in color banding, while also lacking the multimodal capability for intuitive text-driven control. To address these bottlenecks, we propose StatLUT, an innovative multimodal framework for 3D LUT generation. First, we bypass traditional encoders and introduce a Lab-Extractor to derive spatially-agnostic statistical features, fundamentally decoupling color distributions from structural semantics to ensure artifact-free rendering. Second, we formulate LUT generation as a Transformer-based Seq2Seq translation task, utilizing a Multi-dimensional Residual Mapper (MR-Mapper) to predict topologically smooth 3D LUTs. Finally, to break the single-modal barrier, we propose the H-Diffuser, a lightweight Diffusion Transformer that directly synthesizes statistical features from natural language prompts, enabling flexible text-driven color grading. Extensive experiments on standard benchmarks demonstrate that StatLUT significantly outperforms state-of-the-art methods in both visual quality and quantitative metrics, pioneering a highly robust and flexible paradigm for multimodal photorealistic style transfer.
逼真风格迁移(PST)旨在将参考图像的色彩与影调风格迁移至内容图像,同时严格保持其结构完整性。然而,现有基于深度学习的方法由于预训练图像编码器而固有地遭受语义纠缠问题,导致不自然的空间畸变。此外,当前像素级映射范式往往忽略色域拓扑结构,导致色带现象产生,同时缺乏用于直观文本驱动控制的多模态能力。为解决这些瓶颈,我们提出了StatLUT,一个用于3D LUT生成的创新多模态框架。首先,我们绕过传统编码器,引入Lab-Extractor以提取空间无关的统计特征,从根本上将色彩分布与结构语义解耦,确保无伪影渲染。其次,我们将LUT生成建模为基于Transformer的Seq2Seq翻译任务,利用多维残差映射器(MR-Mapper)预测拓扑平滑的3D LUT。最后,为打破单模态壁垒,我们提出了H-Diffuser,这是一个轻量级扩散Transformer,可直接从自然语言提示合成统计特征,实现灵活的文本驱动色彩调校。在标准基准数据集上的广泛实验表明,StatLUT在视觉质量和定量指标上均显著优于最先进的方法,开创了高度鲁棒且灵活的多模态逼真风格迁移新范式。
SkelGen4D: Weakly-Supervised Skeleton-Based 4D Generation for Text-Driven Mesh Animation
中文标题:SkelGen4D:面向文本驱动网格动画的弱监督基于骨骼的4D生成
作者:Hao Feng, Zhi Zuo, Jia-Hui Pan, Ka-Hei Hui, Zhengzhe Liu, Dian Zhang, Haoran Xie, Bin Sheng, Jingyu Hu
We study 4D generation to synthesize temporally coherent sequences of 3D geometry for animation and content creation. In contrast to existing SDS-based optimization methods and video-driven animation approaches, we adopt a skeleton-driven animation framework aligned with standard industrial pipelines, which enables explicit control and editing. To this end, we propose SkelGen4D, a weakly supervised feed-forward framework for text-driven mesh animation that generates explicit skeleton motions without requiring per-frame skeleton annotations. SkelGen4D first recovers temporally consistent pseudo-skeletons from animated meshes via differentiable fitting, and then generates text-conditioned skeleton motion sequences in a feed-forward manner, further refined with Motion-GRPO to ensure temporally coherent, physically plausible, and articulated animation. We evaluate our method on two large-scale benchmarks, Truebones Zoo and Diffusion4D. Our results show that our weakly supervised skeleton modeling matches or surpasses fully supervised baselines while scaling to diverse object categories for high-quality text-driven mesh animation. Further, our method supports flexible motion editing and is aligned with standard animation production pipelines.
我们研究4D生成任务,旨在为动画和内容创作合成时间一致的3D几何序列。与现有的基于SDS的优化方法和视频驱动动画方法不同,我们采用与标准工业管线对齐的骨骼驱动动画框架,从而实现显式控制和编辑。为此,我们提出了SkelGen4D,一个面向文本驱动网格动画的弱监督前馈框架,无需每帧骨骼标注即可生成显式骨骼运动。SkelGen4D首先通过可微拟合从动画网格中恢复时间一致的伪骨骼,然后以前馈方式生成文本条件的骨骼运动序列,并进一步通过Motion-GRPO进行细化,以确保时间一致、物理合理且连贯的动画。我们在Truebones Zoo和Diffusion4D两个大规模基准上评估了我们的方法。结果表明,我们的弱监督骨骼建模在匹配或超越全监督基线的同时,能够扩展到多样化的物体类别,实现高质量的文本驱动网格动画。此外,我们的方法支持灵活的运动编辑,并与标准动画制作管线保持一致。
HumanForge: A Human-Centric Deepfake Video Benchmark with Multi-Agent Forgery Rationales
中文标题:HumanForge:一个以人为中心的深度伪造视频基准测试与多智能体伪造推理
作者:Wenbo Xu, Zhimin Chen, Xiaojie Liang, Hengrui Liu, Wei Lu
Rapid advancements in video diffusion models and temporal editing tools have enabled the generation of highly realistic human-centric videos, posing unprecedented challenges to digital content forensics. Existing benchmarks primarily focus on either face-swapping or global text-to-video synthesis, overlooking the crucial dimensions of human-object or human-human interactions and multi-modal alignment. To address these limitations, we introduce HumanForge, a unified, large-scale, and multi-paradigm human-centric video forgery dataset. To construct and annotate this dataset without labor-intensive manual labeling or hallucinated monolithic prompts, we propose Gen2Anno, a modular active multi-agent pipeline built on LangGraph. Gen2Anno coordinates six specialized agents-ranging from source profiling to MoE-based reference analysis and closed-loop forensic verification-to generate over 18K high-fidelity video segments and produce structured, contrastive omni-annotations containing binary decisions, fine-grained artifact categories, and spatio-temporal localization. Extensive benchmarks using state-of-the-art traditional detectors and Large Multimodal Models (LMMs) demonstrate the significant challenges of zero-shot generalization and fine-grained reasoning on HumanForge. Code and dataset will be publicly released.
视频扩散模型和时间编辑工具的快速发展使得生成高度逼真的人为中心视频成为可能,这对数字内容取证带来了前所未有的挑战。现有的基准测试主要关注换脸或全局文本到视频合成,忽略了人与物体或人与人交互以及多模态对齐的关键维度。为解决这些局限性,我们提出了HumanForge,这是一个统一、大规模、多范式的人为中心视频伪造数据集。为了在无需繁重人工标注或虚构整体提示的情况下构建和标注该数据集,我们提出了Gen2Anno,一个基于LangGraph的模块化主动多智能体流水线。Gen2Anno协调六个专业智能体——从源分析到基于MoE的参考分析,再到闭环取证验证——生成了超过18K个高质量视频片段,并产生了结构化对比全注释,包含二元决策、细粒度伪造类别和时空定位。使用最先进的传统检测器和大型多模态模型(LMMs)进行的广泛基准测试表明,HumanForge上零样本泛化和细粒度推理具有显著挑战。代码和数据集将公开发布。
OPSD-V: On-Policy Self-Distillation for Post-Training Few-Step Autoregressive Video Generators
中文标题:OPSD-V: 面向后训练少步自回归视频生成器的策略自蒸馏方法
作者:Hongyu Liu, Chun Wang, Feng Gao, Xuanhua He, Yue Ma, Ziyu Wan, Yong Zhang, Xiaoming Wei, Qifeng Chen
We propose OPSD-V, an on-policy self-distillation paradigm for post-training few-step autoregressive (AR) video diffusion models. Existing few-step AR video generators can produce long videos with low latency, but still suffer from error accumulation and weakened motion dynamics during long autoregressive rollout. OPSD-V reduces long-horizon degradation while preserving the original few-step inference path. The key idea is to introduce real long-video data as temporal context during training and use it to provide dense trajectory-level supervision. Specifically, the student follows the exact inference-time rollout, generating each chunk conditioned on its own previously generated KV cache. In parallel, the teacher is evaluated at the same student-visited denoising states, but uses a cleaner AR-consistent temporal cache in which older history can be replaced by real-video context. This provides dense denoising-level corrective targets under on-policy AR cache dynamics, without changing the sampler, number of denoising steps, or inference-time cache mechanism. We apply OPSD-V to representative few-step AR video models, including Self-Forcing and LongLive. Experiments show consistent improvements in visual quality, motion dynamics, and VBenchLong scores. A user study with 10 participants comparing 20 video pairs shows that OPSD-V is preferred over the base models in 66.0% of overall-preference judgments (82.5% excluding ties).
本文提出OPSD-V,一种面向后训练少步自回归视频扩散模型的策略自蒸馏范式。现有的少步自回归视频生成器虽然能够以低延迟生成较长的视频,但在长时自回归展开过程中仍面临误差累积和运动动态减弱的问题。OPSD-V在保留原有少步推理路径的同时,降低了长程退化问题。其核心思想是在训练过程中引入真实长视频数据作为时序上下文,并利用其提供密集的轨迹级监督。具体而言,学生模型遵循推理时的展开方式,基于自身生成的KV缓存来生成每个片段;与此同时,教师模型在学生访问的相同去噪状态下进行评估,但使用更清晰的自回归一致时序缓存,其中较旧的历史可以被真实视频上下文所替代。这在保持采样器、去噪步数及推理时缓存机制不变的情况下,基于策略自回归缓存动态提供了密集的去噪级纠正目标。本文将OPSD-V应用于代表性的少步自回归视频模型,包括Self-Forcing和LongLive。实验表明,OPSD-V在视觉质量、运动动态和VBenchLong分数方面均取得一致的提升。针对20对视频进行的10人用户研究表明,在总体偏好判断中,OPSD-V在66.0%的案例中更受青睐(排除平局后为82.5%)。
LongE2V: Long-Horizon Event-based Video Reconstruction, Prediction, and Frame Interpolation with Video Diffusion Models
中文标题:LongE2V:基于视频扩散模型的长时序事件相机视频重建、预测与帧插值
作者:Cheng-De Fan, Chun-Wei Tuan Mu, Chen-Wei Chang, Chin-Yang Lin, Kun-Ru Wu, Yu-Chee Tseng, Yu-Lun Liu
Recovering high-quality video from sparse event streams is a challenging task. Regression methods often blur textures, while existing generative models struggle with long-term stability. We propose LongE2V, a novel approach that leverages pre-trained video diffusion priors to jointly handle event-based video reconstruction, prediction, and frame interpolation. By fine-tuning a foundational video model, our approach achieves high data efficiency and superior perceptual quality. We introduce Autoregressive Unrolling and Adaptive Context Switching to mitigate temporal drift in extremely long sequences. We also propose Reencoding Alignment with Cross Residual Correction to ensure precise bidirectional consistency during frame interpolation. Furthermore, Event Voxel Density Augmentation ensures robustness across varying sensor resolutions. Extensive experiments on real-world benchmarks demonstrate that LongE2V outperforms state-of-the-art methods across all three tasks, exhibiting exceptional temporal coherence and zero-shot generalization. Project page: https://cdfan0627.github.io/LongE2V-page/
从稀疏事件流中恢复高质量视频是一项具有挑战性的任务。回归方法往往会导致纹理模糊,而现有生成模型在长期稳定性方面存在不足。我们提出LongE2V,这是一种利用预训练视频扩散先验的新方法,可联合处理基于事件的视频重建、预测和帧插值。通过微调基础视频模型,我们的方法实现了较高的数据效率和卓越的感知质量。我们引入自回归展开和自适应上下文切换来缓解极长序列中的时间漂移问题。此外,我们还提出重新编码对齐与交叉残差校正,以确保帧插值过程中的精确双向一致性。事件体素密度增强则确保了模型在不同传感器分辨率下的鲁棒性。在真实世界基准数据集上的广泛实验表明,LongE2V在所有三项任务上均优于最先进的方法,表现出优异的时间一致性和零样本泛化能力。项目主页:https://cdfan0627.github.io/LongE2V-page/
ARDY: Autoregressive Diffusion with Hybrid Representation for Interactive Human Motion Generation
中文标题:ARDY:用于交互式人体运动生成的自回归扩散混合表示框架
作者:Kaifeng Zhao, Mathis Petrovich, Haotian Zhang, Tingwu Wang, Siyu Tang, Davis Rempe
Generating realistic 3D human motions in real-time within interactive applications is key for animation, simulation, and humanoid robotics. While recent offline motion generation approaches offer precise control via text and kinematic constraints, they lack the inference speed required for interactive settings. Conversely, existing online methods enable real-time synthesis but often sacrifice controllability or struggle with complex text semantics and long-horizon goals due to limited context windows. In this work, we introduce ARDY, a streaming generation framework that bridges this gap by enabling high-fidelity motion generation controllable via online text prompts and flexible kinematic constraints. ARDY employs a hybrid representation that combines explicit root features with a latent body embedding, balancing precise trajectory control with efficient generative learning. We propose a two-stage autoregressive transformer denoiser that features variable history context and supports conditioning on flexible, long-horizon kinematic constraints. By training on a large-scale motion capture dataset and being directly conditioned on text labels and kinematic constraints sampled from ground truth poses, ARDY natively learns controllable generation that supports online prompting and flexible long-horizon goals. Extensive evaluations on the HumanML3D benchmark and the large-scale, high-fidelity Bones Rigplay dataset demonstrate ARDY's high motion quality and constraint adherence, validating the efficacy of our key architectural decisions. Finally, we demonstrate the method&x27;s practical versatility through an interactive demo featuring dynamic text control, diverse keyframe pose constraints, path following, and interactive locomotion control via mouse and keyboard. Supplementary video results, code, and model releases can be found at https://research.nvidia.com/labs/sil/projects/ardy/.
在交互式应用中实时生成逼真的3D人体运动对于动画、仿真和类人机器人学至关重要。尽管近期离线运动生成方法可通过文本和运动学约束实现精确控制,但缺乏交互场景所需的推理速度。相反,现有在线方法虽能实现实时合成,但往往牺牲了可控性,或因有限的时间上下文窗口而难以处理复杂文本语义和长时序目标。本工作提出ARDY,一个流式生成框架,通过支持高保真运动生成并可通过在线文本提示和灵活的运动学约束进行控制来弥合这一差距。ARDY采用结合显式根节点特征与潜在身体嵌入的混合表示,在精确轨迹控制与高效生成学习之间取得平衡。我们提出了一种具有可变历史上下文并支持灵活长时序运动学约束条件的两阶段自回归Transformer去噪器。通过在大规模动作捕捉数据集上训练,并直接以文本标签和从真实姿态采样的运动学约束为条件,ARDY原生学习可控生成,支持在线提示和灵活的长时序目标。在HumanML3D基准和大规模高保真Bones Rigplay数据集上的广泛评估表明,ARDY具有高运动质量和约束依从性,验证了我们关键架构决策的有效性。最后,我们通过一个交互式演示展示了该方法的实际通用性,包括动态文本控制、多样化关键帧姿态约束、路径跟随以及通过鼠标和键盘实现的交互式运动控制。补充视频结果、代码和模型发布于https://research.nvidia.com/labs/sil/projects/ardy/。
DeltaDeno: Zero-Shot Anomaly Generation via Delta-Denoising Attribution
中文标题:DeltaDeno:基于Delta去噪归因的零样本异常生成
作者:Chaoran Xu, Chengkan Lv, Qiyu Chen, Yunkang Cao, Feng Zhang, Zhengtao Zhang
Anomaly generation is often framed as few-shot fine-tuning with anomalous samples, which contradicts the scarcity that motivates generation and tends to overfit category priors. We tackle the setting where no real anomaly samples or training are available. We propose Delta-Denoising (\textbf{DeltaDeno}), a training-free zero-shot anomaly generation method that localizes and edits defects by contrasting two diffusion branches driven by a minimal prompt pair under a shared schedule. By accumulating per-step denoising deltas into an image-specific localization map, we obtain a mask to guide the latent inpainting during later diffusion steps and preserve the surrounding context while generating realistic local defects. To improve stability and control, DeltaDeno performs token-level prompt refinement that aligns shared content and strengthens anomaly tokens, and applies a spatial attention bias restricted to anomaly tokens in the predicted region. Experiments on public datasets show that DeltaDeno achieves great generation, realism and consistent gains in downstream detection performance. Code will be made publicly available at https://github.com/CROVO1026/DeltaDeno.
异常生成通常被建模为使用异常样本进行少样本微调,这与生成任务所基于的稀缺性假设相矛盾,且容易导致类别先验过拟合。本研究针对无真实异常样本且无训练可用的场景进行探索。我们提出Delta去噪(DeltaDeno),一种无需训练即可实现零样本异常生成的方法。该方法通过在共享调度下利用最小提示对驱动两个扩散分支进行对比,从而定位并编辑缺陷。通过将每步去噪差分累积为图像特定的定位图,我们获得了一个掩码来引导后续扩散步骤中的潜在修复过程,从而在生成逼真局部缺陷的同时保留周围上下文。为提升稳定性和可控性,DeltaDeno执行标记级提示优化以对齐共享内容并增强异常标记,并在预测区域对异常标记施加空间注意力偏置限制。在公开数据集上的实验表明,DeltaDeno在生成质量、真实感以及下游检测性能提升方面均取得了优异表现。代码将公开于 https://github.com/CROVO1026/DeltaDeno。
Elastic3D: Controllable Stereo Video Conversion with Guided Latent Decoding
中文标题:Elastic3D:基于引导潜在解码的可控立体视频转换
作者:Nando Metzger, Prune Truong, Goutam Bhat, Konrad Schindler, Federico Tombari
The growing demand for immersive 3D content calls for automated monocular-to-stereo video conversion. We present Elastic3D, a controllable, direct end-to-end method for upgrading a conventional video to a binocular one. Our approach, based on (conditional) latent diffusion, avoids artifacts due to explicit depth estimation and warping. The key to its high-quality stereo video output is a novel, guided VAE decoder that ensures sharp and epipolar-consistent stereo video output. Moreover, our method gives the user control over the strength of the stereo effect (more precisely, the disparity range) at inference time, via an intuitive, scalar tuning knob. Experiments on three different datasets of real-world stereo videos show that our method outperforms both traditional warping-based and recent warping-free baselines and sets a new standard for reliable, controllable stereo video conversion. Please check the project page for the video samples https://elastic3d.github.io.
日益增长的沉浸式3D内容需求催生了自动化单目到立体视频转换的研究。我们提出了Elastic3D,一种可控的、直接的端到端方法,用于将传统视频升级为双目视频。我们的方法基于(条件)潜在扩散模型,避免了显式深度估计和变形操作导致的伪影。其高质量立体视频输出的关键在于一种新颖的引导VAE解码器,能够确保输出清晰且对极一致的立体视频。此外,我们的方法通过一个直观的标量调节旋钮,使用户能够在推理时控制立体效果强度(更准确地说是视差范围)。在三个真实世界立体视频数据集上的实验表明,我们的方法优于传统基于变形的基线方法以及近期无变形方法,为可靠、可控的立体视频转换树立了新的标准。请访问项目页面查看视频样本 https://elastic3d.github.io。
HairWeaver: Few-Shot Photorealistic Hair Motion Synthesis with Sim-to-Real Guided Video Diffusion
中文标题:HairWeaver:基于仿真到真实引导视频扩散的少样本照片级真实感头发运动合成
作者:Di Chang, Ji Hou, Aljaz Bozic, Assaf Neuberger, Felix Juefei-Xu, Olivier Maury, Gene Wei-Chin Lin, Tuur Stuyck, Doug Roble, Mohammad Soleymani, Stephane Grabli
We present HairWeaver, a diffusion-based pipeline that animates a single human image with realistic and expressive hair dynamics. While existing methods successfully control body pose, they lack specific control over hair, and as a result, fail to capture the intricate hair motions, resulting in stiff and unrealistic animations. HairWeaver overcomes this limitation using two specialized modules: a Motion-Context-LoRA to integrate motion conditions and a Style-Alignment-LoRA to preserve the subject's photoreal appearance across different data domains. These lightweight components are designed to guide a video diffusion backbone while maintaining its core generative capabilities. By training on a specialized dataset of dynamic human motion generated from a CG simulator, HairWeaver affords fine control over hair motion and ultimately learns to produce highly realistic hair that responds naturally to movement. Comprehensive evaluations demonstrate that our approach sets a new state of the art, producing lifelike human hair animations with dynamic details.
我们提出了HairWeaver,一个基于扩散的管道,能够为单个人物图像添加真实且富有表现力的头发动态效果。现有方法虽然成功控制了身体姿态,但缺乏对头发的具体控制,因此无法捕捉复杂的头发运动,导致动画生硬且不真实。HairWeaver通过两个专门模块克服了这一限制:Motion-Context-LoRA用于整合运动条件,以及Style-Alignment-LoRA用于在不同数据域中保持拍摄对象的照片级真实感。这些轻量级组件旨在引导视频扩散主干网络,同时保持其核心生成能力。通过在由CG模拟器生成的动态人物运动专业数据集上进行训练,HairWeaver实现了对头发运动的精细控制,并最终学会了生成对人物动作自然响应的高度逼真头发。全面评估表明,我们的方法达到了新的先进水平,能够生成具有动态细节的逼真人物头发动画。
Deep Sprite-based Image Models: An Analysis
中文标题:基于深度精灵的图像模型:分析与研究
作者:Zeynep Sonat Baltac{\i}, Romain Loiseau, Mathieu Aubry
While foundation models drive steady progress in image segmentation and diffusion algorithms compose always more realistic images, the seemingly simple problem of identifying recurrent patterns in a collection of images remains very much open. In this paper, we focus on sprite-based image decomposition models, which have shown some promise for clustering and image decomposition and are appealing because of their high interpretability. These models come in different flavors, need to be tailored to specific datasets, and struggle to scale to images with many objects. We dive into the details of their design, identify their core components, and perform an extensive analysis on clustering benchmarks. We leverage this analysis to propose a deep sprite-based image decomposition method that performs on par with state-of-the-art unsupervised class-aware image segmentation methods on the standard CLEVR benchmark, scales linearly with the number of objects, identifies explicitly object categories, and fully models images in an easily interpretable way.
尽管基础模型推动了图像分割的稳步发展,扩散算法也能够合成越来越逼真的图像,但识别图像集合中重复模式这一看似简单的问题仍然悬而未决。本文聚焦于基于精灵的图像分解模型,该类模型在聚类和图像分解方面展现出了一定的潜力,因其高度可解释性而颇具吸引力。这些模型具有多种类型,需要针对特定数据集进行定制,且难以扩展到包含大量对象的图像。我们深入研究其设计细节,识别核心组件,并在聚类基准上进行了广泛的分析。基于这一分析,我们提出了一种深度基于精灵的图像分解方法,该方法在标准CLEVR基准上与最先进的无监督类别感知图像分割方法性能相当,能够随对象数量线性扩展,明确识别对象类别,并以完全可解释的方式对图像进行建模。
PanoImager: Geometry-Guided Novel View Synthesis and Reconstruction from Sparse Panoramic Views
中文标题:PanoImager:基于几何引导的稀疏全景视图新视图合成与重建
作者:Zhisong Xu, Takeshi Oishi
Panoramic sensing offers wide field-of-view coverage, yet 3D reconstruction from sparse panoramas remains challenging under rotation-dominant, weak-parallax motion. In such regimes, SfM/SLAM initialization is often ill-conditioned and unreliable. We present PanoImager, an SfM-free framework that combines feed-forward pose/depth priors, geometry-conditioned diffusion view completion, and depth-guided 3DGS optimization. Given only a few panoramic images, PanoImager decomposes them into local perspective views, synthesizes auxiliary observations to enrich sparse evidence, and stabilizes Gaussian optimization for improved cross-view consistency. Experiments on multiple benchmarks show improved stability under extreme sparsity, suggesting PanoImager as an offline/background component for map refinement when SfM/SLAM fails to initialize.
全景感知提供了广阔的视场覆盖,然而在旋转主导、弱视差的运动条件下,从稀疏全景图进行3D重建仍然具有挑战性。在这类场景中,SfM/SLAM的初始化往往病态且不可靠。我们提出了PanoImager,一个无需SfM的框架,结合了前馈姿态/深度先验、几何条件扩散视图补全以及深度引导的3DGS优化。给定少量全景图像,PanoImager将它们分解为局部透视视图,合成辅助观测以丰富稀疏证据,并稳定高斯优化以改善跨视图一致性。在多个基准数据集上的实验表明,在极端稀疏条件下具有更好的稳定性,证明了PanoImager可作为SfM/SLAM初始化失败时的离线/后台地图优化组件。
The Surprising Effectiveness of Video Diffusion Models for Hand Motion Reconstruction
中文标题:视频扩散模型在手部运动重建中的惊人有效性
作者:Yuxi Wang, Chengkai Jin, Yufei Liu, Wenqi Ouyang, Tianyi Wei, Zhiwei Zeng, Siyuan Huang, Zhiqi Shen, Xingang Pan
4D hand motion reconstruction from egocentric video is bottlenecked by clear limitations of existing methods: image-based pipelines depend on a detector that fails under heavy occlusion, while video-based methods rely on temporal modules learned only from scarce hand-pose annotations, a narrow signal insufficient to model motion dynamics, occlusion reasoning, and hand-object interaction. These capabilities, however, are exactly what video generative models must implicitly acquire when trained to synthesize coherent video at internet scale. Motivated by this, we present ViDiHand, which leverages the representations of a pretrained video diffusion model to reconstruct 4D two-hand pose. We adapt it via a hand-overlay rendering objective that specializes its features for hands while preserving its world priors. A decoder then recovers metric-scale pose from the adapted features. The whole pipeline operates directly on full frames--no detector, no infiller, and no test-time optimization. On ARCTIC, HOT3D, and HOI4D, ViDiHand substantially outperforms prior methods, establishing video diffusion models as a powerful new foundation for hand motion reconstruction and a promising route to scalable in-the-wild data collection for embodied AI. Project page: https://vidihand.github.io.
从第一人称视频进行4D手部运动重建受限于现有方法的明显缺陷:基于图像的流程依赖于在严重遮挡下失效的检测器,而基于视频的方法仅依赖稀缺的手部姿态标注来学习时序模块,这种狭隘的信号不足以建模运动动力学、遮挡推理和手-物体交互。然而,这些能力恰恰是视频生成模型在互联网规模数据上训练以合成连贯视频时必须隐式获取的。受此启发,我们提出了ViDiHand,它利用预训练视频扩散模型的表示来重建4D双手姿态。我们通过一种手部叠加渲染目标对其进行适配,使其特征专门化用于手部,同时保留其世界先验。随后,一个解码器从适配后的特征中恢复度量尺度的姿态。整个流程直接在完整帧上运行——无需检测器,无需填充器,无需测试时优化。在ARCTIC、HOT3D和HOI4D数据集上,ViDiHand显著优于之前的方法,确立了视频扩散模型作为手部运动重建的强大新基础,以及为具身智能进行可扩展的野外数据收集的有前景途径。
您好!我需要指出一个小问题:您提供的这两篇论文实际上并不属于Image Compression(图像压缩)类别。
- Real-World Blind Super-Resolution via Feature Matching with Implicit High-Resolution Priors — 这是一篇关于盲超分辨率(Blind Super-Resolution)的论文,聚焦于真实场景下的图像超分辨率重建问题。
- Collaborative Synthetic Data Generation for Knowledge Transfer in Federated Learning — 这是一篇关于联邦学习(Federated Learning)的论文,讨论如何通过合成数据进行知识迁移。
很抱歉,由于列表中没有Image Compression相关的论文,我无法按要求撰写该分类的每日总览和推荐列表。
如果您能提供正确的Image Compression类别论文列表,我很乐意帮您完成这项任务!
Collaborative Synthetic Data Generation for Knowledge Transfer in Federated Learning
中文标题:联邦学习中用于知识迁移的协作式合成数据生成
作者:Maximilian Andreas Hoefler, Karsten Mueller, Wojciech Samek
One-shot federated learning (OSFL) addresses the communication overhead of federated learning by limiting training to a single round, but doing so without sacrificing model quality is non-trivial, particularly when client data distributions diverge. Recent work has addressed this challenge by aggregating client knowledge on the server through the construction of transferable synthetic datasets or distillates. However, most of these methods lack formal privacy guarantees, leaving a gap in jointly achieving low communication, robustness to heterogeneity, and rigorous privacy. We propose FedKT-CSD (Federated Knowledge Transfer via Collaborative Synthetic Data), a framework inspired by neural image compression that closes this gap by leveraging publicly pretrained autoencoders as a shared latent space. Each client encodes its private data in a single forward pass, computes class-conditional latent statistics, and transmits these to the server. The server aggregates these statistics via secure aggregation, adds calibrated differential privacy noise, and decodes a synthetic dataset for training a global model and further downstream tasks. This design provides formal $(\varepsilon,\delta)$-differential privacy by construction, while keeping client-side computation and communication lightweight. Despite operating under privacy constraints, FedKT-CSD is competitive with and even outperforms non-private baselines across diverse datasets and heterogeneity settings, and scales to a large number of clients. Our code is available at: https://github.com/an7123/FedKT-CSD
单轮联邦学习(OSFL)通过将训练限制在单轮来解决通信开销问题,但在不牺牲模型质量的前提下实现这一点并非易事,尤其当客户端数据分布存在差异时。最近的研究通过在服务器上聚合客户端知识来解决这一挑战,方法是构建可迁移的合成数据集或蒸馏物。然而,这些方法大多数缺乏正式的隐私保证,在实现低通信、异构性鲁棒性和严格隐私的联合目标方面存在差距。我们提出了FedKT-CSD(通过协作式合成数据的联邦知识迁移),这是一个受神经图像压缩启发的框架,利用公开预训练的自编码器作为共享潜在空间来弥补这一差距。每个客户端在单次前向传播中编码其私有数据,计算类别条件潜在统计量,并将其传输到服务器。服务器通过安全聚合聚合这些统计量,添加校准的差分隐私噪声,并解码出合成数据集用于训练全局模型和下游任务。该设计在架构上提供了正式的(ε,δ)-差分隐私保证,同时保持客户端的计算和通信轻量级。尽管在隐私约束下运行,FedKT-CSD在多样化的数据集和异构性设置中与基线方法相比具有竞争力,甚至表现更优,并能扩展到大量客户端。我们的代码可访问:https://github.com/an7123/FedKT-CSD
Real-World Blind Super-Resolution via Feature Matching with Implicit High-Resolution Priors
中文标题:基于特征匹配与隐式高分辨率先验的实用盲超分辨率
作者:Chaofeng Chen, Xinyu Shi, Yipeng Qin, Xiaoming Li, Xiaoguang Han, Tao Yang, Shihui Guo
A key challenge of real-world image super-resolution (SR) is to recover the missing details in low-resolution (LR) images with complex unknown degradations (e.g., downsampling, noise and compression). Most previous works restore such missing details in the image space. To cope with the high diversity of natural images, they either rely on the unstable GANs that are difficult to train and prone to artifacts, or resort to explicit references from high-resolution (HR) images that are usually unavailable. In this work, we propose Feature Matching SR (FeMaSR), which restores realistic HR images in a much more compact feature space. Unlike image-space methods, our FeMaSR restores HR images by matching distorted LR image features to their distortion-free HR counterparts in our pretrained HR priors, and decoding the matched features to obtain realistic HR images. Specifically, our HR priors contain a discrete feature codebook and its associated decoder, which are pretrained on HR images with a Vector Quantized Generative Adversarial Network (VQGAN). Notably, we incorporate a novel semantic regularization in VQGAN to improve the quality of reconstructed images. For the feature matching, we first extract LR features with an LR encoder consisting of several Swin Transformer blocks and then follow a simple nearest neighbour strategy to match them with the pretrained codebook. In particular, we equip the LR encoder with residual shortcut connections to the decoder, which is critical to the optimization of feature matching loss and also helps to complement the possible feature matching errors. Experimental results show that our approach produces more realistic HR images than previous methods. Codes are released at https://github.com/chaofengc/FeMaSR.
真实世界图像超分辨率(SR)的一个关键挑战是在具有复杂未知退化(如下采样、噪声和压缩)的低分辨率(LR)图像中恢复缺失的细节。以往的大多数方法在图像空间中恢复这些缺失细节。为了应对自然图像的高度多样性,这些方法要么依赖难以训练且容易产生伪影的不稳定生成对抗网络(GANs),要么借助通常不可用的高分辨率(HR)图像作为显式参考。本工作提出了特征匹配超分辨率(FeMaSR),它在更为紧凑的特征空间中恢复真实的HR图像。与图像空间方法不同,FeMaSR通过将退化的LR图像特征与预训练的HR先验中的无退化HR对应特征进行匹配,并将匹配后的特征解码以获得真实的HR图像来恢复HR图像。具体而言,我们的HR先验包含一个离散的feature码本及其关联解码器,这些是在HR图像上通过向量量化生成对抗网络(VQGAN)预训练得到的。值得注意的是,我们在VQGAN中引入了一种新的语义正则化方法来提高重建图像的质量。对于特征匹配,我们首先使用由多个Swin Transformer块组成的LR编码器提取LR特征,然后采用简单的最近邻策略将它们与预训练的码本进行匹配。特别地,我们为LR编码器配备了到解码器的残差shortcut连接,这对于特征匹配损失的优化至关重要,同时有助于弥补可能的特征匹配误差。实验结果表明,我们的方法比以往方法能够生成更加真实的HR图像。代码已发布于https://github.com/chaofengc/FeMaSR。
今日未找到该分类的匹配论文。
今日未找到该分类的匹配论文。
今日未找到该分类的匹配论文。
今日未找到该分类的匹配论文。