https://arxiv.org/api/8AgvpyA48z1+PitI6FUGoKy9HE0 2026-06-14T23:24:09Z 9323 585 15 http://arxiv.org/abs/2603.27573v1 SPREAD: Spatial-Physical REasoning via geometry Aware Diffusion 2026-03-29T08:15:17Z Automated 3D scene generation is pivotal for applications spanning virtual reality, digital content creation, and Embodied AI. While computer graphics prioritizes aesthetic layouts, vision and robotics demand scenes that mirror real-world complexity which current data-driven methods struggle to achieve due to limited unstructured training data and insufficient spatial and physical modeling. We propose SPREAD, a diffusion-based framework that jointly learns spatial and physical relationships through a graph transformer, explicitly conditioning on posed scene point clouds for geometric awareness. Moreover, our model integrates differentiable guidance for collision avoidance, relational constraint, and gravity, ensuring physically coherent scenes without sacrificing relational context. Our experiments on 3D-FRONT and ProcTHOR datasets demonstrate state-of-the-art performance in spatial-relational reasoning and physical metrics. Moreover, \ours{} outperforms baselines in scene consistency and stability during pre- and post-physics simulation, proving its capability to generate simulation-ready environments for embodied AI agents. 2026-03-29T08:15:17Z Minzhang Li Kuixiang Shao Xuebing Li Yuyang Jiao Yinuo Bai Hengan Zhou Sixian Shen Jiayuan Gu Jingyi Yu http://arxiv.org/abs/2509.13688v3 CraftMesh: High-Fidelity Generative Mesh Manipulation via Poisson Seamless Fusion 2026-03-29T05:32:32Z Controllable, high-fidelity mesh editing remains a significant challenge in 3D content creation. Existing generative methods often struggle with complex geometries and fail to produce detailed results. We propose CraftMesh, a novel framework for high-fidelity generative mesh manipulation via Poisson Seamless Fusion. Our key insight is to decompose mesh editing into a pipeline that leverages the strengths of 2D and 3D generative models: we edit a 2D reference image, then generate a region-specific 3D mesh, and seamlessly fuse it into the original model. We introduce two core techniques: Poisson Geometric Fusion, which utilizes a hybrid SDF/Mesh representation with normal blending to achieve harmonious geometric integration, and Poisson Texture Harmonization for visually consistent texture blending. Experimental results demonstrate that CraftMesh outperforms state-of-the-art methods, delivering superior global consistency and local detail in complex editing tasks. 2025-09-17T04:35:48Z James Jincheng Yuxiao Wu Youcheng Cai Ligang Liu http://arxiv.org/abs/2603.02190v2 Sketch2Colab: Sketch-Conditioned Multi-Human Animation via Controllable Flow Distillation 2026-03-29T03:13:10Z We present Sketch2Colab, which turns storyboard-style 2D sketches into coherent, object-aware 3D multi-human motion with fine-grained control over agents, joints, timing, and contacts. Diffusion-based motion generators offer strong realism but often rely on costly guidance for multi-entity control and degrade under strong conditioning. Sketch2Colab instead learns a sketch-conditioned diffusion prior and distills it into a rectified-flow student in latent space for fast, stable sampling. To make motion follow storyboards closely, we guide the student with differentiable objectives that enforce keyframes, paths, contacts, and physical consistency. Collaborative motion naturally involves discrete changes in interaction, such as converging, forming contact, cooperative transport, or disengaging, and a continuous flow alone struggles to sequence these shifts cleanly. We address this with a lightweight continuous-time Markov chain (CTMC) planner that tracks the active interaction regime and modulates the flow to produce clearer, synchronized coordination in human-object-human motion. Experiments on CORE4D and InterHuman show that Sketch2Colab outperforms baselines in constraint adherence and perceptual quality while sampling substantially faster than diffusion-only alternatives. 2026-03-02T18:52:51Z Accepted to CVPR 2026 Main Conference (11 pages, 8 figures) Divyanshu Daiya Aniket Bera http://arxiv.org/abs/2603.07455v2 Image Generation Models: A Technical History 2026-03-29T02:48:56Z Image generation has advanced rapidly over the past decade, yet the literature seems fragmented across different models and application domains. This paper aims to offer a comprehensive survey of breakthrough image generation models, including variational autoencoders (VAEs), generative adversarial networks (GANs), normalizing flows, autoregressive and transformer-based generators, and diffusion-based methods. We provide a detailed technical walkthrough of each model type, including their underlying objectives, architectural building blocks, and algorithmic training steps. For each model type, we present the optimization techniques as well as common failure modes and limitations. We also go over recent developments in video generation and present the research works that made it possible to go from still frames to high quality videos. Lastly, we cover the growing importance of robustness and responsible deployment of these models, including deepfake risks, detection, artifacts, and watermarking. 2026-03-08T04:11:01Z Rouzbeh Shirvani http://arxiv.org/abs/2512.02172v2 SplatSuRe: Selective Super-Resolution for Multi-view Consistent 3D Gaussian Splatting 2026-03-28T19:26:09Z 3D Gaussian Splatting (3DGS) enables high-quality novel view synthesis, motivating interest in generating higher-resolution renders than those available during training. A natural strategy is to apply super-resolution (SR) to low-resolution (LR) input views, but independently enhancing each image introduces multi-view inconsistencies, leading to blurry renders. Prior methods attempt to mitigate these inconsistencies through learned neural components, temporally consistent video priors, or joint optimization on LR and SR views, but all uniformly apply SR across every image. In contrast, our key insight is that close-up LR views may contain high-frequency information for regions also captured in more distant views and that we can use the camera pose relative to scene geometry to inform where to add SR content. Building on this insight, we propose SplatSuRe, a method that selectively applies SR content only in undersampled regions lacking high-frequency supervision, yielding sharper and more consistent results. Across Tanks & Temples, Deep Blending, and Mip-NeRF 360, our approach surpasses baselines in both fidelity and perceptual quality. Notably, our gains are most significant in localized foreground regions where higher detail is desired. 2025-12-01T20:08:39Z Project Page: https://splatsure.github.io/ Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2026 Pranav Asthana Alex Hanson Allen Tu Tom Goldstein Matthias Zwicker Amitabh Varshney http://arxiv.org/abs/2603.27209v1 LightMover: Generative Light Movement with Color and Intensity Controls 2026-03-28T09:51:59Z We present LightMover, a framework for controllable light manipulation in single images that leverages video diffusion priors to produce physically plausible illumination changes without re-rendering the scene. We formulate light editing as a sequence-to-sequence prediction problem in visual token space: given an image and light-control tokens, the model adjusts light position, color, and intensity together with resulting reflections, shadows, and falloff from a single view. This unified treatment of spatial (movement) and appearance (color, intensity) controls improves both manipulation and illumination understanding. We further introduce an adaptive token-pruning mechanism that preserves spatially informative tokens while compactly encoding non-spatial attributes, reducing control sequence length by 41% while maintaining editing fidelity. To train our framework, we construct a scalable rendering pipeline that generates large numbers of image pairs across varied light positions, colors, and intensities while keeping the scene content consistent with the original image. LightMover enables precise, independent control over light position, color, and intensity, and achieves high PSNR and strong semantic consistency (DINO, CLIP) across different tasks. 2026-03-28T09:51:59Z CVPR 2026. 10 pages, 5 figures, 6 tables in main paper; supplementary material included Gengze Zhou Tianyu Wang Soo Ye Kim Zhixin Shu Xin Yu Yannick Hold-Geoffroy Sumit Chaturvedi Qi Wu Zhe Lin Scott Cohen http://arxiv.org/abs/2603.27162v1 BrainRing: An Interactive Web-Based Tool for Brain Connectivity Chord Diagram Visualization 2026-03-28T06:58:21Z Visualizing brain functional connectivity (FC) patterns is essential for understanding neural organization, yet existing tools such as Circos and BrainNet Viewer require complex configuration files or proprietary software environments. We present BrainRing, a free, open-source, browser-based interactive tool for generating publication-quality chord diagrams of brain connectivity data. BrainRing requires no installation, backend server, or programming knowledge. Users simply open a single HTML file in any modern browser. The tool supports 8 widely-used brain atlases (Brainnetome 246, AAL-90/116, Schaefer 100/200/400, Power 264, and Dosenbach 160), provides real-time parameter adjustment through an intuitive graphical interface, and offers comprehensive edge management including click-to-connect, per-edge color customization, and Circos link file import. BrainRing supports both Chinese and English interfaces and enables researchers to produce publication-ready SVG and PNG figures with full control over visual styling, all within seconds rather than the minutes-to-hours workflow typical of script-based approaches. BrainRing is freely available at https://github.com/XiuFan719/brain-connectivity-viz with a live demo at https://XiuFan719.github.io/brain-connectivity-viz/. 2026-03-28T06:58:21Z Xiao Fan Yi Zhang http://arxiv.org/abs/2510.12901v3 SimULi: Real-Time LiDAR and Camera Simulation with Unscented Transforms 2026-03-28T06:05:17Z Rigorous testing of autonomous robots, such as self-driving vehicles, is essential to ensure their safety in real-world deployments. This requires building high-fidelity simulators to test scenarios beyond those that can be safely or exhaustively collected in the real-world. Existing neural rendering methods based on NeRF and 3DGS hold promise but suffer from low rendering speeds or can only render pinhole camera models, hindering their suitability to applications that commonly require high-distortion lenses and LiDAR data. Multi-sensor simulation poses additional challenges as existing methods handle cross-sensor inconsistencies by favoring the quality of one modality at the expense of others. To overcome these limitations, we propose SimULi, the first method capable of rendering arbitrary camera models and LiDAR data in real-time. Our method extends 3DGUT, which natively supports complex camera models, with LiDAR support, via an automated tiling strategy for arbitrary spinning LiDAR models and ray-based culling. To address cross-sensor inconsistencies, we design a factorized 3D Gaussian representation and anchoring strategy that reduces mean camera and depth error by up to 40% compared to existing methods. SimULi renders 10-20x faster than ray tracing approaches and 1.5-10x faster than prior rasterization-based work (and handles a wider range of camera models). When evaluated on two widely benchmarked autonomous driving datasets, SimULi matches or exceeds the fidelity of existing state-of-the-art methods across numerous camera and LiDAR metrics. 2025-10-14T18:22:45Z ICLR 2026 - project page: https://research.nvidia.com/labs/sil/projects/simuli Haithem Turki Qi Wu Xin Kang Janick Martinez Esturo Shengyu Huang Ruilong Li Zan Gojcic Riccardo de Lutio http://arxiv.org/abs/2603.27151v1 DiffSoup: Direct Differentiable Rasterization of Triangle Soup for Extreme Radiance Field Simplification 2026-03-28T06:00:57Z Radiance field reconstruction aims to recover high-quality 3D representations from multi-view RGB images. Recent advances, such as 3D Gaussian splatting, enable real-time rendering with high visual fidelity on sufficiently powerful graphics hardware. However, efficient online transmission and rendering across diverse platforms requires drastic model simplification, reducing the number of primitives by several orders of magnitude. We introduce DiffSoup, a radiance field representation that employs a soup (i.e., a highly unstructured set) of a small number of triangles with neural textures and binary opacity. We show that this binary opacity representation is directly differentiable via stochastic opacity masking, enabling stable training without a mollifier (i.e., smooth rasterization). DiffSoup can be rasterized using standard depth testing, enabling seamless integration into traditional graphics pipelines and interactive rendering on consumer-grade laptops and mobile devices. Code is available at https://github.com/kenji-tojo/diffsoup. 2026-03-28T06:00:57Z Kenji Tojo Bernd Bickel Nobuyuki Umetani http://arxiv.org/abs/2502.03330v3 ControlGUI: Guiding Generative GUI Exploration through Perceptual Visual Flow 2026-03-28T03:13:20Z During the early stages of interface design, designers need to produce multiple sketches to explore a design space. Design tools often fail to support this critical stage, because they insist on specifying more details than necessary. Although recent advances in generative AI have raised hopes of solving this issue, in practice they fail because expressing loose ideas in a prompt is impractical. In this paper, we propose a diffusion-based approach to the low-effort generation of interface sketches. It breaks new ground by allowing flexible control of the generation process via three types of inputs: A) prompts, B) wireframes, and C) visual flows. The designer can provide any combination of these as input at any level of detail, and will get a diverse gallery of low-fidelity solutions in response. The unique benefit is that large design spaces can be explored rapidly with very little effort in input-specification. We present qualitative results for various combinations of input specifications. Additionally, we demonstrate that our model aligns more accurately with these specifications than other models. 2025-02-05T16:25:35Z Aryan Garg Yue Jiang Antti Oulasvirta http://arxiv.org/abs/2506.05398v2 2ndMatch: Finetuning Pruned Diffusion Models via Second-Order Jacobian Matching 2026-03-28T00:58:33Z Diffusion models achieve remarkable performance across diverse generative tasks in computer vision, but their high computational cost remains a major barrier to deployment. Model pruning offers a promising way to reduce inference cost and enable lightweight models. However, pruning leads to quality drop due to reduced capacity. A key limitation of existing pruning approaches is that pruned models are finetuned using the same objective as the dense model (denoising score matching). Since the dense model is accessible during finetuning, it warrants a more effective approach for knowledge transfer from the dense to the pruned model. Motivated by this, we propose \textbf{2ndMatch} (\textbf{2ndM}), a general-purpose finetuning framework that introduces a \textbf{2nd}-order Jacobian ($J^{\top} J$) \textbf{M}atching loss inspired by Finite-Time Lyapunov Exponents. \textbf{2ndM} teaches the pruned model to mimic the sensitivity of the dense teacher, i.e., how to respond to small perturbations over time, through scalable random projections. The framework is architecture-agnostic and applies to both U-Net- and Transformer-based diffusion models. Experiments on CIFAR-10, CelebA, LSUN, ImageNet, and MSCOCO demonstrate that \textbf{2ndM} reduces the performance gap between pruned and dense models, substantially improving output quality. 2025-06-03T20:04:53Z Accepted to CVPR 2026 Caleb Zheng Eli Shlizerman http://arxiv.org/abs/2603.27013v1 PhySkin: Physics-based Bone-driven Neural Garment Simulation 2026-03-27T22:03:38Z Recent advances in digital avatar technology have enabled the generation of compelling virtual characters, but deploying these avatars on compute-constrained devices poses significant challenges for achieving realistic garment deformations. While physics-based simulations yield accurate results, they are computationally prohibitive for real-time applications. Conversely, linear blend skinning offers efficiency but fails to capture the complex dynamics of loose-fitting garments, resulting in unrealistic motion and visual artifacts. Neural methods have shown promise, yet they struggle to animate loose clothing plausibly under strict performance constraints. In this work, we present a novel approach for fast and physically plausible garment draping tailored for resource-constrained environments. Our method leverages a reduced-space quasi-static neural simulation, mapping the garment's full degrees of freedom to a set of bone handles that drive deformation. A neural deformation model is trained in a fully self-supervised manner, eliminating the need for costly simulation data. At runtime, a lightweight neural network modulates the handle deformations based on body shape and pose, enabling realistic garment behavior that respects physical properties such as gravity, fabric stretching, bending, and collision avoidance. Experimental results demonstrate that our method achieves physically plausible garment drapes while generalizing across diverse poses and body shapes, supporting zero-shot evaluation and mesh topology independence. Our method's runtime significantly outperforms past works, as it runs in microseconds per frame using single-threaded CPU inference, offering a practical solution for real-time avatar animation on low-compute devices. 2026-03-27T22:03:38Z Astitva Srivastava Hsiao-yu Chen Ryan Goldade Philipp Herholz Zhongshi Jiang Gene Wei-Chin Lin Lingchen Yang Nikolaos Sarafianos Tuur Stuyck Egor Larionov http://arxiv.org/abs/2506.07917v4 SpeeDe3DGS: Speedy Deformable 3D Gaussian Splatting with Temporal Pruning and Motion Grouping 2026-03-27T20:02:43Z Dynamic extensions of 3D Gaussian Splatting (3DGS) achieve high-quality reconstructions through neural motion fields, but per-Gaussian neural inference makes these models computationally expensive. Building on DeformableGS, we introduce Speedy Deformable 3D Gaussian Splatting (SpeeDe3DGS), which bridges this efficiency-fidelity gap through three complementary modules: Temporal Sensitivity Pruning (TSP) removes low-impact Gaussians via temporally aggregated sensitivity analysis, Temporal Sensitivity Sampling (TSS) perturbs timestamps to suppress floaters and improve temporal coherence, and GroupFlow distills the learned deformation field into shared SE(3) transformations for efficient groupwise motion. On the 50 dynamic scenes in MonoDyGauBench, integrating TSP and TSS into DeformableGS accelerates rendering by 6.78$\times$ on average while maintaining neural-field fidelity and using 10$\times$ fewer primitives. Adding GroupFlow culminates in 13.71$\times$ faster rendering and 2.53$\times$ shorter training, surpassing all baselines in speed while preserving superior image quality. 2025-06-09T16:30:48Z Project Page: https://speede3dgs.github.io/ Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2026 Allen Tu Haiyang Ying Alex Hanson Yonghan Lee Tom Goldstein Matthias Zwicker http://arxiv.org/abs/2603.26926v1 TopoCtrl: Post-Optimization Topology Editing Toward Target Structural Characteristics 2026-03-27T19:04:05Z Topology optimization can generate high-performance structures, but designers often need to revise the resulting topology in ways that reflect fabrication preferences, structural intuition, or downstream design constraints. In particular, they may wish to explicitly control interpretable structural characteristics such as member thickness, characteristic member length, the number of joints, or the number of members connected to a joint. These quantities are often discrete, non-smooth, or only available through a forward evaluation procedure, making them difficult to impose within conventional optimization pipelines. We present TopoCtrl, a post-optimization control framework that repurposes the latent space of a pre-trained topology foundation model for explicit characteristic-guided editing. Given an optimized topology, TopoCtrl encodes it into the latent space of a latent diffusion model, applies partial noising to preserve instance similarity while creating room for modification, and then performs regression-guided denoising toward a prescribed target characteristic. The concept is to train a lightweight regression model on latent representations annotated with evaluated structural characteristics, and to use its gradient as a differentiable guidance signal during reverse diffusion. This avoids the need for characteristic-specific reformulations, hand-derived sensitivities, or iterative optimization. Because the method operates through partial noising of an existing topology latent, it preserves overall structural similarity while still enabling characteristic controls. Across representative control tasks involving both continuous and discrete structural characteristics, TopoCtrl produces target-aligned topology modifications while better preserving structural coherence and design intent than indirect parameter tuning or naive geometric post-processing. 2026-03-27T19:04:05Z Hongrui Chen Dat Quoc Ha Josephine V. Carstensen Faez Ahmed http://arxiv.org/abs/2512.11798v2 Particulate: Feed-Forward 3D Object Articulation 2026-03-27T16:33:22Z We introduce Particulate, a feed-forward model that, given a 3D mesh of an object, infers its articulations, including its 3D parts, their kinematic structure, and the motion constraints. The model is based on a transformer network, the Part Articulation Transformer, which predicts all these parameters for all joints. We train the network end-to-end on a diverse collection of articulated 3D assets from public datasets. During inference, Particulate maps the output of the network back to the input mesh, yielding a fully articulated 3D model in seconds, much faster than prior approaches that require per-object optimization. Particulate also works on AI-generated 3D assets, enabling the generation of articulated 3D objects from a single (real or synthetic) image when combined with an off-the-shelf image-to-3D model. We further introduce a new challenging benchmark for 3D articulation estimation curated from high-quality public 3D assets, and redesign the evaluation protocol to be more consistent with human preferences. Empirically, Particulate significantly outperforms state-of-the-art approaches. 2025-12-12T18:59:51Z CVPR 2026. Project page: https://ruiningli.com/particulate Ruining Li Yuxin Yao Chuanxia Zheng Christian Rupprecht Joan Lasenby Shangzhe Wu Andrea Vedaldi