https://arxiv.org/api/c/p8WVjOgmJvJoFeT+4Xdsu7dxY 2026-09-12T21:00:02Z 9740 30 15 http://arxiv.org/abs/2609.07414v1 RelightFormer: Feed-forward Generative Transformer for Multiview Object Relighting 2026-09-07T12:25:58Z Image relighting is traditionally tackled via complex inverse rendering pipelines, which suffer from ill-posed optimization, or single-image generative models that ignore crucial multi-view cues necessary for understanding 3D geometry and material interactions. To address these limitations, we introduce a feed-forward generative Transformer for direct single- and multi-view image relighting that entirely bypasses explicit intrinsic property estimation. Adapted from a video foundation model, our architecture features a latent illumination module that dynamically injects target environment maps into spatial features via cross-attention. Furthermore, we employ permutation-invariant positional encodings to symmetrically process unordered multi-view inputs without sequential bias. To train this robust data-driven model, we construct the massive Laval Objaverse Dataset (LOD), comprising 90K objects and 39K unique illuminations. Extensive experiments demonstrate state-of-the-art visual quality, photorealistic relighting quality, and strong zero-shot generalization across single-view, multi-view, and novel-view relighting tasks. 2026-09-07T12:25:58Z SIGGRAPH Asia 2026. Hejun and Jinxi are co-first authors. Code and data are available at: https://github.com/vLAR-group/RelightFormer Hejun Wang Jinxi Li Junwei Jiang Shiwei Mao Hu Cheng Shouwang Huang Bo Yang http://arxiv.org/abs/2604.08547v2 GaussiAnimate: Rig Animatable Categories with Level of Dynamics 2026-09-07T09:23:23Z We propose Skelebones, a Scaffold-Skin Rigging System built on three steps: (1) Bones compress temporally consistent Gaussian or mesh sequences into free-form bones with smooth skinning weights, approximating non-rigid deformations via linear blend skinning (LBS); (2) Skeleton extracts the Mean Curvature Skeleton (MCS) from the canonical shape and temporally refines its topology and kinematics into a compact skeletal structure; and (3) Binding connects the skeleton and bones through non-parametric Partwise Motion Matching (PartMM), which synthesizes novel bone motions by matching, retrieving, and blending existing ones. Together, these steps compress the dynamics of 4D shapes into compact skelebones that are simultaneously controllable and expressive. The resulting representation is category-agnostic, meaning template-free; motion-adaptive, with dynamic topology; and topology-correct, with a skeleton consistent with the surface geometry. PartMM requires no learning. We validate our method on both synthetic and real-world datasets, achieving substantial reanimation improvements on unseen poses: a 17.3 dB PSNR gain over LBS on DNA-Rendering and a 45.6 dB gain over Bag-of-Bones on ActorsHQ, while preserving high rendering fidelity for characters with complex non-rigid dynamics. PartMM generalizes robustly to both Gaussian and mesh representations, excelling in low-data regimes of approximately 1,000 frames, with a 48.4 RMSE improvement over LBS and improvements of more than 20 over GRU- and MLP-based methods. Code will be publicly released at https://cookmaker.cn/gaussianimate/. 2026-04-09T17:59:59Z Accepted by SIGGRAPH Asia 2026, Page: https://cookmaker.cn/gaussianimate Jiaxin Wang Dongxin Lyu Zeyu Cai Zhiyang Dou Cheng Lin Anpei Chen Yuliang Xiu http://arxiv.org/abs/2609.06950v1 CT2Yarn: Yarn-Level Reconstruction of Crochet from Computed Tomography 2026-09-07T02:41:20Z We introduce CT2Yarn, a human-in-the-loop framework for recovering a single continuous yarn path from micro-computed tomography (micro-CT) scans of real crochet objects. Crochet is a craft that creates complex three-dimensional shapes by interlocking loops formed from a single yarn. Recovering the underlying yarn path from external observations is challenging because of severe self-occlusion. While micro-CT reveals the full internal structure of a crochet object, the volumetric scan alone does not explicitly encode how the yarn traverses the object. This challenge stems from the hierarchical structure of yarn: a yarn consists of multiple twisted plies, and each ply itself consists of twisted fibers. Consequently, local fiber orientations observed in micro-CT scans are not aligned with the overall yarn direction. To recover yarn-level orientations from ply-level fiber orientations, we first estimate local fiber directions using Gabor filtering and convert the volume into an oriented point cloud. We then introduce an anisotropic mean-shift procedure that aggregates local fiber orientations within a neighborhood of the yarn radius into yarn-level orientation estimates. Combined with an automatic topology skeletonization strategy, our method extracts yarn-path fragments. Subsequent fragment linking, junction cleaning, and loop detection merge these trees into a small number of long curves. Where the automatic reconstruction remains ambiguous, a sketch-based user interface enables users to interactively complete the single continuous yarn path. The recovered yarn path enables downstream applications including physically based simulation, ply-level rendering, and stitch-pattern extraction. 2026-09-07T02:41:20Z Accepted to the Conference Track of Pacific Graphics 2026 Chang Luo Nobuyuki Umetani http://arxiv.org/abs/2609.06929v1 Sub-Pixel Affine Registration of Space Debris Images via the Radon Point Spread Function 2026-09-07T01:56:13Z Inter-frame affine misalignment caused by platform jitter and attitude adjustments poses a fundamental challenge for multi-frame analysis of point targets in optical surveillance. Conventional registration methods rely on spatial intensity correlations or distinctive image features, both of which are largely absent in low-signal-to-noise-ratio point target imagery. We introduce the Radon Point Spread Function (RPSF) to characterize point targets in the Radon-transformed domain, and derive a closed-form framework that jointly estimates inter-frame translation and rotation from as few as four scalar RPSF samples per frame pair. The method requires no iterative optimization, feature extraction or interpolation, which is suitable for resource-constrained onboard processing. Simulation results confirm sub-pixel translation accuracy and a mean rotation error of 0.2556° at 1° Radon angular resolution. Validation on five real space debris datasets including both ground-based and in-orbit observations yields a mean calibration error below 0.5 pixels, substantially exceeding the precision required for reliable multi-frame processing. 2026-09-07T01:56:13Z Shenshen Luan Miaomiao Tian Shuai Jiang Yan Yang Shuguo Xie Zezhou Sun http://arxiv.org/abs/2609.06766v1 Illustrating Hyperbolic Surfaces with Mesh Embeddings 2026-09-06T18:17:48Z Hyperbolic geometry exhibits geometric phenomena, such as fast area growth, that are difficult to visualize faithfully in Euclidean space, and which standard models like the Poincaré disk can obscure. To bring hyperbolic geometry to life, we embed hyperbolic surfaces in Euclidean space by discretizing the surfaces into meshes, and minimizing a distortion energy so that the edge lengths in the embeddings match those in the hyperbolic plane. The resulting surfaces buckle and ruffle to accommodate the extra area, making visible what flat models hide. We present exemplary illustrations, such as embedded disks, equidistant strips, diverging geodesics, and also artistic organic-like renders. We discuss our use of these models, as renders and 3D prints, in research talks, public engagement, outreach, and education. 2026-09-06T18:17:48Z 11 pages, 13 figures. Bridges 2026 Conference Proceedings Proceedings of Bridges 2026: Mathematics, Art, Music, Architecture, Culture, Galway, Ireland, Aug. 5-9, 2026 Fabian Lander Erik Löffelholtz Diaaeldin Taha Steve Trettel Anna Wienhard http://arxiv.org/abs/2609.06723v1 ADELE - Adaptive Delaunay Grids for High-Fidelity Mesh-Native Reconstruction 2026-09-06T16:55:21Z Meshes remain the most practical representation for geometry reasoning and integration into graphics pipelines, yet existing reconstruction methods struggle to produce high-quality meshes. Most state-of-the-art approaches initially learn an intermediate representation (NeRF/3DGS) and treat mesh extraction as a post-processing step, which often leads to oversmoothed surfaces or poor quality meshes with excessive triangle counts.Existing mesh-native optimization methods alleviate some of these issues but suffer from fixed-resolution discretizations and unstable optimization behavior. In this paper, we introduce an adaptive mesh-based optimization framework and a practical mesh rendering technique to address these challenges. Our representation combines an optimizable Delaunay-triangulated tetrahedral grid with a multi-resolution hash grid. The former is refined through point pruning and insertion, while the latter provides latent features for SDF/appearance value predictions. We use volumetric rendering to bootstrap a coarse geometry while leveraging mesh-based rendering for recovering fine-grained details. Additionally, we propose a differentiable, rasterization-based depth-offset rendering formulation, reducing geometric artifacts and improving reconstruction quality. Our method significantly outperforms existing mesh optimization approaches across a variety of object-centric benchmarks while being competitive with state-of-the-art NeRF/3DGS methods. 2026-09-06T16:55:21Z Accepted to SIGGRAPH Asia 2026 Conference Papers | Project page: https://johannes-weidenfeller.github.io/adele | Code: https://github.com/johannes-weidenfeller/adele Johannes Weidenfeller Shaofei Wang Philipp Fürnstahl Siyu Tang 10.1145/3829340.3842214 http://arxiv.org/abs/2605.11696v2 WildRelight: A Real-World Benchmark and Physics-Guided Adaptation for Single-Image Relighting 2026-09-06T15:35:28Z Recent single-image relighting methods, powered by advanced generative models, have achieved impressive photorealism on synthetic benchmarks. However, their effectiveness in the complex visual landscape of the real world remains largely unverified. A critical gap exists, as current datasets are typically designed for multi-view reconstruction and fail to address the unique challenges of single-image relighting. To bridge this synthetic-to-real gap, we introduce WildRelight, the first in-the-wild dataset specifically created for evaluating single-image relighting models. WildRelight features a diverse collection of high-resolution outdoor scenes, captured under strictly aligned, temporally varying natural illuminations, each paired with a high-dynamic-range environment map. Using this data, we establish a rigorous benchmark revealing that state-of-the-art models trained on synthetic data suffer from severe domain shifts. The strictly aligned temporal structure of WildRelight enables a new paradigm for domain adaptation. We demonstrate this by introducing a physics-guided inference framework that leverages the captured natural light evolution as a self-supervised constraint. By integrating Diffusion Posterior Sampling (DPS) with temporal Sampling-Aware Test-Time Adaptation (TTA), we show that the dataset allows synthetic models to align with real-world statistics on-the-fly, transforming the intractable sim-to-real challenge into a tractable self-supervised task. The dataset and code will be made publicly available to foster robust, physically-grounded relighting research. 2026-05-12T07:53:27Z Numerical typo fixes. Project Page: https://lez-s.github.io/wildrelight_proj/ Lezhong Wang Mehmet Onurcan Kaya Siavash Bigdeli Jeppe Revall Frisvad http://arxiv.org/abs/2609.06591v1 Unifying Physics-Based Humanoid Interaction with a Context-Conditioned Interaction Prior 2026-09-06T13:16:48Z Developing unified physics-based humanoid controllers that can navigate complex 3D scenes and manipulate objects remains a longstanding challenge. Existing approaches are often specialized for either locomotion or object-centric manipulation, or rely on task-specific reward engineering that does not scale well across diverse behaviors. We present CHIP, a unified, physics-grounded framework for learning reusable humanoid interaction skills from heterogeneous motion data. Central to our approach is a conditional interaction prior that models a context-dependent distribution over these skills within a shared discrete space. Our method is trained in three stages. We first learn physics-based motion-imitation policies that acquire grounded teacher behaviors from heterogeneous interaction data. We then distill these behaviors into a context-conditioned interaction prior that captures reusable motion structure across locomotion and manipulation. Finally, we initialize downstream task policies from the pretrained prior and adapt them through prior-regularized online RL post-training. Experiments on a diverse suite of humanoid interaction tasks show that our approach supports scene-aware locomotion, contact-rich object manipulation, and compositional behaviors such as environment-aware object transport and long-horizon skill sequencing, while producing smooth transitions and physically plausible motion. 2026-09-06T13:16:48Z Accepted at SIGGRAPH Asia 2026. Project page: https://jiann-li.github.io/chip-project/ Jianan Li Xiao Chen Tien-Tsin Wong http://arxiv.org/abs/2609.06517v1 Skinned Motion Retargeting via Artifact-driven Kinematic Prior Refinement 2026-09-06T10:18:17Z Motion retargeting aims to transfer a source motion to target characters with different skeletal structures, proportions, and body shapes. Although recent neural retargeting methods have improved flexibility across diverse skeletons, target-side geometric artifacts such as self-penetration remain difficult to resolve. Specifically, existing geometry-aware approaches often rely on fixed skeleton templates or implicit geometry-conditioned prediction, requiring a single network to account for target geometry deformation, detect target-side artifacts, and predict the corresponding correction from target geometry alone, which limits their ability to generalize across diverse skeleton structures and body shapes. In this paper, we present a geometry-aware motion retargeting framework that explicitly connects artifacts observed in the posed character geometry to motion refinement while preserving the flexibility of skeleton-agnostic neural retargeting. Our method first learns a motion embedding shared across different skeletons using a transformer-based retargeting autoencoder that transfers motion across arbitrary source--target skeleton pairs. Building on this kinematic motion prior, we introduce an artifact-driven refinement module that observes self-penetration on the posed target mesh and converts it into a corrective cue through a motion-to-vertex Jacobian. We further condition motion decoding on target geometry using skinning weight-based joint-aligned geometry features derived from the rest pose mesh. This design combines explicit target-side artifact reasoning with flexible geometry-aware decoding in a unified framework. Experiments on both fixed and arbitrary skeleton structure settings show that our method improves kinematic retargeting accuracy and reduces geometric artifacts, producing plausible motions across seen and unseen target characters. 2026-09-06T10:18:17Z Accepted SIGGRAPH Asia 2026 (Journal Track); Project page https://seokhyeonhong.github.io/projects/kinematic-refinement/ Seokhyeon Hong Chaelin Kim Inseo Jang Soojin Choi Junyong Noh http://arxiv.org/abs/2609.06436v1 PLSR: Progressive and Localized Super-Resolution of 3D Objects via Localized Latent Voxel Diffusion 2026-09-06T07:21:36Z High-resolution 3D asset generation is vital in various 3D applications. Existing state-of-the-art diffusion-based models remain constrained by fixed resolutions, limiting their ability to produce details. In this paper, we tackle the challenge of generating more detailed, higher-resolution 3D objects by introducing a 3D super-resolution (SR) framework built on existing 3D generative foundation models. To this end, we design PLSR, a progressive and localized super-resolution solution to achieve this goal effectively and memory efficiently. Technically, given a coarse geometry from a pretrained 3D generator, we decompose the global SR task into localized sub-tasks via an associative input decomposition scheme, adapt a flow-based 3D generator into a localized super-resolution model through low-cost finetuning, and unify them in an iterative patch-wise denoising pipeline for seamless high-resolution output. Experiments on challenging objects show that our approach is able to generate 3D details with new strong fine-detail fidelity while significantly reducing the computational cost, offering a new and practical solution for high-resolution 3D asset generation. 2026-09-06T07:21:36Z Yuxin Liu Minshan Xie Jiawen Liang Runsong Zhu Chi-Wing Fu Tien-Tsin Wong http://arxiv.org/abs/2609.06209v1 RBF Your SDF: Radial Basis Function Interpolation of Signed Distance Fields with Implied Tangent Points 2026-09-05T18:12:32Z Signed distance fields (SDFs) are a popular implicit representation of geometry. Converting a discrete set of SDF samples into an explicit surface is a fundamental problem in geometry processing. Traditional reconstruction methods such as marching cubes and dual contouring ignore the geometric information carried by samples far from the surface. Recently, Sellán et al.[2023] and several follow-up works leveraged the tangent-sphere structure of SDFs; every sample implies a point on a sphere tangent to the surface. However, these approaches extract the zero-level set via surface reconstruction, which considers only points and normals on the surface and ignores the remaining samples. We propose an approach that marries the tangent-sphere observation with radial basis function interpolation of all data, the implied surface points and the original data. By detecting spheres with extremely constrained tangent points, a configuration geometrically forced at sharp surface features, we identify and preserve surface corners that surface reconstruction-based methods systematically round. A partition-of-unity decomposition allows our method to scale efficiently to large grid resolutions. Our reconstructions improve both Chamfer and Hausdorff accuracy at every tested resolution. 2026-09-05T18:12:32Z 18 pages Yong Cheng Yotam Gingold http://arxiv.org/abs/2609.06157v1 From Splats to Silicon: Rethinking Computational Efficiency of 3DGS 2026-09-05T15:57:26Z 3D Gaussian splatting (3DGS) represents scenes with explicit primitives and supports real-time novel-view synthesis, yet its system efficiency varies substantially across scenes, viewpoints, rendering paths, and platform constraints. Existing studies pursue efficiency through representation and algorithm design, GPU runtime optimization, and architectural support, but their reported gains correspond to different points along the rendering and update paths. Connecting these indicators to end-to-end system benefit requires tracing how each optimization changes Gaussian selection, screen-space work, data movement, and stage or frame time. We therefore use a workload-centric framework to connect representation and algorithm research, GPU runtimes, and hardware architectures and to identify recurring workload patterns. We complement literature analysis with reproduced measurements and controlled GPU profiling of selected implementations, relating workload counts to stage time and memory traffic. Together, these comparisons show that system gains depend on workload reductions reaching downstream execution, granularity matching each stage, and the cost of data transfers, synchronization, and cached results, gradients, and optimizer data. Building on these findings, we discuss more consistent evaluation under rendering-quality constraints and identify key directions for future system design. 2026-09-05T15:57:26Z Minnan Pei Qiwei Dong Yihan Zhou Gang Li Yuchen Zhu Wenju Zhao Zhongtian Long Siting Wang Peisong Wang Jian Cheng http://arxiv.org/abs/2606.09606v4 PTIR-GS: Path-Traced Inverse Rendering with Global Illumination in 3D Gaussian Fields 2026-09-05T15:53:35Z Ray tracing enables 3D Gaussian fields to serve as a representation for physically based light transport. Faithful inverse rendering requires forward rendering and backward optimization to be defined within a consistent light-transport pipeline. Existing Gaussian inverse-rendering methods typically rely on splatting-derived G-buffers and screen-space optimization. The local affine approximation of perspective projection can introduce inconsistencies with path tracing, while simplified rendering formulations often neglect or approximate indirect illumination and visibility. Therefore, we propose a splatting-free path-traced inverse-rendering framework for 3D Gaussian fields that unifies forward rendering and backward optimization within the same recursive multi-bounce light-transport pipeline. We formulate light transport over overlapping Gaussian primitives directly in path space, enabling Monte Carlo path tracing and pathwise gradient replay over the same recursively sampled transport paths. The framework jointly optimizes materials and environment light under the full rendering equation, with ray-traced visibility and global illumination explicitly evaluated. Extensive experiments demonstrate competitive material estimation and improved path-traced rendering quality, producing more plausible shadows, reflections, and relighting under global illumination. 2026-06-08T15:15:07Z Project page: https://junkzhu.github.io/project_pages/PTIR/ Junke Zhu Hao Zhang Yutian Zhu Ang Li Chenxiao Hu Meng Gai Fei Zhu Zhangjin Huang Sheng Li http://arxiv.org/abs/2601.02805v2 The Perceptual Cost of Passthrough: How Video See-Through HMDs Degrade Human Visual Perception of Acuity, Contrast, and Color 2026-09-05T15:38:26Z Video see-through (VST) technology aims to seamlessly blend the virtual and physical worlds by reconstructing reality through cameras. However, while manufacturers promise high perceptual fidelity, it remains unclear how closely recent commercial VST systems preserve basic visual functions across environmental conditions. In this work, we present an end-to-end perceptual benchmark for three popular VST headsets: Apple Vision Pro, Meta Quest 3, and Meta Quest Pro. Using adapted psychophysical measures, we evaluated participants' visual acuity, contrast sensitivity, and color vision under both normal and low-light conditions, with naked-eye vision as the reference. Our results show measurable gaps between VST and naked-eye performance, especially for visual acuity and contrast sensitivity in low-light environments. By mapping these perceptual gaps across devices, visual functions, and lighting levels, this work provides a practical benchmark for current commercial VST capabilities and highlights where experience design or device optimization may need to compensate for perceptual loss. 2026-01-06T08:28:23Z 12 pages, 8 figures, 4 tables Jialin Wang Songming Ping Kemu Xu Yue Li Hai-Ning Liang http://arxiv.org/abs/2511.16273v4 TetraSDF: Analytic Isosurface Extraction with Multi-resolution Tetrahedral Grid 2026-09-05T05:50:46Z Extracting an explicit surface that exactly matches the zero-level set of a neural signed distance function (SDF) remains challenging. Sampling-based isosurfacing methods such as Marching Cubes introduce discretization error. In contrast, continuous piecewise affine (CPWA) analytic approaches typically require plain ReLU MLPs, which limits the ability to learn high-frequency SDFs in practice. We present TetraSDF, an analytic isosurface extraction framework for SDFs that retains the expressiveness of grid-based encoders while enabling exact zero-level set extraction, by representing the SDF with a ReLU MLP composed with a multi-resolution tetrahedral positional encoder. Our positional encoder's barycentric interpolation preserves a global CPWA structure, allowing us to track ReLU linear regions within an encoder-induced polyhedral complex. We further introduce a fixed analytic input preconditioner derived from the encoder's metric to reduce directional bias, thereby stabilizing training. Across multiple benchmarks, TetraSDF matches or surpasses existing grid-based encoders in SDF reconstruction accuracy, while faithfully recovering the network's zero-level set as a triangle mesh. 2025-11-20T11:53:52Z Seonghun Oh Youngjung Uh Jin-Hwa Kim