https://arxiv.org/api/c/p8WVjOgmJvJoFeT+4Xdsu7dxY2026-09-12T21:00:02Z97403015http://arxiv.org/abs/2609.07414v1RelightFormer: Feed-forward Generative Transformer for Multiview Object Relighting2026-09-07T12:25:58ZImage relighting is traditionally tackled via complex inverse rendering pipelines, which suffer from ill-posed optimization, or single-image generative models that ignore crucial multi-view cues necessary for understanding 3D geometry and material interactions. To address these limitations, we introduce a feed-forward generative Transformer for direct single- and multi-view image relighting that entirely bypasses explicit intrinsic property estimation. Adapted from a video foundation model, our architecture features a latent illumination module that dynamically injects target environment maps into spatial features via cross-attention. Furthermore, we employ permutation-invariant positional encodings to symmetrically process unordered multi-view inputs without sequential bias. To train this robust data-driven model, we construct the massive Laval Objaverse Dataset (LOD), comprising 90K objects and 39K unique illuminations. Extensive experiments demonstrate state-of-the-art visual quality, photorealistic relighting quality, and strong zero-shot generalization across single-view, multi-view, and novel-view relighting tasks.2026-09-07T12:25:58ZSIGGRAPH Asia 2026. Hejun and Jinxi are co-first authors. Code and data are available at: https://github.com/vLAR-group/RelightFormerHejun WangJinxi LiJunwei JiangShiwei MaoHu ChengShouwang HuangBo Yanghttp://arxiv.org/abs/2604.08547v2GaussiAnimate: Rig Animatable Categories with Level of Dynamics2026-09-07T09:23:23ZWe propose Skelebones, a Scaffold-Skin Rigging System built on three steps: (1) Bones compress temporally consistent Gaussian or mesh sequences into free-form bones with smooth skinning weights, approximating non-rigid deformations via linear blend skinning (LBS); (2) Skeleton extracts the Mean Curvature Skeleton (MCS) from the canonical shape and temporally refines its topology and kinematics into a compact skeletal structure; and (3) Binding connects the skeleton and bones through non-parametric Partwise Motion Matching (PartMM), which synthesizes novel bone motions by matching, retrieving, and blending existing ones.
Together, these steps compress the dynamics of 4D shapes into compact skelebones that are simultaneously controllable and expressive. The resulting representation is category-agnostic, meaning template-free; motion-adaptive, with dynamic topology; and topology-correct, with a skeleton consistent with the surface geometry. PartMM requires no learning.
We validate our method on both synthetic and real-world datasets, achieving substantial reanimation improvements on unseen poses: a 17.3 dB PSNR gain over LBS on DNA-Rendering and a 45.6 dB gain over Bag-of-Bones on ActorsHQ, while preserving high rendering fidelity for characters with complex non-rigid dynamics. PartMM generalizes robustly to both Gaussian and mesh representations, excelling in low-data regimes of approximately 1,000 frames, with a 48.4 RMSE improvement over LBS and improvements of more than 20 over GRU- and MLP-based methods. Code will be publicly released at https://cookmaker.cn/gaussianimate/.2026-04-09T17:59:59ZAccepted by SIGGRAPH Asia 2026, Page: https://cookmaker.cn/gaussianimateJiaxin WangDongxin LyuZeyu CaiZhiyang DouCheng LinAnpei ChenYuliang Xiuhttp://arxiv.org/abs/2609.06950v1CT2Yarn: Yarn-Level Reconstruction of Crochet from Computed Tomography2026-09-07T02:41:20ZWe introduce CT2Yarn, a human-in-the-loop framework for recovering a single continuous yarn path from micro-computed tomography (micro-CT) scans of real crochet objects. Crochet is a craft that creates complex three-dimensional shapes by interlocking loops formed from a single yarn. Recovering the underlying yarn path from external observations is challenging because of severe self-occlusion. While micro-CT reveals the full internal structure of a crochet object, the volumetric scan alone does not explicitly encode how the yarn traverses the object. This challenge stems from the hierarchical structure of yarn: a yarn consists of multiple twisted plies, and each ply itself consists of twisted fibers. Consequently, local fiber orientations observed in micro-CT scans are not aligned with the overall yarn direction. To recover yarn-level orientations from ply-level fiber orientations, we first estimate local fiber directions using Gabor filtering and convert the volume into an oriented point cloud. We then introduce an anisotropic mean-shift procedure that aggregates local fiber orientations within a neighborhood of the yarn radius into yarn-level orientation estimates. Combined with an automatic topology skeletonization strategy, our method extracts yarn-path fragments. Subsequent fragment linking, junction cleaning, and loop detection merge these trees into a small number of long curves. Where the automatic reconstruction remains ambiguous, a sketch-based user interface enables users to interactively complete the single continuous yarn path. The recovered yarn path enables downstream applications including physically based simulation, ply-level rendering, and stitch-pattern extraction.2026-09-07T02:41:20ZAccepted to the Conference Track of Pacific Graphics 2026Chang LuoNobuyuki Umetanihttp://arxiv.org/abs/2609.06929v1Sub-Pixel Affine Registration of Space Debris Images via the Radon Point Spread Function2026-09-07T01:56:13ZInter-frame affine misalignment caused by platform jitter and attitude adjustments poses a fundamental challenge for multi-frame analysis of point targets in optical surveillance. Conventional registration methods rely on spatial intensity correlations or distinctive image features, both of which are largely absent in low-signal-to-noise-ratio point target imagery. We introduce the Radon Point Spread Function (RPSF) to characterize point targets in the Radon-transformed domain, and derive a closed-form framework that jointly estimates inter-frame translation and rotation from as few as four scalar RPSF samples per frame pair. The method requires no iterative optimization, feature extraction or interpolation, which is suitable for resource-constrained onboard processing. Simulation results confirm sub-pixel translation accuracy and a mean rotation error of 0.2556° at 1° Radon angular resolution. Validation on five real space debris datasets including both ground-based and in-orbit observations yields a mean calibration error below 0.5 pixels, substantially exceeding the precision required for reliable multi-frame processing.2026-09-07T01:56:13ZShenshen LuanMiaomiao TianShuai JiangYan YangShuguo XieZezhou Sunhttp://arxiv.org/abs/2609.06766v1Illustrating Hyperbolic Surfaces with Mesh Embeddings2026-09-06T18:17:48ZHyperbolic geometry exhibits geometric phenomena, such as fast area growth, that are difficult to visualize faithfully in Euclidean space, and which standard models like the Poincaré disk can obscure. To bring hyperbolic geometry to life, we embed hyperbolic surfaces in Euclidean space by discretizing the surfaces into meshes, and minimizing a distortion energy so that the edge lengths in the embeddings match those in the hyperbolic plane. The resulting surfaces buckle and ruffle to accommodate the extra area, making visible what flat models hide. We present exemplary illustrations, such as embedded disks, equidistant strips, diverging geodesics, and also artistic organic-like renders. We discuss our use of these models, as renders and 3D prints, in research talks, public engagement, outreach, and education.2026-09-06T18:17:48Z11 pages, 13 figures. Bridges 2026 Conference ProceedingsProceedings of Bridges 2026: Mathematics, Art, Music, Architecture, Culture, Galway, Ireland, Aug. 5-9, 2026Fabian LanderErik LöffelholtzDiaaeldin TahaSteve TrettelAnna Wienhardhttp://arxiv.org/abs/2609.06723v1ADELE - Adaptive Delaunay Grids for High-Fidelity Mesh-Native Reconstruction2026-09-06T16:55:21ZMeshes remain the most practical representation for geometry reasoning and integration into graphics pipelines, yet existing reconstruction methods struggle to produce high-quality meshes. Most state-of-the-art approaches initially learn an intermediate representation (NeRF/3DGS) and treat mesh extraction as a post-processing step, which often leads to oversmoothed surfaces or poor quality meshes with excessive triangle counts.Existing mesh-native optimization methods alleviate some of these issues but suffer from fixed-resolution discretizations and unstable optimization behavior. In this paper, we introduce an adaptive mesh-based optimization framework and a practical mesh rendering technique to address these challenges. Our representation combines an optimizable Delaunay-triangulated tetrahedral grid with a multi-resolution hash grid. The former is refined through point pruning and insertion, while the latter provides latent features for SDF/appearance value predictions. We use volumetric rendering to bootstrap a coarse geometry while leveraging mesh-based rendering for recovering fine-grained details. Additionally, we propose a differentiable, rasterization-based depth-offset rendering formulation, reducing geometric artifacts and improving reconstruction quality. Our method significantly outperforms existing mesh optimization approaches across a variety of object-centric benchmarks while being competitive with state-of-the-art NeRF/3DGS methods.2026-09-06T16:55:21ZAccepted to SIGGRAPH Asia 2026 Conference Papers | Project page: https://johannes-weidenfeller.github.io/adele | Code: https://github.com/johannes-weidenfeller/adeleJohannes WeidenfellerShaofei WangPhilipp FürnstahlSiyu Tang10.1145/3829340.3842214http://arxiv.org/abs/2605.11696v2WildRelight: A Real-World Benchmark and Physics-Guided Adaptation for Single-Image Relighting2026-09-06T15:35:28ZRecent single-image relighting methods, powered by advanced generative models, have achieved impressive photorealism on synthetic benchmarks. However, their effectiveness in the complex visual landscape of the real world remains largely unverified. A critical gap exists, as current datasets are typically designed for multi-view reconstruction and fail to address the unique challenges of single-image relighting. To bridge this synthetic-to-real gap, we introduce WildRelight, the first in-the-wild dataset specifically created for evaluating single-image relighting models. WildRelight features a diverse collection of high-resolution outdoor scenes, captured under strictly aligned, temporally varying natural illuminations, each paired with a high-dynamic-range environment map. Using this data, we establish a rigorous benchmark revealing that state-of-the-art models trained on synthetic data suffer from severe domain shifts. The strictly aligned temporal structure of WildRelight enables a new paradigm for domain adaptation. We demonstrate this by introducing a physics-guided inference framework that leverages the captured natural light evolution as a self-supervised constraint. By integrating Diffusion Posterior Sampling (DPS) with temporal Sampling-Aware Test-Time Adaptation (TTA), we show that the dataset allows synthetic models to align with real-world statistics on-the-fly, transforming the intractable sim-to-real challenge into a tractable self-supervised task. The dataset and code will be made publicly available to foster robust, physically-grounded relighting research.2026-05-12T07:53:27ZNumerical typo fixes. Project Page: https://lez-s.github.io/wildrelight_proj/Lezhong WangMehmet Onurcan KayaSiavash BigdeliJeppe Revall Frisvadhttp://arxiv.org/abs/2609.06591v1Unifying Physics-Based Humanoid Interaction with a Context-Conditioned Interaction Prior2026-09-06T13:16:48ZDeveloping unified physics-based humanoid controllers that can navigate complex 3D scenes and manipulate objects remains a longstanding challenge. Existing approaches are often specialized for either locomotion or object-centric manipulation, or rely on task-specific reward engineering that does not scale well across diverse behaviors. We present CHIP, a unified, physics-grounded framework for learning reusable humanoid interaction skills from heterogeneous motion data. Central to our approach is a conditional interaction prior that models a context-dependent distribution over these skills within a shared discrete space. Our method is trained in three stages. We first learn physics-based motion-imitation policies that acquire grounded teacher behaviors from heterogeneous interaction data. We then distill these behaviors into a context-conditioned interaction prior that captures reusable motion structure across locomotion and manipulation. Finally, we initialize downstream task policies from the pretrained prior and adapt them through prior-regularized online RL post-training. Experiments on a diverse suite of humanoid interaction tasks show that our approach supports scene-aware locomotion, contact-rich object manipulation, and compositional behaviors such as environment-aware object transport and long-horizon skill sequencing, while producing smooth transitions and physically plausible motion.2026-09-06T13:16:48ZAccepted at SIGGRAPH Asia 2026. Project page: https://jiann-li.github.io/chip-project/Jianan LiXiao ChenTien-Tsin Wonghttp://arxiv.org/abs/2609.06517v1Skinned Motion Retargeting via Artifact-driven Kinematic Prior Refinement2026-09-06T10:18:17ZMotion retargeting aims to transfer a source motion to target characters with different skeletal structures, proportions, and body shapes. Although recent neural retargeting methods have improved flexibility across diverse skeletons, target-side geometric artifacts such as self-penetration remain difficult to resolve. Specifically, existing geometry-aware approaches often rely on fixed skeleton templates or implicit geometry-conditioned prediction, requiring a single network to account for target geometry deformation, detect target-side artifacts, and predict the corresponding correction from target geometry alone, which limits their ability to generalize across diverse skeleton structures and body shapes. In this paper, we present a geometry-aware motion retargeting framework that explicitly connects artifacts observed in the posed character geometry to motion refinement while preserving the flexibility of skeleton-agnostic neural retargeting. Our method first learns a motion embedding shared across different skeletons using a transformer-based retargeting autoencoder that transfers motion across arbitrary source--target skeleton pairs. Building on this kinematic motion prior, we introduce an artifact-driven refinement module that observes self-penetration on the posed target mesh and converts it into a corrective cue through a motion-to-vertex Jacobian. We further condition motion decoding on target geometry using skinning weight-based joint-aligned geometry features derived from the rest pose mesh. This design combines explicit target-side artifact reasoning with flexible geometry-aware decoding in a unified framework. Experiments on both fixed and arbitrary skeleton structure settings show that our method improves kinematic retargeting accuracy and reduces geometric artifacts, producing plausible motions across seen and unseen target characters.2026-09-06T10:18:17ZAccepted SIGGRAPH Asia 2026 (Journal Track); Project page https://seokhyeonhong.github.io/projects/kinematic-refinement/Seokhyeon HongChaelin KimInseo JangSoojin ChoiJunyong Nohhttp://arxiv.org/abs/2609.06436v1PLSR: Progressive and Localized Super-Resolution of 3D Objects via Localized Latent Voxel Diffusion2026-09-06T07:21:36ZHigh-resolution 3D asset generation is vital in various 3D applications. Existing state-of-the-art diffusion-based models remain constrained by fixed resolutions, limiting their ability to produce details. In this paper, we tackle the challenge of generating more detailed, higher-resolution 3D objects by introducing a 3D super-resolution (SR) framework built on existing 3D generative foundation models. To this end, we design PLSR, a progressive and localized super-resolution solution to achieve this goal effectively and memory efficiently. Technically, given a coarse geometry from a pretrained 3D generator, we decompose the global SR task into localized sub-tasks via an associative input decomposition scheme, adapt a flow-based 3D generator into a localized super-resolution model through low-cost finetuning, and unify them in an iterative patch-wise denoising pipeline for seamless high-resolution output. Experiments on challenging objects show that our approach is able to generate 3D details with new strong fine-detail fidelity while significantly reducing the computational cost, offering a new and practical solution for high-resolution 3D asset generation.2026-09-06T07:21:36ZYuxin LiuMinshan XieJiawen LiangRunsong ZhuChi-Wing FuTien-Tsin Wonghttp://arxiv.org/abs/2609.06209v1RBF Your SDF: Radial Basis Function Interpolation of Signed Distance Fields with Implied Tangent Points2026-09-05T18:12:32ZSigned distance fields (SDFs) are a popular implicit representation of geometry. Converting a discrete set of SDF samples into an explicit surface is a fundamental problem in geometry processing. Traditional reconstruction methods such as marching cubes and dual contouring ignore the geometric information carried by samples far from the surface. Recently, Sellán et al.[2023] and several follow-up works leveraged the tangent-sphere structure of SDFs; every sample implies a point on a sphere tangent to the surface. However, these approaches extract the zero-level set via surface reconstruction, which considers only points and normals on the surface and ignores the remaining samples. We propose an approach that marries the tangent-sphere observation with radial basis function interpolation of all data, the implied surface points and the original data. By detecting spheres with extremely constrained tangent points, a configuration geometrically forced at sharp surface features, we identify and preserve surface corners that surface reconstruction-based methods systematically round. A partition-of-unity decomposition allows our method to scale efficiently to large grid resolutions. Our reconstructions improve both Chamfer and Hausdorff accuracy at every tested resolution.2026-09-05T18:12:32Z18 pagesYong ChengYotam Gingoldhttp://arxiv.org/abs/2609.06157v1From Splats to Silicon: Rethinking Computational Efficiency of 3DGS2026-09-05T15:57:26Z3D Gaussian splatting (3DGS) represents scenes with explicit primitives and supports real-time novel-view synthesis, yet its system efficiency varies substantially across scenes, viewpoints, rendering paths, and platform constraints. Existing studies pursue efficiency through representation and algorithm design, GPU runtime optimization, and architectural support, but their reported gains correspond to different points along the rendering and update paths. Connecting these indicators to end-to-end system benefit requires tracing how each optimization changes Gaussian selection, screen-space work, data movement, and stage or frame time. We therefore use a workload-centric framework to connect representation and algorithm research, GPU runtimes, and hardware architectures and to identify recurring workload patterns. We complement literature analysis with reproduced measurements and controlled GPU profiling of selected implementations, relating workload counts to stage time and memory traffic. Together, these comparisons show that system gains depend on workload reductions reaching downstream execution, granularity matching each stage, and the cost of data transfers, synchronization, and cached results, gradients, and optimizer data. Building on these findings, we discuss more consistent evaluation under rendering-quality constraints and identify key directions for future system design.2026-09-05T15:57:26ZMinnan PeiQiwei DongYihan ZhouGang LiYuchen ZhuWenju ZhaoZhongtian LongSiting WangPeisong WangJian Chenghttp://arxiv.org/abs/2606.09606v4PTIR-GS: Path-Traced Inverse Rendering with Global Illumination in 3D Gaussian Fields2026-09-05T15:53:35ZRay tracing enables 3D Gaussian fields to serve as a representation for physically based light transport. Faithful inverse rendering requires forward rendering and backward optimization to be defined within a consistent light-transport pipeline. Existing Gaussian inverse-rendering methods typically rely on splatting-derived G-buffers and screen-space optimization. The local affine approximation of perspective projection can introduce inconsistencies with path tracing, while simplified rendering formulations often neglect or approximate indirect illumination and visibility. Therefore, we propose a splatting-free path-traced inverse-rendering framework for 3D Gaussian fields that unifies forward rendering and backward optimization within the same recursive multi-bounce light-transport pipeline. We formulate light transport over overlapping Gaussian primitives directly in path space, enabling Monte Carlo path tracing and pathwise gradient replay over the same recursively sampled transport paths. The framework jointly optimizes materials and environment light under the full rendering equation, with ray-traced visibility and global illumination explicitly evaluated. Extensive experiments demonstrate competitive material estimation and improved path-traced rendering quality, producing more plausible shadows, reflections, and relighting under global illumination.2026-06-08T15:15:07ZProject page: https://junkzhu.github.io/project_pages/PTIR/Junke ZhuHao ZhangYutian ZhuAng LiChenxiao HuMeng GaiFei ZhuZhangjin HuangSheng Lihttp://arxiv.org/abs/2601.02805v2The Perceptual Cost of Passthrough: How Video See-Through HMDs Degrade Human Visual Perception of Acuity, Contrast, and Color2026-09-05T15:38:26ZVideo see-through (VST) technology aims to seamlessly blend the virtual and physical worlds by reconstructing reality through cameras. However, while manufacturers promise high perceptual fidelity, it remains unclear how closely recent commercial VST systems preserve basic visual functions across environmental conditions. In this work, we present an end-to-end perceptual benchmark for three popular VST headsets: Apple Vision Pro, Meta Quest 3, and Meta Quest Pro. Using adapted psychophysical measures, we evaluated participants' visual acuity, contrast sensitivity, and color vision under both normal and low-light conditions, with naked-eye vision as the reference. Our results show measurable gaps between VST and naked-eye performance, especially for visual acuity and contrast sensitivity in low-light environments. By mapping these perceptual gaps across devices, visual functions, and lighting levels, this work provides a practical benchmark for current commercial VST capabilities and highlights where experience design or device optimization may need to compensate for perceptual loss.2026-01-06T08:28:23Z12 pages, 8 figures, 4 tablesJialin WangSongming PingKemu XuYue LiHai-Ning Lianghttp://arxiv.org/abs/2511.16273v4TetraSDF: Analytic Isosurface Extraction with Multi-resolution Tetrahedral Grid2026-09-05T05:50:46ZExtracting an explicit surface that exactly matches the zero-level set of a neural signed distance function (SDF) remains challenging. Sampling-based isosurfacing methods such as Marching Cubes introduce discretization error. In contrast, continuous piecewise affine (CPWA) analytic approaches typically require plain ReLU MLPs, which limits the ability to learn high-frequency SDFs in practice. We present TetraSDF, an analytic isosurface extraction framework for SDFs that retains the expressiveness of grid-based encoders while enabling exact zero-level set extraction, by representing the SDF with a ReLU MLP composed with a multi-resolution tetrahedral positional encoder. Our positional encoder's barycentric interpolation preserves a global CPWA structure, allowing us to track ReLU linear regions within an encoder-induced polyhedral complex. We further introduce a fixed analytic input preconditioner derived from the encoder's metric to reduce directional bias, thereby stabilizing training. Across multiple benchmarks, TetraSDF matches or surpasses existing grid-based encoders in SDF reconstruction accuracy, while faithfully recovering the network's zero-level set as a triangle mesh.2025-11-20T11:53:52ZSeonghun OhYoungjung UhJin-Hwa Kim