https://arxiv.org/api/9d9wDaZnxgy9Gj7xIS9QTPSBvA4 2026-10-02T23:59:50Z 9866 2310 15 http://arxiv.org/abs/2507.23454v2 Breaking the mould of Social Mixed Reality - State-of-the-Art and Glossary 2025-08-01T06:05:12Z This article explores a critical gap in Mixed Reality (MR) technology: while advances have been made, MR still struggles to authentically replicate human embodiment and socio-motor interaction. For MR to enable truly meaningful social experiences, it needs to incorporate multi-modal data streams and multi-agent interaction capabilities. To address this challenge, we present a comprehensive glossary covering key topics such as Virtual Characters and Autonomisation, Responsible AI, Ethics by Design, and the Scientific Challenges of Social MR within Neuroscience, Embodiment, and Technology. Our aim is to drive the transformative evolution of MR technologies that prioritize human-centric innovation, fostering richer digital connections. We advocate for MR systems that enhance social interaction and collaboration between humans and virtual autonomous agents, ensuring inclusivity, ethical design and psychological safety in the process. 2025-07-31T11:31:12Z pre-print Marta Bieńkiewicz Julia Ayache Panayiotis Charalambous Cristina Becchio Marco Corragio Bertram Taetz Francesco De Lellis Antonio Grotta Anna Server Daniel Rammer Richard Kulpa Franck Multon Azucena Garcia-Palacios Jessica Sutherland Kathleen Bryson Stéphane Donikian Didier Stricker Benoît Bardy http://arxiv.org/abs/2412.17812v2 FaceLift: Learning Generalizable Single Image 3D Face Reconstruction from Synthetic Heads 2025-08-01T03:09:31Z We present FaceLift, a novel feed-forward approach for generalizable high-quality 360-degree 3D head reconstruction from a single image. Our pipeline first employs a multi-view latent diffusion model to generate consistent side and back views from a single facial input, which then feeds into a transformer-based reconstructor that produces a comprehensive 3D Gaussian splats representation. Previous methods for monocular 3D face reconstruction often lack full view coverage or view consistency due to insufficient multi-view supervision. We address this by creating a high-quality synthetic head dataset that enables consistent supervision across viewpoints. To bridge the domain gap between synthetic training data and real-world images, we propose a simple yet effective technique that ensures the view generation process maintains fidelity to the input by learning to reconstruct the input image alongside the view generation. Despite being trained exclusively on synthetic data, our method demonstrates remarkable generalization to real-world images. Through extensive qualitative and quantitative evaluations, we show that FaceLift outperforms state-of-the-art 3D face reconstruction methods on identity preservation, detail recovery, and rendering quality. 2024-12-23T18:59:49Z ICCV 2025 Camera-Ready Version. Project Page: https://weijielyu.github.io/FaceLift Weijie Lyu Yi Zhou Ming-Hsuan Yang Zhixin Shu http://arxiv.org/abs/2508.00950v1 Investigating Crossing Perception in 3D Graph Visualisation 2025-07-31T21:57:07Z Human perception of graph drawings is influenced by a variety of impact factors for which quality measures are used as a proxy indicator. The investigation of those impact factors and their effects is important to evaluate and improve quality measures and drawing algorithms. The number of edge crossings in a 2D graph drawing has long been a main quality measure for drawing evaluation. The use of stereoscopic 3D graph visualisations has gained attraction over the last years, and results from several studies indicate that they can improve analysis efficiency for a range of analysis scenarios. While edge crossings can also occur in 3D, there are edge configurations in space that are not crossings but might be perceived as such from a specific viewpoint. Such configurations create crossings when projected on the corresponding 2D image plane and could impact readability similar to 2D crossings. In 3D drawings, the additional depth aspect and the subsequent impact factors of edge distance and relative edge direction in space might further influence the importance of those configurations for readability. We investigate the impact of such factors in an empirical study and report on findings of difference between major factor categories. 2025-07-31T21:57:07Z Ying Zhang Niklas Groene Karsten Klein Giuseppe Liotta Falk Schreiber http://arxiv.org/abs/2401.13639v2 Winding Clearness for Differentiable Point Cloud Optimization 2025-07-31T16:22:50Z We propose to explore the properties of raw point clouds through the \emph{winding clearness}, a concept we first introduce for measuring the clarity of the interior/exterior relationships represented by the winding number field of the point cloud. In geometric modeling, the winding number is a powerful tool for distinguishing the interior and exterior of a given surface $\partial Ω$, and it has been previously used for point normal orientation and surface reconstruction. In this work, we introduce a novel approach to evaluate and optimize the quality of point clouds based on the winding clearness. We observe that point clouds with less noise generally exhibit better winding clearness. Accordingly, we propose an objective function that quantifies the error in winding clearness, solely utilizing the coordinates of the point clouds. Moreover, we demonstrate that the winding clearness error is differentiable and can serve as a loss function in point cloud processing. We present this observation from two aspects: 1) We update the coordinates of the points by back-propagating the loss function for individual point clouds, resulting in an overall improvement without involving a neural network. 2) We incorporate winding clearness as a geometric constraint in the diffusion-based 3D generative model and update the network parameters to generate point clouds with less noise. Experimental results demonstrate the effectiveness of optimizing the winding clearness in enhancing the point cloud quality. Notably, our method exhibits superior performance in handling noisy point clouds with thin structures, highlighting the benefits of the global perspective enabled by the winding number. 2024-01-24T18:09:16Z Accepted by Computer-Aided Design through SPM 2025 Dong Xiao Yueji Ma Zuoqiang Shi Shiqing Xin Wenping Wang Bailin Deng Bin Wang 10.1016/j.cad.2025.103930 http://arxiv.org/abs/2507.01631v2 Tile and Slide : A New Framework for Scaling NeRF from Local to Global 3D Earth Observation 2025-07-31T13:32:03Z Neural Radiance Fields (NeRF) have recently emerged as a paradigm for 3D reconstruction from multiview satellite imagery. However, state-of-the-art NeRF methods are typically constrained to small scenes due to the memory footprint during training, which we study in this paper. Previous work on large-scale NeRFs palliate this by dividing the scene into NeRFs. This paper introduces Snake-NeRF, a framework that scales to large scenes. Our out-of-core method eliminates the need to load all images and networks simultaneously, and operates on a single device. We achieve this by dividing the region of interest into NeRFs that 3D tile without overlap. Importantly, we crop the images with overlap to ensure each NeRFs is trained with all the necessary pixels. We introduce a novel $2\times 2$ 3D tile progression strategy and segmented sampler, which together prevent 3D reconstruction errors along the tile edges. Our experiments conclude that large satellite images can effectively be processed with linear time complexity, on a single GPU, and without compromise in quality. 2025-07-02T11:59:36Z Accepted at ICCV 2025 Workshop 3D-VAST (From street to space: 3D Vision Across Altitudes). Our code will be made public after the conference at https://github.com/Ellimac0/Snake-NeRF Camille Billouard Dawa Derksen Alexandre Constantin Bruno Vallet http://arxiv.org/abs/2507.15230v3 GALE: Leveraging Heterogeneous Systems for Efficient Unstructured Mesh Data Analysis 2025-07-30T20:08:13Z Unstructured meshes present challenges in scientific data analysis due to irregular distribution and complex connectivity. Computing and storing connectivity information is a major bottleneck for visualization algorithms, affecting both time and memory performance. Recent task-parallel data structures address this by precomputing connectivity information at runtime while the analysis algorithm executes, effectively hiding computation costs and improving performance. However, existing approaches are CPU-bound, forcing the data structure and analysis algorithm to compete for the same computational resources, limiting potential speedups. To overcome this limitation, we introduce a novel task-parallel approach optimized for heterogeneous CPU-GPU systems. Specifically, we offload the computation of mesh connectivity information to GPU threads, enabling CPU threads to focus on executing the visualization algorithm. Following this paradigm, we propose GALE (GPU-Aided Localized data structurE), the first open-source CUDA-based data structure designed for heterogeneous task parallelism. Experiments on two 20-core CPUs and an NVIDIA V100 GPU show that GALE achieves up to 2.7x speedup over state-of-the-art localized data structures while maintaining memory efficiency. 2025-07-21T04:20:12Z Accepted at IEEE VIS 2025 Guoxi Liu Thomas Randall Rong Ge Federico Iuricich http://arxiv.org/abs/2507.23002v1 Noise-Coded Illumination for Forensic and Photometric Video Analysis 2025-07-30T18:08:34Z The proliferation of advanced tools for manipulating video has led to an arms race, pitting those who wish to sow disinformation against those who want to detect and expose it. Unfortunately, time favors the ill-intentioned in this race, with fake videos growing increasingly difficult to distinguish from real ones. At the root of this trend is a fundamental advantage held by those manipulating media: equal access to a distribution of what we consider authentic (i.e., "natural") video. In this paper, we show how coding very subtle, noise-like modulations into the illumination of a scene can help combat this advantage by creating an information asymmetry that favors verification. Our approach effectively adds a temporal watermark to any video recorded under coded illumination. However, rather than encoding a specific message, this watermark encodes an image of the unmanipulated scene as it would appear lit only by the coded illumination. We show that even when an adversary knows that our technique is being used, creating a plausible coded fake video amounts to solving a second, more difficult version of the original adversarial content creation problem at an information disadvantage. This is a promising avenue for protecting high-stakes settings like public events and interviews, where the content on display is a likely target for manipulation, and while the illumination can be controlled, the cameras capturing video cannot. 2025-07-30T18:08:34Z ACM Transactions on Graphics (2025), presented at SIGGRAPH 2025 ACM Trans. Graph. 44, 5, Article 165 (October 2025), 16 pages Peter F. Michael Zekun Hao Serge Belongie Abe Davis 10.1145/3742892 http://arxiv.org/abs/2402.00186v3 Distance and Collision Probability Estimation from Gaussian Surface Models 2025-07-30T17:10:59Z This paper describes continuous-space methodologies to estimate the collision probability, Euclidean distance and gradient between an ellipsoidal robot model and an environment surface modeled as a set of Gaussian distributions. Continuous-space collision probability estimation is critical for uncertainty-aware motion planning. Most collision detection and avoidance approaches assume the robot is modeled as a sphere, but ellipsoidal representations provide tighter approximations and enable navigation in cluttered and narrow spaces. State-of-the-art methods derive the Euclidean distance and gradient by processing raw point clouds, which is computationally expensive for large workspaces. Recent advances in Gaussian surface modeling (e.g. mixture models, splatting) enable compressed and high-fidelity surface representations. Few methods exist to estimate continuous-space occupancy from such models. They require Gaussians to model free space and are unable to estimate the collision probability, Euclidean distance and gradient for an ellipsoidal robot. The proposed methods bridge this gap by extending prior work in ellipsoid-to-ellipsoid Euclidean distance and collision probability estimation to Gaussian surface models. A geometric blending approach is also proposed to improve collision probability estimation. The approaches are evaluated with numerical 2D and 3D experiments using real-world point cloud data. Methods for efficient calculation of these quantities are demonstrated to execute within a few microseconds per ellipsoid pair using a single-thread on low-power CPUs of modern embedded computers 2024-01-31T21:28:40Z Accepted at IROS 2025 Kshitij Goel Wennie Tabib http://arxiv.org/abs/2503.20999v2 Text-Driven Voice Conversion via Latent State-Space Modeling 2025-07-30T17:02:18Z Text-driven voice conversion allows customization of speaker characteristics and prosodic elements using textual descriptions. However, most existing methods rely heavily on direct text-to-speech training, limiting their flexibility in controlling nuanced style elements or timbral features. In this paper, we propose a novel \textbf{Latent State-Space} approach for text-driven voice conversion (\textbf{LSS-VC}). Our method treats each utterance as an evolving dynamical system in a continuous latent space. Drawing inspiration from mamba, which introduced a state-space model for efficient text-driven \emph{image} style transfer, we adapt a loosely related methodology for \emph{voice} style transformation. Specifically, we learn a voice latent manifold where style and content can be manipulated independently by textual style prompts. We propose an adaptive cross-modal fusion mechanism to inject style information into the voice latent representation, enabling interpretable and fine-grained control over speaker identity, speaking rate, and emphasis. Extensive experiments show that our approach significantly outperforms recent baselines in both subjective and objective quality metrics, while offering smoother transitions between styles, reduced artifacts, and more precise text-based style control. 2025-03-26T21:30:29Z arXiv admin note: This paper has been withdrawn by arXiv due to disputed and unverifiable authorship and affiliation Wen Li Sofia Martinez Priyanka Shah http://arxiv.org/abs/2503.20992v2 ReverBERT: A State Space Model for Efficient Text-Driven Speech Style Transfer 2025-07-30T17:02:04Z Text-driven speech style transfer aims to mold the intonation, pace, and timbre of a spoken utterance to match stylistic cues from text descriptions. While existing methods leverage large-scale neural architectures or pre-trained language models, the computational costs often remain high. In this paper, we present \emph{ReverBERT}, an efficient framework for text-driven speech style transfer that draws inspiration from a state space model (SSM) paradigm, loosely motivated by the image-based method of Wang and Liu~\cite{wang2024stylemamba}. Unlike image domain techniques, our method operates in the speech space and integrates a discrete Fourier transform of latent speech features to enable smooth and continuous style modulation. We also propose a novel \emph{Transformer-based SSM} layer for bridging textual style descriptors with acoustic attributes, dramatically reducing inference time while preserving high-quality speech characteristics. Extensive experiments on benchmark speech corpora demonstrate that \emph{ReverBERT} significantly outperforms baselines in terms of naturalness, expressiveness, and computational efficiency. We release our model and code publicly to foster further research in text-driven speech style transfer. 2025-03-26T21:11:17Z arXiv admin note: This paper has been withdrawn by arXiv due to disputed and unverifiable authorship and affiliation Michael Brown Sofia Martinez Priya Singh http://arxiv.org/abs/2503.20988v2 Cross-Modal State-Space Graph Reasoning for Structured Summarization 2025-07-30T17:01:45Z The ability to extract compact, meaningful summaries from large-scale and multimodal data is critical for numerous applications, ranging from video analytics to medical reports. Prior methods in cross-modal summarization have often suffered from high computational overheads and limited interpretability. In this paper, we propose a \textit{Cross-Modal State-Space Graph Reasoning} (\textbf{CSS-GR}) framework that incorporates a state-space model with graph-based message passing, inspired by prior work on efficient state-space models. Unlike existing approaches relying on purely sequential models, our method constructs a graph that captures inter- and intra-modal relationships, allowing more holistic reasoning over both textual and visual streams. We demonstrate that our approach significantly improves summarization quality and interpretability while maintaining computational efficiency, as validated on standard multimodal summarization benchmarks. We also provide a thorough ablation study to highlight the contributions of each component. 2025-03-26T21:06:56Z arXiv admin note: This paper has been withdrawn by arXiv due to disputed and unverifiable authorship and affiliation Hannah Kim Sofia Martinez Jason Lee http://arxiv.org/abs/2503.16133v2 Multi-Prompt Style Interpolation for Fine-Grained Artistic Control 2025-07-29T20:13:57Z Text-driven image style transfer has seen remarkable progress with methods leveraging cross-modal embeddings for fast, high-quality stylization. However, most existing pipelines assume a \emph{single} textual style prompt, limiting the range of artistic control and expressiveness. In this paper, we propose a novel \emph{multi-prompt style interpolation} framework that extends the recently introduced \textbf{StyleMamba} approach. Our method supports blending or interpolating among multiple textual prompts (eg, ``cubism,'' ``impressionism,'' and ``cartoon''), allowing the creation of nuanced or hybrid artistic styles within a \emph{single} image. We introduce a \textit{Multi-Prompt Embedding Mixer} combined with \textit{Adaptive Blending Weights} to enable fine-grained control over the spatial and semantic influence of each style. Further, we propose a \emph{Hierarchical Masked Directional Loss} to refine region-specific style consistency. Experiments and user studies confirm our approach outperforms single-prompt baselines and naive linear combinations of styles, achieving superior style fidelity, text-image alignment, and artistic flexibility, all while maintaining the computational efficiency offered by the state-space formulation. 2025-03-20T13:29:32Z arXiv admin note: This paper has been withdrawn by arXiv due to disputed and unverifiable authorship and affiliation Lei Chen Hao Li Yuxin Zhang Chao Li Kai Wen http://arxiv.org/abs/2503.16129v2 Controllable Segmentation-Based Text-Guided Style Editing 2025-07-29T20:13:40Z We present a novel approach for controllable, region-specific style editing driven by textual prompts. Building upon the state-space style alignment framework introduced by \emph{StyleMamba}, our method integrates a semantic segmentation model into the style transfer pipeline. This allows users to selectively apply text-driven style changes to specific segments (e.g., ``turn the building into a cyberpunk tower'') while leaving other regions (e.g., ``people'' or ``trees'') unchanged. By incorporating region-wise condition vectors and a region-specific directional loss, our method achieves high-fidelity transformations that respect both semantic boundaries and user-driven style descriptions. Extensive experiments demonstrate that our approach can flexibly handle complex scene stylizations in real-world scenarios, improving control and quality over purely global style transfer methods. 2025-03-20T13:24:41Z arXiv admin note: This paper has been withdrawn by arXiv due to disputed and unverifiable authorship and affiliation Jingwen Li Aravind Chandrasekar Mariana Rocha Chao Li Yuqing Chen http://arxiv.org/abs/2503.12291v2 Text-Driven Video Style Transfer with State-Space Models: Extending StyleMamba for Temporal Coherence 2025-07-29T20:13:17Z StyleMamba has recently demonstrated efficient text-driven image style transfer by leveraging state-space models (SSMs) and masked directional losses. In this paper, we extend the StyleMamba framework to handle video sequences. We propose new temporal modules, including a \emph{Video State-Space Fusion Module} to model inter-frame dependencies and a novel \emph{Temporal Masked Directional Loss} that ensures style consistency while addressing scene changes and partial occlusions. Additionally, we introduce a \emph{Temporal Second-Order Loss} to suppress abrupt style variations across consecutive frames. Our experiments on DAVIS and UCF101 show that the proposed approach outperforms competing methods in terms of style consistency, smoothness, and computational efficiency. We believe our new framework paves the way for real-time text-driven video stylization with state-of-the-art perceptual results. 2025-03-15T23:09:03Z arXiv admin note: This paper has been withdrawn by arXiv due to disputed and unverifiable authorship and affiliation Chao Li Minsu Park Cristina Rossi Zhuang Li http://arxiv.org/abs/2411.17489v3 Puzzle Similarity: A Perceptually-guided Cross-Reference Metric for Artifact Detection in 3D Scene Reconstructions 2025-07-29T14:56:41Z Modern reconstruction techniques can effectively model complex 3D scenes from sparse 2D views. However, automatically assessing the quality of novel views and identifying artifacts is challenging due to the lack of ground truth images and the limitations of no-reference image metrics in predicting reliable artifact maps. The absence of such metrics hinders assessment of the quality of novel views and limits the adoption of post-processing techniques, such as inpainting, to enhance reconstruction quality. To tackle this, recent work has established a new category of metrics (cross-reference), predicting image quality solely by leveraging context from alternate viewpoint captures (arXiv:2404.14409). In this work, we propose a new cross-reference metric, Puzzle Similarity, which is designed to localize artifacts in novel views. Our approach utilizes image patch statistics from the training views to establish a scene-specific distribution, later used to identify poorly reconstructed regions in the novel views. Given the lack of good measures to evaluate cross-reference methods in the context of 3D reconstruction, we collected a novel human-labeled dataset of artifact and distortion maps in unseen reconstructed views. Through this dataset, we demonstrate that our method achieves state-of-the-art localization of artifacts in novel views, correlating with human assessment, even without aligned references. We can leverage our new metric to enhance applications like automatic image restoration, guided acquisition, or 3D reconstruction from sparse inputs. Find the project page at https://nihermann.github.io/puzzlesim/ . 2024-11-26T14:57:30Z Nicolai Hermann Jorge Condor Piotr Didyk