https://arxiv.org/api/Gf/JkTCgF0Gso6I8sKw/bBO/x7w2026-09-10T17:25:50Z223481515http://arxiv.org/abs/2609.09903v1SphereVAE: Hyperspherical Latent Autoencoders for Robust Autoregressive Speech Representation Modeling2026-09-09T09:01:15ZWith the rapid development of speech generation technology, discrete codec representations have been widely used because they provide a stable prediction paradigm. In expressive speech generation, however, the quantization bottleneck of discrete codecs results in information gaps in fine-grained prosody, timbre, pronunciation, and frame-to-frame continuity. Continuous representations (e.g., VAE latents), by eliminating this constraint, have emerged as a more effective alternative for autoregressive modeling. Yet when continuous representations are used as autoregressive prediction targets, prediction errors can accumulate along the generation chain, causing latent drift and degrading long-form stability. To mitigate this problem, we propose SphereVAE, which constrains the VAE latent space to the unit hypersphere. SphereVAE defines a Power Spherical posterior on the hypersphere and regularizes the latent distribution toward a uniform prior, so that information is encoded mainly by directional variation, providing a bounded geometric target for autoregressive prediction and reducing the risk of norm drift. SphereVAE underperforms the standard VAE on reconstruction metrics due to reduced latent freedom. However, when integrated into VoxCPM for zero-shot TTS and long-text generation, it yields lower content error rates with comparable speaker similarity, and shows more stable long-range speaker consistency. These results indicate that an appropriate latent geometric constraint can effectively mitigate autoregressive error accumulation and drift in speech generation.2026-09-09T09:01:15Z15 pages, 4 figures. Accepted to NCMMSC 2026Haoyu ZhangJingbin HuHanke XieQirui ZhanWenhao LiZiyu ZhangXiaming RenYue LiXunyu ZhuZhipeng ChenLei Xiehttp://arxiv.org/abs/2609.09866v1UniStream: Multi-Expert Residual Vector Quantization for 48 kHz Causal Streaming Audio Coding2026-09-09T08:16:52ZWe present UniStream, a fully causal 48 kHz neural audio codec for streaming speech, music, and environmental sounds. At its core is Multi-Expert Residual Vector Quantization (ME-RVQ), which replaces the single shared codebook in each residual quantization layer with four expert codebooks controlled by a deterministic Top-K router. Because routing decisions are derived solely from previously decoded quantized states, the decoder can reproduce the selected experts without transmitting expert identifiers, thereby expanding quantization capacity while adding 5.5M parameters. We further introduce an auxiliary Optimal Transport Conditional Flow Matching (OT-CFM) objective to regularize the quantized latent space during training. The flow module is removed entirely at inference and therefore incurs no runtime overhead. UniStream supports a 12 kbps Top-1 mode and a 22.5 kbps Top-2 mode within a causal 48 kHz encoder-decoder framework, while achieving real-time GPU inference. To complement narrow-band speech metrics, we report 48 kHz ViSQOL audio mode, ViSQOL speech mode, standard VGGish-FAD, DNSMOS P.835, and higher-rate reference comparisons with Opus and EnCodec. At 12 kbps, UniStream-Top1 achieves PESQ and UTMOS scores comparable to EnCodec while reducing speech Mel-D from 13.07 to 8.21. At 22.5 kbps, UniStream-Top2 achieves a ViSQOL speech-mode score of 4.67 and an environmental audio-mode score of 3.96, exceeding all evaluated systems operating at 12 kbps or below in the latter setting. It also comes within 0.03 MOS-LQO of Opus at 24 kbps on speech in ViSQOL audio mode. Ablation studies confirm that ME-RVQ is the primary source of quality improvement, whereas OT-CFM provides perceptual gains on speech with a mild trade-off in spectral distortion.2026-09-09T08:16:52Z15 pages, 1 figure, 5 tables. Accepted at NCMMSC 2026Mingyu ZhaoZhiyong Wuhttp://arxiv.org/abs/2609.08899v2From Scores to Evidence: Auditable Decisions Can Improve Speech Deepfake Detection2026-09-09T06:27:03ZSpeech deepfakes can mimic a speaker's voice convincingly enough to deceive listeners and automated systems. This has driven strong progress in speech deepfake detection, but most detectors still end with one score per utterance. That score is useful for ranking systems, yet it says little about why a borderline item should be trusted, deferred, or reviewed. Two utterances can fall in the same score band for different reasons, for example because passive and retrieval evidence disagree or because the keyed probe is unavailable. We ask whether the final decision can remain scalar without discarding that provenance. We answer this question with an auditable decision record that carries four aligned cues into a late calibration step: a passive detector score, a conditional keyed-probe score on a marked derivative, retrieval support, and a speaker-profile margin, together with explicit disagreement coordinates. On the 4,080-example ASVspoof 5 Track 1 matched subset, the fixed retrieval-augmented rule improves on retrieval-only evidence, reducing equal error rate (EER) from 15.84% to 11.91%, and late calibration over the full record reaches 8.43% EER. At a 33.75% review budget, the exposed cue union covers 82.85% of the calibrated model's errors. The best passive WavLM run still reaches 6.71% EER, so we do not present the decision record as a stronger standalone detector. Its contribution is to preserve the evidence behind each surfaced utterance while still producing one operating score for thresholding and review.2026-09-08T15:38:57ZMengzhe GengYujia LuPatrick LittellManuela KunzXie Chenhttp://arxiv.org/abs/2609.09719v1StreamAlign: Streaming Text-Aligned Speech Tokenization2026-09-09T05:03:12ZText-aligned speech tokenization methods have emerged to better align speech tokens with LLM token spaces, enabling more effective utilization of pretrained LLMs. However, they rely on offline automatic speech recognition (ASR), leading to two key limitations: (i) the need for complete utterances before tokenization, precluding real-time streaming, and (ii) vocabulary mismatch between ASR and LLMs, which reduces acoustic granularity from the subword to the word level. We introduce StreamAlign, a text-aligned speech tokenization framework that enables streaming tokenization for real-time speech-text joint modeling. StreamAlign performs online speech-text alignment by combining character-level RNN-Transducer alignment with word-level ASR guidance, mitigating ASR-LLM vocabulary mismatch while preserving recognition accuracy. A proactive word boundary classifier anticipates word completion at chunk boundaries, reducing tokenization latency from 560 ms to 270 ms. On LibriSpeech, StreamAlign achieves the lowest WER and highest UTMOS among evaluated tokenizers. Furthermore, StreamAlign-SLM, a spoken language model trained on StreamAlign units, outperforms other end-to-end spoken language models in speech continuation while achieving the strongest overall consistency on SALMon and spoken StoryCloze.2026-09-09T05:03:12ZFindings of EMNLP 2026. Project page: https://ishlove77.github.io/StreamAlign/Kang-wook KimJinyoung ParkJinsoo KimSehun LeeSang Hoon WooGunhee Kimhttp://arxiv.org/abs/2609.09656v1Why Learning Rediscovers the Closed-Form Diagonal Regularizer2026-09-09T03:10:42ZWe identify a diagonal saturation principle in modal inverse problems: when truncation noise is isotropic, the Bayes-optimal Tikhonov shape is a closed-form power law Gamma_k proportional to lambda_k^|s| set by the prior alone, independent of the domain. Berry's random-wave conjecture decorrelates the truncation noise across modes, and Weyl's eigenvalue counting law supplies enough modes for the conclusion to survive empirical Berry violations. Together they predict an approximately flat loss landscape across the per-mode family, leaving narrow scope for a diagonal regularizer to robustly beat the closed form. On FEM-simulated acoustic rooms, the closed form is near-optimal relative to per-room oracle tuning across observation windows, and three diagonal architectures trained on the same data match its reconstruction error within 1 pp despite learning qualitatively different spectra. The framework extends to heat diffusion via a known exponential Green's function correction with no new free parameters. Saturation is restricted to the diagonal family: Learned Iterative Ridge crosses the boundary by exploiting cross-mode coupling, locating where learning starts to help.2026-09-09T03:10:42Zmain paper: 9 pages, 3 figures appendixJeahn HanPyojin Kimhttp://arxiv.org/abs/2608.25384v4Mandarin Humorous Homophone Recognition and Disambiguation in Automatic Speech Recognition2026-09-09T02:30:50ZMandarin homophones remain a key challenge to improving automatic speech recognition (ASR) accuracy due to the amount of potential homophones. Mandarin speakers use this feature casually to convey emotions such as humour. Recent homophone-aware ASR studies have improved recognition accuracy, but intentional homophone twists in speech remain underexplored. In this paper, we identify patterns of homophone-based rhetorical wordplay in Mandarin, referred to as HumourPhone, and propose an ASR Adapter for homophone and HumourPhone recovery. Experimental results show that the proposed approach improves the recall of recognising HumourPhone by over 5\% and achieves a 4.35\% drop in target-span character error rate for homophone correction compared to baseline. These results highlight the need for homophone-aware modelling of lexical ambiguity and rhetorical wordplay in Mandarin speech recognition.2026-08-26T05:18:49ZCamera-Ready Version at ISCSLP 2026Sicheng JinJinghao ChenLiuheng ZhouMostafa ShahinBeena AhmedAditya Joshihttp://arxiv.org/abs/2608.28981v2V2TATC: Joint Voice-Trajectory Embedding and Dataset for Air Traffic Controller Situational Awareness2026-09-09T00:19:26ZAs air traffic volumes in the National Airspace System continue to expand, in particular at low altitude, the need for scalable decision support tools used by air traffic controllers will also require more development. This article introduces Voice-to-Trajectory for Air Traffic Control, a joint voice communication-flight trajectory data embedding framework, that can be a component of situational awareness in congested airspaces, and assist the development of tools for ATC as they reason in real-time over Automatic Dependent Surveillance-Broadcast trajectories, or the intent expressed by pilots in natural language. We show that these data modalities are not independent and represent a common physical referent: an aircraft flying through the airspace. V2TATC maps a voice instruction and the trajectory of the addressed aircraft to nearby points in a single latent space that can be queried in both directions. It combines a self-supervised trajectory encoder, a frozen speech encoder, a contrastive joint embedding, and a bijective lifting via normalizing flows. We demonstrate V2TATC's effectiveness on the San Francisco Bay Area, for its concentration of major airports, and its mix of commercial and general aviation traffic. Lastly, we release a novel paired voice-trajectory dataset, and report experiments on cross-modal retrieval, ablations, and latent-space analysis.2026-08-29T01:22:52Z41 pages, 21 figures, 9 tablesLouis BrussetMathurin PetitJordan KamAlexandre Bayenhttp://arxiv.org/abs/2609.09499v1Language Orthogonalization of Self-Supervised Speech Representations for Cross-lingual Parkinson's Detection2026-09-08T22:25:55ZSelf-supervised speech models (S3Ms) provide powerful representations for Parkinson's disease (PD) detection, making cross-lingual transfer attractive for languages lacking labeled patient speech. However, these representations also encode language identity, which can confound this transfer: without target-language PD speech, classifiers may separate languages rather than pathology, yielding high specificity but low sensitivity on target patients. We propose \emph{language orthogonalization}, a closed-form ridge residualization of S3M features against external VoxLingua107 language embeddings, fitted using only healthy-control (HC) speech. By removing language-predictable components while retaining pathology-related variation, it produces a less language-dependent geometry in which HC representations concentrate while PD representations disperse. Across five S3M backbones, three speech tasks, and three target languages, our method consistently improves cross-lingual PD-detection performance while correcting the high-specificity/low-sensitivity failure.2026-09-08T22:25:55ZIEEE SLT 2026 submissionMinu KimEunjung YeoKwanghee ChoiJune-Woo Kimhttp://arxiv.org/abs/2603.23723v4Autoregressive Guidance of Deep Spatially Selective Filters using Bayesian Tracking for Efficient Extraction of Moving Speakers2026-09-08T21:54:32ZDeep spatially selective filters achieve high-quality enhancement with real-time capable architectures for stationary speakers of known directions. To retain this level of performance in dynamic scenarios where only the speakers' initial directions are given, accurate, yet computationally lightweight tracking algorithms become necessary. Assuming a frame-wise causal processing style, temporal feedback allows for leveraging the enhanced speech signal to improve tracking performance. In this work, we investigate strategies to incorporate the enhanced signal into lightweight tracking algorithms and autoregressively guide deep spatial filters. Our proposed Bayesian tracking algorithms are compatible with arbitrary deep spatial filters. To increase the realism of simulated trajectories during development and evaluation, we develop a synthetic data generation framework based on the social force model. Results validate that the autoregressive incorporation significantly improves the accuracy of our Bayesian trackers, resulting in superior enhancement with none or only negligibly increased computational overhead. Real-world recordings complement these findings and demonstrate the generalizability of our methods to unseen acoustic conditions.2026-03-24T21:19:45ZAccepted for publication in IEEE/ACM Transactions on Audio, Speech, and Language ProcessingJakob KieneggerTimo Gerkmannhttp://arxiv.org/abs/2608.24558v3Array-Agnostic Ambisonics Encoding via Diffusion Posterior Sampling2026-09-08T15:48:00ZSpatial audio enhances user immersion by reproducing 3D sound fields, with Ambisonics being a widely adopted representation. While Ambisonics is theoretically independent of the recording setup, practical microphone arrays introduce hardware-dependent encoding artifacts. Moreover, existing data-driven solutions lack flexibility, as they are typically restricted to fixed array geometries. To overcome these limitations, we propose ADEPS, a generative framework that explicitly embeds the physical acquisition model into the inference process. By leveraging this formulation, ADEPS effectively compensates for array-specific distortions while enabling zero-shot encoding across arbitrary array topologies. We train the underlying generative prior in an unsupervised manner solely on target Ambisonic representations. Extensive evaluations across diverse simulated and real microphone arrays demonstrate that ADEPS consistently outperforms both traditional linear and parametric baselines in spatial fidelity and spectral quality.2026-08-25T13:43:31ZAmit MilsteinNir ShlezingerBoaz Rafaelyhttp://arxiv.org/abs/2609.08795v1Interpreting Dolphin Vocal Sequences via Multiple Sequence Alignment2026-09-08T14:23:10ZDolphin communication understanding is essential for uncovering the linguistic complexity and social structures of wild pods. We adapt the ClustalW bioinformatics algorithm to analyze continuous acoustic data, treating vocalizations as high-dimensional spectral feature vectors. By replacing discrete scoring with a continuous Gaussian kernel similarity measure, our framework generates Multiple Sequence Alignment (MSA) visualizations that reveal shared structural patterns. These alignments highlight temporal motifs such as synchronized burst pulses in aggressive contexts that are difficult to detect through standard spectrogram inspection.2026-09-08T14:23:10ZDaniel KohlsdorfDenise HerzingThad Starnerhttp://arxiv.org/abs/2608.10878v3X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction2026-09-08T12:28:34ZAccurate and responsive turn-taking is essential for spoken dialogue systems, which must distinguish in real time between user interruptions, backchannels that should be ignored, and the completion of an utterance. Prior modular approaches typically optimize turn state prediction at the utterance or fixed-chunk level, creating a mismatch with the continuous turn state estimate, and often depend on an auxiliary ASR model, which limits responsiveness and increases overall system complexity. Therefore, we present X2-Turn, a frame-synchronous turn state prediction method via delayed-stream modeling. Specifically, building on the pretrained Voxtral Realtime model, we introduce a frame-synchronous turn state head that operates in parallel with the ASR head on shared streaming representations, jointly predicting ASR tokens and fine-grained turn states at the frame level. Experiments on bilingual EasyTurn and Full-Duplex-Bench demonstrate that the proposed method achieves an effective trade-off between turn state accuracy and decision latency.2026-08-11T12:54:52ZKaiqi FuRime WenAltman LinShawn QinRoy GanHao WangQian Wanghttp://arxiv.org/abs/2609.08587v1Rescuing Performance from the Demo: Co-Designing Drum Gesture Mappings with a Percussionist2026-09-08T11:23:33ZAugmenting instruments with sensors and neural network mappings is a well-explored digital musical instrument design approach. While augmentations can create new expressive opportunities, they also exert aesthetic influence and can constrain musicians' gestural language, which, if left unchecked, can lead to technological capture. To examine this, we conducted a study with a professional percussionist, co-developing a gesture mapping toolkit and recording a ten-track album. Drawing on the concept of productive dissonance, our study aimed to hold the musician's aesthetic in tension with technological constraints. This, along with a practice-based reflective approach, supported the development of a continuous gesture recognition method for percussive mapping and surfaced insights into the design process. We identify knowing-when as a form of tacit knowledge that supported productive dissonance, and raise an open question: absent a musician's broader social context, how do we know whether a technology's influence is genuinely supporting their practice?2026-09-08T11:23:33ZAccepted for publication at AIMC 2026Jordie ShierTeresa PelinskiCharalampos SaitisAndrew RobertsonAndrew McPhersonhttp://arxiv.org/abs/2605.12036v2Towards Fine-Grained Multi-Dimensional Speech Understanding: Data Pipeline, Benchmark, and Model2026-09-08T10:58:33ZWhile speech Large Language Models (LLMs) excel at conventional tasks like basic speech recognition, they lack fine-grained, multi-dimensional perception. This deficiency is evident in their struggle to disentangle complex features like micro-acoustic cues, acoustic scenes, and paralinguistic signals. This resulting incomplete comprehension of real-world speech fundamentally bottlenecks the development of perceptive and empathetic next-generation speech systems. At its core, this persistent perceptual limitation primarily stems from three interacting factors: scarce high-quality expressive data, absent fine-grained modeling for multi-dimensional attributes, and reliance on restricted coverage, coarse-grained benchmarks. We address these challenges through three pillars: First, our robust data curation pipeline resolves complex acoustic environments and long-audio timestamp alignment challenges to extract a high-quality spontaneous speech corpus from audiovisual sources. Second, we construct FMSU-Bench, a pioneering benchmark covering 14 speech attribute dimensions to rigorously assess the fine-grained, multi-dimensional speech understanding capabilities of current models. Third, empowered by our curated corpus, we introduce FM-Speech. Driven by a decoupled attribute modeling and progressive curriculum fine-tuning framework, it substantially elevates fine-grained, multi-dimensional acoustic perception. Extensive evaluations on FMSU-Bench reveal that current speech LLMs still require significant improvement in multi-dimensional, fine-grained understanding. In contrast, FM-Speech substantially outperforms current open-source models, establishing a robust paradigm for real-world speech understanding.2026-05-12T12:19:33ZGuojian LiZhixian ZhaoZhennan LinJingbin HuQirui ZhanYuang CaoPengyuan XieChuan XieJie LiuQiang ZhangZhonghua FuLei Xiehttp://arxiv.org/abs/2609.08542v1Spatial Audio Coding Through Relative Room Impulse Response Estimation2026-09-08T10:26:01ZImmersive virtual listening relies on spatial audio technologies such as Higher-Order Ambisonics (HOA), which represent sound scenes as multichannel signals. As the desired spatial resolution increases, so does the number of channels, making efficient compression essential for transmission over bandwidth-limited networks. Moreover, to facilitate deployment by network operators, the target bitrate for immersive audio coding should ideally remain close to the 25 kbps currently allocated to VoLTE audio services. State-of-the-art parametric codecs, such as the recently standardized Immersive Voice and Audio Services (IVAS) codec, achieve compression by transmitting spatial metadata together with a reduced number of transport channels. However, recent studies have shown that IVAS performance degrades on reverberant content, particularly at low bitrates, a limitation that suggests its inability to accurately model room acoustics. In this paper, we propose a novel HOA coding scheme based on the explicit and blind estimation of the Relative Spatial Room Impulse Response (ReSRIR), using a beamformed version of the HOA signal as a reference signal. By exploiting the structure and sparsity of the estimated ReSRIR, we derive an efficient parametric representation for immersive audio coding. Experimental evaluations show that the proposed method achieves higher compression than IVAS in the single-transport-channel regime, while maintaining comparable to slightly better quality.2026-09-08T10:26:01ZSubmitted to International Networked Immersive Audio 2026 (satellite Event of IEEE IS2 2026)Nour BouayedAdrien LlaveJérôme DanielPascal Scalart