https://arxiv.org/api/uDYQwQomFbr3is54uAZBr1EfJIA 2026-09-11T19:58:47Z 22362 30 15 http://arxiv.org/abs/2609.09965v1 Over-Tightening-Aware Pseudo-Labeling for Tight-Boundary Speaker Diarization 2026-09-09T09:52:27Z Training speaker diarization models on loose labels, such as speech segments with padded boundaries or filled pauses, often results in similarly loose model outputs. To obtain tighter boundaries, pseudo-labeling based on the averaged outputs of causal and anticausal models has been proposed. However, since the pseudo-labels are estimation-based, they can suffer from over-tightening, which increases missed detections that can propagate as unrecoverable errors to downstream tasks. This paper carefully analyzes the causes of over-tightening and proposes three approaches to address them: (i) removing pause filling rather than padding, (ii) introducing a burn-in phase to mitigate missed detections near the beginning of causal and anticausal predictions, and (iii) making pseudo-label-based co-training aware of the non-causal model used for final inference. Experimental results show that the proposed method reduces missed detections caused by over-tightening and improves both diarization accuracy and downstream multi-talker ASR performance. 2026-09-09T09:52:27Z Accepted to IEEE SLT 2026 Shota Horiguchi Takanori Ashihara Marc Delcroix Naohiro Tawara Alexis Plaquet http://arxiv.org/abs/2609.08977v2 Omni Interaction Agent Technical Report 2026-09-09T09:47:31Z In this work, we present Gander, an end-to-end model that unifies omni perception, realtime interaction, and agentic capabilities within a single framework. In contrast to turn-based conventional paradigms, Gander continuously receives streaming inputs across multiple modalities, including video, speech, and text, enabling natural full-duplex interaction in both everyday conversations and complex workflow-oriented agent scenarios. Users can interrupt the model at any time, while the model can also proactively provide intermediate feedback or ask follow up questions. To natively support these capabilities, Gander adopts two key architectural designs: 1) It employs a Cerebellum-Brain collaborative framework, in which the Cerebellum is responsible for realtime interaction and omni conversational capabilities, while the Brain handles complex reasoning and higher-level agentic tasks. The two components interact continuously through tool calling and the agent orchestration runtime. 2) The Cerebellum is built upon a streaming Thinker-Talker architecture, user inputs and model outputs are further flattened into an ordered token stream at the chunk level, providing a unified representation for low latency, continuous interaction. We conduct comprehensive evaluations of Gander across four dimensions: conversational ability, omni understanding, interactive capability, and agentic intelligence. Internal human evaluations demonstrate that Gander maintains the natural and expressive spoken dialogue capabilities of SOTA open source models while achieving competitive performance in omni interaction. Gander also demonstrates robustness in challenging real-world scenarios, including background noise interference, multi-party interactions, and backchannel communication. We release Gander together with its models, code, and data to facilitate further research and development in the community. 2026-09-08T16:22:23Z Project Page: https://Omni-Interaction-Gander.github.io/Omni-Interaction-Agent Orantqing Shengpeng Ji Junlong Tong Jialong Zuo Dongjie Fu Di Cao Yangzhuo Li Shangda Wu Franz Evan Theron Veyra Changhao Pan Jingyu Lu Dongchao Yang Zhifei Xie Yang Tan Xiaoyu Shen Xiaoda Yang Wenfu Wang Teddy Sun Steve Yves Zhou Zhao http://arxiv.org/abs/2609.03481v2 Geometric Ceilings on Time-Frequency Masking for Single-Channel Separation 2026-09-09T09:40:30Z Most single-channel separators estimate a source by applying a real gain to the mixture in each time-frequency bin. The optimum of that format, which the oracle masks used as bounds do not attain, is the orthogonal projection of the source onto the line spanned by the mixture, its residual set by the angle between them. Locating an estimator reduces to the block structure of a real-linear operator on stacked spectra, giving a chain of four nested classes whose three larger terms match three assumptions on the prior: zero means, circularity and absence of inter-frequency coupling. Held fixed the chain is a cascade of four orthogonal projections; refitted per frame it collapses onto its first term, attributing the whole residual to one missing real parameter per bin, the phase. When the phase posterior is symmetric about the mixture direction, the minimum mean-square estimate falls back onto the line, with gain the posterior mean of the oracle gain and excess error its variance. On MUSDB18 a posterior mean under a non-circular Gaussian-mixture prior leaves the class yet stays 11.44 dB under the per-frame ceiling, which four times as many components and 7.5x the data do not close; a closed-form gate attributes some 70% of it, in decibels, to the predicted variance. The widest fixed class stays 6.70 dB under the same ceiling. Leaving the class and minimising squared error are conflicting requests: the barrier lies in the criterion rather than in the prior. 2026-09-03T07:38:02Z Submitted to Digital Signal Processing (Elsevier), manuscript DSP-S-26-06261. 33 pages, 6 figures. Companion paper: What Selects, What Reconstructs (submitted to IEEE/ACM TASLP) Maxime Baelde http://arxiv.org/abs/2609.09947v1 SpeechAnnotator: A Context-Aware Multi-Agent Framework and Benchmark for Multidimensional Speech Annotation 2026-09-09T09:37:29Z Recent controllable speech generation requires training data with fine-grained annotations of speaker traits, prosody, emotion, paralinguistic cues, acoustic scenes, and context. Existing workflows often rely on manual correction, paid hosted multimodal services, or fixed processing chains, which limits large-scale data processing through annotation cost, external-service dependence, or weak cross-stage recovery. We introduce SpeechAnnotator, a locally deployable, context-aware multi-agent framework built entirely from open-source models and tools. Supporting frontend modules first obtain speaker-aware segments and final segment transcripts, while prior evidence extractors attach heterogeneous segment-level cues. Three specialist agents then collaborate through shared state: the Planning Agent converts local audio evidence, speaker history, neighboring segments, and recording-level context into field-specific contracts; the Labeling Agent performs contract-guided multimodal prediction for directly observable attributes; and the Review Agent runs a bounded review loop that checks evidence support and cross-segment consistency, triggering relabeling only for unsupported or inconsistent fields. To address the fragmentation of existing evaluation resources across isolated tasks and narrow-domain test sets, we introduce SpeechAnnotator-Bench (SA-Bench), containing 8.87 hours of human-annotated audio across nine source formats, together with SpeechAnnotator-Eval (SA-Eval), which separates Timeline-Eval for speaker-aware timeline recovery, Closed-Eval for finite-set attributes, and Open-Eval for open-ended attributes. Experiments and ablations show that SpeechAnnotator provides a locally deployable alternative to commercial audio-capable systems, while the bounded review loop improves multidimensional annotation through evidence- and context-aware field-level recovery. 2026-09-09T09:37:29Z 15 pages, 3 figures, to be published in NCMMSC 2026 Qirui Zhan Shuiyuan Wang Jingbin Hu Haoyu Zhang Xiaming Ren Jinrui Liang Chaoren Yu Bengu Wu Yunxiang Chen Houdun Liu Su Feng Liumeng Xue Lei Xie http://arxiv.org/abs/2609.09940v1 NVV-Locator: From Transcript Tags to Acoustic Boundaries for Fine-Grained Nonverbal Vocalization Grounding 2026-09-09T09:30:48Z Human speech includes nonverbal vocalizations (NVVs), such as laughter, sighs, breaths, and coughs, which convey affective and interactional information. Existing approaches typically represent NVVs as transcript-level tags, providing limited supervision for their waveform-time boundaries. We present NVV-Locator for fine-grained NVV temporal grounding. We first unify 26 NVV categories across public resources and construct large-scale timestamp-supervised training data through dual-LLM verification, transcript-guided forced alignment, and energy-based boundary refinement. We further introduce NVV-TimeBench, an expert-refined benchmark with 667 utterances and 1,094 events. NVV-Locator uses a non-autoregressive slot-filling architecture to jointly predict lexical timestamps, NVV categories, and event boundaries. On NVV-TimeBench, it achieves 71.0% Micro F1, 70.2% Macro F1, 80.4% Macro mIoU, and 59.6 ms Macro mMAE, outperforming the evaluated large audio model counterparts. Evaluation on an external corpus further demonstrates the cross-corpus generalization of NVV-Locator. 2026-09-09T09:30:48Z Yuang Cao Bingshen Mu Zhennan Lin Guojian Li Haoyue Zhan Jie Liu Chuan Xie Qiang Zhang Liumeng Xue Lei Xie http://arxiv.org/abs/2609.09929v1 Source-Adaptive Data Curation for Bilingual NVV-Aware ASR 2026-09-09T09:19:15Z Nonverbal vocalizations (NVVs), such as laughter, sighs, breaths, and coughs, convey affective and interactional information that conventional automatic speech recognition (ASR) systems often discard. We present a bilingual Mandarin-English system for Track 1 of the NVVSpeech Challenge at ISCSLP 2026, which requires joint transcription of lexical content and 16 NVV categories at their transcript-relative positions. Our NVV-Aware Whisper adapts Whisper-medium through checkpoint-compatible vocabulary remapping, enabling lexical tokens and inline NVV tags to be decoded within a unified autoregressive sequence without expanding the vocabulary. To provide reliable and diverse supervision, we further introduce a source-adaptive data curation strategy that refines public NVV corpora through acoustic augmentation and multimodal LLM filtering, while mining spontaneous NVVs from in-the-wild media through automated preprocessing and annotation. Under the official bilingual evaluation protocol, the proposed system improves final score from 33.32 to 53.61, with ablations confirming the complementary benefits of the proposed data-curation components. 2026-09-09T09:19:15Z ISCSLP 2026 Yuang Cao Qirui Zhan Jingbin Hu Ziyu Zhang Yunxiang Chen Houdun Liu Shuo Feng Bengu Wu Lei Xie Liumeng Xue http://arxiv.org/abs/2609.09903v1 SphereVAE: Hyperspherical Latent Autoencoders for Robust Autoregressive Speech Representation Modeling 2026-09-09T09:01:15Z With the rapid development of speech generation technology, discrete codec representations have been widely used because they provide a stable prediction paradigm. In expressive speech generation, however, the quantization bottleneck of discrete codecs results in information gaps in fine-grained prosody, timbre, pronunciation, and frame-to-frame continuity. Continuous representations (e.g., VAE latents), by eliminating this constraint, have emerged as a more effective alternative for autoregressive modeling. Yet when continuous representations are used as autoregressive prediction targets, prediction errors can accumulate along the generation chain, causing latent drift and degrading long-form stability. To mitigate this problem, we propose SphereVAE, which constrains the VAE latent space to the unit hypersphere. SphereVAE defines a Power Spherical posterior on the hypersphere and regularizes the latent distribution toward a uniform prior, so that information is encoded mainly by directional variation, providing a bounded geometric target for autoregressive prediction and reducing the risk of norm drift. SphereVAE underperforms the standard VAE on reconstruction metrics due to reduced latent freedom. However, when integrated into VoxCPM for zero-shot TTS and long-text generation, it yields lower content error rates with comparable speaker similarity, and shows more stable long-range speaker consistency. These results indicate that an appropriate latent geometric constraint can effectively mitigate autoregressive error accumulation and drift in speech generation. 2026-09-09T09:01:15Z 15 pages, 4 figures. Accepted to NCMMSC 2026 Haoyu Zhang Jingbin Hu Hanke Xie Qirui Zhan Wenhao Li Ziyu Zhang Xiaming Ren Yue Li Xunyu Zhu Zhipeng Chen Lei Xie http://arxiv.org/abs/2609.09866v1 UniStream: Multi-Expert Residual Vector Quantization for 48 kHz Causal Streaming Audio Coding 2026-09-09T08:16:52Z We present UniStream, a fully causal 48 kHz neural audio codec for streaming speech, music, and environmental sounds. At its core is Multi-Expert Residual Vector Quantization (ME-RVQ), which replaces the single shared codebook in each residual quantization layer with four expert codebooks controlled by a deterministic Top-K router. Because routing decisions are derived solely from previously decoded quantized states, the decoder can reproduce the selected experts without transmitting expert identifiers, thereby expanding quantization capacity while adding 5.5M parameters. We further introduce an auxiliary Optimal Transport Conditional Flow Matching (OT-CFM) objective to regularize the quantized latent space during training. The flow module is removed entirely at inference and therefore incurs no runtime overhead. UniStream supports a 12 kbps Top-1 mode and a 22.5 kbps Top-2 mode within a causal 48 kHz encoder-decoder framework, while achieving real-time GPU inference. To complement narrow-band speech metrics, we report 48 kHz ViSQOL audio mode, ViSQOL speech mode, standard VGGish-FAD, DNSMOS P.835, and higher-rate reference comparisons with Opus and EnCodec. At 12 kbps, UniStream-Top1 achieves PESQ and UTMOS scores comparable to EnCodec while reducing speech Mel-D from 13.07 to 8.21. At 22.5 kbps, UniStream-Top2 achieves a ViSQOL speech-mode score of 4.67 and an environmental audio-mode score of 3.96, exceeding all evaluated systems operating at 12 kbps or below in the latter setting. It also comes within 0.03 MOS-LQO of Opus at 24 kbps on speech in ViSQOL audio mode. Ablation studies confirm that ME-RVQ is the primary source of quality improvement, whereas OT-CFM provides perceptual gains on speech with a mild trade-off in spectral distortion. 2026-09-09T08:16:52Z 15 pages, 1 figure, 5 tables. Accepted at NCMMSC 2026 Mingyu Zhao Zhiyong Wu http://arxiv.org/abs/2609.08899v2 From Scores to Evidence: Auditable Decisions Can Improve Speech Deepfake Detection 2026-09-09T06:27:03Z Speech deepfakes can mimic a speaker's voice convincingly enough to deceive listeners and automated systems. This has driven strong progress in speech deepfake detection, but most detectors still end with one score per utterance. That score is useful for ranking systems, yet it says little about why a borderline item should be trusted, deferred, or reviewed. Two utterances can fall in the same score band for different reasons, for example because passive and retrieval evidence disagree or because the keyed probe is unavailable. We ask whether the final decision can remain scalar without discarding that provenance. We answer this question with an auditable decision record that carries four aligned cues into a late calibration step: a passive detector score, a conditional keyed-probe score on a marked derivative, retrieval support, and a speaker-profile margin, together with explicit disagreement coordinates. On the 4,080-example ASVspoof 5 Track 1 matched subset, the fixed retrieval-augmented rule improves on retrieval-only evidence, reducing equal error rate (EER) from 15.84% to 11.91%, and late calibration over the full record reaches 8.43% EER. At a 33.75% review budget, the exposed cue union covers 82.85% of the calibrated model's errors. The best passive WavLM run still reaches 6.71% EER, so we do not present the decision record as a stronger standalone detector. Its contribution is to preserve the evidence behind each surfaced utterance while still producing one operating score for thresholding and review. 2026-09-08T15:38:57Z Mengzhe Geng Yujia Lu Patrick Littell Manuela Kunz Xie Chen http://arxiv.org/abs/2609.09719v1 StreamAlign: Streaming Text-Aligned Speech Tokenization 2026-09-09T05:03:12Z Text-aligned speech tokenization methods have emerged to better align speech tokens with LLM token spaces, enabling more effective utilization of pretrained LLMs. However, they rely on offline automatic speech recognition (ASR), leading to two key limitations: (i) the need for complete utterances before tokenization, precluding real-time streaming, and (ii) vocabulary mismatch between ASR and LLMs, which reduces acoustic granularity from the subword to the word level. We introduce StreamAlign, a text-aligned speech tokenization framework that enables streaming tokenization for real-time speech-text joint modeling. StreamAlign performs online speech-text alignment by combining character-level RNN-Transducer alignment with word-level ASR guidance, mitigating ASR-LLM vocabulary mismatch while preserving recognition accuracy. A proactive word boundary classifier anticipates word completion at chunk boundaries, reducing tokenization latency from 560 ms to 270 ms. On LibriSpeech, StreamAlign achieves the lowest WER and highest UTMOS among evaluated tokenizers. Furthermore, StreamAlign-SLM, a spoken language model trained on StreamAlign units, outperforms other end-to-end spoken language models in speech continuation while achieving the strongest overall consistency on SALMon and spoken StoryCloze. 2026-09-09T05:03:12Z Findings of EMNLP 2026. Project page: https://ishlove77.github.io/StreamAlign/ Kang-wook Kim Jinyoung Park Jinsoo Kim Sehun Lee Sang Hoon Woo Gunhee Kim http://arxiv.org/abs/2609.09656v1 Why Learning Rediscovers the Closed-Form Diagonal Regularizer 2026-09-09T03:10:42Z We identify a diagonal saturation principle in modal inverse problems: when truncation noise is isotropic, the Bayes-optimal Tikhonov shape is a closed-form power law Gamma_k proportional to lambda_k^|s| set by the prior alone, independent of the domain. Berry's random-wave conjecture decorrelates the truncation noise across modes, and Weyl's eigenvalue counting law supplies enough modes for the conclusion to survive empirical Berry violations. Together they predict an approximately flat loss landscape across the per-mode family, leaving narrow scope for a diagonal regularizer to robustly beat the closed form. On FEM-simulated acoustic rooms, the closed form is near-optimal relative to per-room oracle tuning across observation windows, and three diagonal architectures trained on the same data match its reconstruction error within 1 pp despite learning qualitatively different spectra. The framework extends to heat diffusion via a known exponential Green's function correction with no new free parameters. Saturation is restricted to the diagonal family: Learned Iterative Ridge crosses the boundary by exploiting cross-mode coupling, locating where learning starts to help. 2026-09-09T03:10:42Z main paper: 9 pages, 3 figures appendix Jeahn Han Pyojin Kim http://arxiv.org/abs/2608.25384v4 Mandarin Humorous Homophone Recognition and Disambiguation in Automatic Speech Recognition 2026-09-09T02:30:50Z Mandarin homophones remain a key challenge to improving automatic speech recognition (ASR) accuracy due to the amount of potential homophones. Mandarin speakers use this feature casually to convey emotions such as humour. Recent homophone-aware ASR studies have improved recognition accuracy, but intentional homophone twists in speech remain underexplored. In this paper, we identify patterns of homophone-based rhetorical wordplay in Mandarin, referred to as HumourPhone, and propose an ASR Adapter for homophone and HumourPhone recovery. Experimental results show that the proposed approach improves the recall of recognising HumourPhone by over 5\% and achieves a 4.35\% drop in target-span character error rate for homophone correction compared to baseline. These results highlight the need for homophone-aware modelling of lexical ambiguity and rhetorical wordplay in Mandarin speech recognition. 2026-08-26T05:18:49Z Camera-Ready Version at ISCSLP 2026 Sicheng Jin Jinghao Chen Liuheng Zhou Mostafa Shahin Beena Ahmed Aditya Joshi http://arxiv.org/abs/2608.28981v2 V2TATC: Joint Voice-Trajectory Embedding and Dataset for Air Traffic Controller Situational Awareness 2026-09-09T00:19:26Z As air traffic volumes in the National Airspace System continue to expand, in particular at low altitude, the need for scalable decision support tools used by air traffic controllers will also require more development. This article introduces Voice-to-Trajectory for Air Traffic Control, a joint voice communication-flight trajectory data embedding framework, that can be a component of situational awareness in congested airspaces, and assist the development of tools for ATC as they reason in real-time over Automatic Dependent Surveillance-Broadcast trajectories, or the intent expressed by pilots in natural language. We show that these data modalities are not independent and represent a common physical referent: an aircraft flying through the airspace. V2TATC maps a voice instruction and the trajectory of the addressed aircraft to nearby points in a single latent space that can be queried in both directions. It combines a self-supervised trajectory encoder, a frozen speech encoder, a contrastive joint embedding, and a bijective lifting via normalizing flows. We demonstrate V2TATC's effectiveness on the San Francisco Bay Area, for its concentration of major airports, and its mix of commercial and general aviation traffic. Lastly, we release a novel paired voice-trajectory dataset, and report experiments on cross-modal retrieval, ablations, and latent-space analysis. 2026-08-29T01:22:52Z 41 pages, 21 figures, 9 tables Louis Brusset Mathurin Petit Jordan Kam Alexandre Bayen http://arxiv.org/abs/2609.09499v1 Language Orthogonalization of Self-Supervised Speech Representations for Cross-lingual Parkinson's Detection 2026-09-08T22:25:55Z Self-supervised speech models (S3Ms) provide powerful representations for Parkinson's disease (PD) detection, making cross-lingual transfer attractive for languages lacking labeled patient speech. However, these representations also encode language identity, which can confound this transfer: without target-language PD speech, classifiers may separate languages rather than pathology, yielding high specificity but low sensitivity on target patients. We propose \emph{language orthogonalization}, a closed-form ridge residualization of S3M features against external VoxLingua107 language embeddings, fitted using only healthy-control (HC) speech. By removing language-predictable components while retaining pathology-related variation, it produces a less language-dependent geometry in which HC representations concentrate while PD representations disperse. Across five S3M backbones, three speech tasks, and three target languages, our method consistently improves cross-lingual PD-detection performance while correcting the high-specificity/low-sensitivity failure. 2026-09-08T22:25:55Z IEEE SLT 2026 submission Minu Kim Eunjung Yeo Kwanghee Choi June-Woo Kim http://arxiv.org/abs/2603.23723v4 Autoregressive Guidance of Deep Spatially Selective Filters using Bayesian Tracking for Efficient Extraction of Moving Speakers 2026-09-08T21:54:32Z Deep spatially selective filters achieve high-quality enhancement with real-time capable architectures for stationary speakers of known directions. To retain this level of performance in dynamic scenarios where only the speakers' initial directions are given, accurate, yet computationally lightweight tracking algorithms become necessary. Assuming a frame-wise causal processing style, temporal feedback allows for leveraging the enhanced speech signal to improve tracking performance. In this work, we investigate strategies to incorporate the enhanced signal into lightweight tracking algorithms and autoregressively guide deep spatial filters. Our proposed Bayesian tracking algorithms are compatible with arbitrary deep spatial filters. To increase the realism of simulated trajectories during development and evaluation, we develop a synthetic data generation framework based on the social force model. Results validate that the autoregressive incorporation significantly improves the accuracy of our Bayesian trackers, resulting in superior enhancement with none or only negligibly increased computational overhead. Real-world recordings complement these findings and demonstrate the generalizability of our methods to unseen acoustic conditions. 2026-03-24T21:19:45Z Accepted for publication in IEEE/ACM Transactions on Audio, Speech, and Language Processing Jakob Kienegger Timo Gerkmann