https://arxiv.org/api/wcjds00Hr1XInd2KDqdvL9qaT9U 2026-09-11T21:00:47Z 22362 60 15 http://arxiv.org/abs/2606.10738v2 Spatial-Omni: Spatial Audio Understanding Integration in Multimodal LLMs via FOA Encoding 2026-09-07T10:43:36Z Recent multimodal large language models mainly process audio as monaural signals, thereby discarding the spatial cues contained in spatial audio for sound localization, spatial relation reasoning, and spatial scene understanding. We propose Spatial-Omni, a lightweight method that implements SO-Encoder to inject First-Order Ambisonics (FOA) spatial audio into existing Omni LLMs as an independent modality, without modifying their original audio encoders. SO-Encoder provides spatial tokens with limited additional context cost and improves spatial audio understanding through efficient staged training. To support training and evaluation, we construct SO-Dataset, SO-QA, and SO-Bench from open-source data, real recordings, and simulations, containing 400K FOA spatial audio clips and 2.1M spatial question answering pairs. SO-Bench covers 16 spatial audio understanding subtasks, including basic detection and location estimation, spatial relation understanding, and complex spatial reasoning. Experiments show that Spatial-Omni outperforms existing open-source Large Audio-Language Models (LALMs) and Omni LLM models on spatial audio understanding tasks while retaining a reasonable level of general audio understanding. Code and data are available at https://github.com/dieKarotte/Spatial-Omni. 2026-06-09T11:50:06Z Zhiyuan Zhu Yixuan Chen Yiwen Shao Wenxiang Guo Changhao Pan Yu Zhang Yuxiang Wang Wei Liu Houhua Zhang Chengkuan Zeng Wenbo Cheng Yunxi Liu Rui Yang Steve Yves Liefeng Bo Zhou Zhao http://arxiv.org/abs/2609.07173v1 Direction-Preserving Active Noise Control with a Conditional Control-Filter Estimation Network 2026-09-07T08:06:46Z Conventional active noise control (ANC) minimizes the total disturbance at the error microphone without distinguishing desired sound from noise. Direction-preserving ANC (DP-ANC) instead aims to attenuate a noise component arriving from a direction other than the specified desired direction while preserving sound naturally arriving from that direction. Existing approaches typically either require analytical optimization to be repeated for each new observation or estimate and reproduce the desired component through a hear-through secondary-source path. To address these limitations, this paper formulates DP-ANC as a direction-conditioned cancellation-preservation optimization problem. A component-separated objective jointly penalizes residual noise energy and the control response induced by the desired component, with a scalar weighting parameter controlling the cancellation-preservation trade-off. A convolutional network conditioned on the specified desired direction through feature-wise linear modulation (FiLM) is trained using a differentiable secondary-path-aware forward model. At deployment, the network estimates the complete multichannel finite impulse response (FIR) control-filter bank directly from a mixed-reference observation and the specified desired direction in a single forward pass, while retaining the conventional feedforward ANC signal path. Over 3300 evaluation cases, the selected operating point achieves 22.8 dB mean noise reduction with a desired-signal distortion of -11.4 dB. Validation using measured in-ear-device transfer functions further demonstrates consistent performance under measured acoustic configurations. 2026-09-07T08:06:46Z 12 pages, 7 figures, 4 tables Ziyi Yang Zhengding Luo Boxiang Wang Libin Zhang Woon-Seng Gan http://arxiv.org/abs/2607.25351v2 Extracting Voice Styles from Frozen TTS Models via Gradient-Based Inverse Optimization 2026-09-07T06:05:04Z Some text-to-speech systems ship a synthesis model and preset style vectors but withhold the reference encoder that turns a recording into a style vector, so a user cannot obtain a style for a new voice. We recover that vector without the encoder by inverting the released pipeline with gradient descent: all weights stay frozen and only the style vector is optimized, against time-pooled WavLM statistics of one recording of the target. The objective discards the time axis, so no transcript is needed. Two constraints from the release make the result a drop-in style: every row of the timbre style stays on the unit sphere where the presets lie, and the small duration style, which a time-pooled loss cannot see, is fitted to the recording's speaking rate. On SupertonicTTS, over 147 speakers and 100 held-out sentences each, ECAPA-TDNN similarity rises from 0.129 to 0.419, every recovered style starts from a preset and ends closer to its target than that preset was, and the pooled word error rate stays below that of the presets. A verifier at its equal-error point accepts 52% of the recovered voices as the target speaker, against 1% of the presets. 2026-07-28T06:54:59Z 5 pages, 2 figures, 2 tables. Code and demo: https://github.com/kdrkdrkdr/supertonic.embed Gyeongmin Kim http://arxiv.org/abs/2609.03620v2 ToolDF: Tool-Integrated Reasoning for Mixed-Authenticity Audio Deepfake Detection 2026-09-07T05:41:50Z Audio deepfake detection is commonly formulated as clip-level binary classification of single-domain audio. However, real-world manipulated audio can exhibit mixed authenticity, where genuine and manipulated cues coexist across temporal transitions, overlapping sources, or both. This setting requires not only detecting manipulated audio but also localizing the components that provide evidence for the decision. We propose ToolDF, a tool-integrated reasoning framework for mixed-authenticity audio deepfake detection. ToolDF employs an audio large language model as an orchestrator trained with supervised tool-use trajectories. It adaptively analyzes the audio scene, selectively performs source separation, routes components to domain-specific experts, and aggregates their evidence into an interpretable verdict. We further introduce a mixed-authenticity ADD benchmark covering temporal transitions, acoustic overlaps, and hybrid mixtures. Experimental results show that ToolDF achieves the best overall performance on composite-type detection, achieving macro-F1 gains of 3.72 and 14.39 points over the strongest monolithic baseline and a fixed pipeline, respectively, while providing interpretable evidence localized to temporal regions and acoustic sources. Our source code and dataset are publicly available online. 2026-09-03T10:05:45Z To appear in Findings of the Association for Computational Linguistics: EMNLP 2026 Taewoo Kim Young Han Lee Nam In Park Chanwoo Kim http://arxiv.org/abs/2606.22473v2 Interleaved Speech Language Models Latently Work In Text 2026-09-06T09:21:17Z Speech language models (SLMs) increasingly combine speech and text, often by interleaving their tokens within a single sequence. Yet how these two modalities interact in the model's latent space remains unclear. In this work, we analyze interleaved speech--text LMs from different model families and training configurations using three complementary methods. We reveal that these models pass through an implicit latent transcription phase in which the text token matching the spoken word becomes decodable in intermediate layers, despite not being trained for speech recognition. This phenomenon occurs in diverse, natural speech, and intermediate representations also encode likely text continuations. We further show that implicit transcription emerges most when combining text-LM pretraining and speech--text interleaving, and that its prevalence is positively associated with spoken factual-knowledge retrieval. Our analysis sheds light on the internal interaction between speech and text modalities in interleaved SLMs. 2026-06-21T12:33:44Z Preprint. 23 pages, 20 figures, 5 tables Talia Sternberg Gallil Maimon Yossi Adi http://arxiv.org/abs/2609.06488v1 Lead Vocal Separation from Vocal Ensemble Mixtures Using Phoneme Alignment 2026-09-06T09:00:58Z Contemporary a cappella singing often has a lead-and-accompaniment texture, where the lead vocal (Vo) part carries the main melody and the remaining vocal parts provide accompaniment. Owing to their distinct roles, separating the Vo part from the remaining vocal parts, referred to as Vo separation, enables downstream applications such as lyric recognition and minus-one accompaniment generation for vocal ensemble music. Despite these potential applications, acoustic cues for this task are limited because the target and interfering sources are all singing voices with similar acoustic characteristics and often overlap in time, making Vo separation challenging. In this paper, we propose a Vo separation model that uses phoneme alignment of the Vo part as auxiliary information. The proposed model is based on band-split RoPE Transformer (BS-RoFormer), a state-of-the-art music source separation model, and introduces frame-level phoneme labels into its intermediate representations using feature-wise linear modulation (FiLM). Experimental results show that phoneme-alignment conditioning improves Vo separation performance over an audio-only baseline and yields larger average gains than conditioning only on Vo singing/silence activity. Further analysis suggests that the advantage of phoneme-label information is larger when fewer remaining vocal parts share the same phoneme as Vo. 2026-09-06T09:00:58Z Accepted for APSIPA Annual Summit and Conference 2026 Yuma Narahata Tomohiko Nakamura Yuki Saito Hiroshi Saruwatari http://arxiv.org/abs/2608.30974v2 CoJEPA: Combining Contrastive Learning and JEPA for Global-Local Music Representations 2026-09-06T08:38:39Z Joint-Embedding Predictive Architecture (JEPA) has shown strong performance in learning rich representations through self-supervised prediction in latent space. However, it typically relies on teacher--student architecture with an EMA to stabilise training, and can tend to yield uninformative representations. Contrastive learning is stable to train and produces strong global representations, but remains limited on local tasks by the global nature of its objective. In this work, we combine both into CoJEPA: a single shared backbone jointly trained with a JEPA objective on masked sequence tokens and a contrastive objective on the class token. The contrastive gradient provides stability, removing the need for an EMA teacher entirely, while JEPA enriches the sequence tokens via local predictions that contrastive learning alone cannot provide. Crucially, no extra parameters are added to the backbone: the same model is guided towards richer representations purely through the design of its training signal. CoJEPA takes the best of both worlds, outperforming or matching both individual methods across global and local MIR tasks, with a particularly strong advantage on tonal and harmonic understanding, and without any task-specific architectural changes. CoJEPA shows that combining objectives with complementary inductive biases can substitute for scale, encouraging future work to invest in smarter training objectives over ever-larger models. 2026-08-31T15:36:13Z Gabriel Meseguer-Brocal Yuexuan Kong Romain Hennequin http://arxiv.org/abs/2509.23454v2 AudioFuse: Unified Spectral-Temporal Learning via a Hybrid ViT-1D CNN Architecture for Robust Phonocardiogram Classification 2026-09-05T19:38:50Z Biomedical audio signals, such as phonocardiograms (PCG), are inherently rhythmic and contain diagnostic information in both their spectral (tonal) and temporal domains. Standard 2D spectrograms provide rich spectral features but compromise the phase information and temporal precision of the 1D waveform. We propose AudioFuse, an architecture that simultaneously learns from both complementary representations to classify PCGs. To mitigate the overfitting risk common in fusion models, we integrate a custom, wide-and-shallow Vision Transformer (ViT) for spectrograms with a shallow 1D CNN for raw waveforms. On the PhysioNet 2016 dataset, AudioFuse achieves a state-of-the-art competitive ROC-AUC of 0.8608 when trained from scratch, outperforming its spectrogram (0.8066) and waveform (0.8223) baselines. Moreover, it demonstrates superior robustness to domain shift on the challenging PASCAL dataset, maintaining an ROC-AUC of 0.7181 while the spectrogram baseline collapses (0.4873). Fusing complementary representations thus provides a strong inductive bias, enabling the creation of efficient, generalizable classifiers without requiring large-scale pre-training. The complete source code for our baseline and AudioFuse models is publicly available on GitHub at: https://github.com/Saiful185/AudioFuse. 2025-09-27T18:52:50Z Published in the Proceedings of 2026 IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP) main conference. Find the published version at: https://doi.org/10.1109/ICASSP55912.2026.11464477 M. S. B. Siddiqui and U. Saha, "AudioFuse: Unified Spectral-Temporal Learning Via A Hybrid VIT-1D CNN Architecture for Phonocardiogram Classification," ICASSP 2026, Barcelona, Spain, 2026, pp. 5711-5715 Md. Saiful Bari Siddiqui Utsab Saha 10.1109/ICASSP55912.2026.11464477 http://arxiv.org/abs/2605.29948v3 HoliTok: A Continuous Holistic Tokenization with Robust Dual Capabilities of Speech Generation and Understanding 2026-09-05T15:01:58Z Unified speech foundation models require a holistic tokenization space that is both learnable by language models and decodable into high-quality waveforms. Existing speech tokenizers, however, often fail to satisfy these requirements simultaneously, leading to increased architectural complexity and more involved training designs. We propose HoliTok, a continuous Holistic speech Tokenization model designed for unified generation-understanding modeling. HoliTok encodes 48~kHz speech into a compact 25~Hz sequence of 128-dimensional latents. It is trained with a progressive strategy that jointly preserves signal-level fidelity, incorporates semantic information, and maintains strong latent learnability. Based on this tokenization, we build a unified AR+DiT model for speech synthesis and recognition, where the same latent sequence supports both generation-specific and unified generation-understanding tasks. Experiments show that HoliTok achieves competitive reconstruction fidelity, improves generative learnability for high-quality and controllable synthesis, and, among the evaluated representations, is the only one that operates robustly in our unified generation-understanding architecture without additional optimization tricks. These results suggest that HoliTok serves as an effective speech tokenizer and a foundational representation interface for unified spoken language modeling. The code is available at: https://github.com/bovod-sjtu/HoliTok. 2026-05-28T13:55:19Z 14 pages, 2 figures, 8 tables; Accepted by EMNLP 2026 Main Conference Bohan Li Shi Lian Hankun Wang Yiwei Guo Yu Xi Zhihan Li Da Zheng Colin Zhang Kai Yu http://arxiv.org/abs/2609.06101v1 SETEAB: Multiscale approach with Squeeze-and-Excitation Temporal Enhanced Aware Block for Speech Emotion Recognition 2026-09-05T13:53:27Z This paper proposes a novel lightweight multiscale architecture for speech emotion recognition (SER) with three key innovations. First, a depthwise convolution-based subsampling module is introduced to reduce model size and computation while preserving salient emotional cues. Second, a Squeeze-and-Excitation block is integrated to enhance channel-wise recalibration and improve representation robustness. Third, a new Temporal Enhanced Aware Block is designed to strengthen temporal dependency modeling and produce more discriminative emotion-aware features. The proposed model is explicitly designed to jointly improve compactness, recognition performance, and generalizability. Experiments on benchmark SER datasets show that our method achieves higher accuracy with reduced computational complexity, while also delivering stronger cross-corpus performance than most recent advanced networks for SER. 2026-09-05T13:53:27Z Accepted to INTERSPEECH 2026 Duy Vo Kiet Anh Hoang Hao Do http://arxiv.org/abs/2507.23266v4 CUHK-EE Systems for the vTAD Challenge at NCMMSC 2025 2026-09-05T12:03:16Z This paper presents the Voice Timbre Attribute Detection (vTAD) systems developed by the Digital Signal Processing & Speech Technology Laboratory (DSP&STL) of the Department of Electronic Engineering (EE) at The Chinese University of Hong Kong (CUHK) for the 20th National Conference on Human-Computer Speech Communication (NCMMSC 2025) vTAD Challenge. The proposed systems leverage WavLM-Large embeddings with attentive statistical pooling (ASTP) to extract robust speaker representations, followed by two variants of Diff-Net, i.e., Feed-Forward Neural Network (FFN) and Squeeze-and-Excitation-enhanced Residual FFN (SE-ResFFN), to compare timbre attribute intensities between utterance pairs. Experimental results demonstrate that the WavLM-Large+FFN system generalises better to unseen speakers, achieving 77.96% accuracy and 21.79% equal error rate (EER), while the WavLM-Large+SE-ResFFN model excels in the 'Seen' setting with 94.42% accuracy and 5.49% EER. These findings highlight a trade-off between model complexity and generalisation, and underscore the importance of architectural choices in fine-grained speaker modelling. Our analysis also reveals the impact of speaker identity, annotation subjectivity, and data imbalance on system performance, pointing to future directions for improving robustness and fairness in timbre attribute detection. 2025-07-31T05:55:16Z Accepted to the 20th National Conference on Man-Machine Speech Communication (NCMMSC 2025) Man-Machine Speech Communication. NCMMSC 2025. Communications in Computer and Information Science, vol 2662. Springer, Singapore, pp. 290-298, 2026 Aemon Yat Fei Chiu Jingyu Li Yusheng Tian Guangyan Zhang Tan Lee 10.1007/978-981-95-5382-2_23 http://arxiv.org/abs/2411.11353v2 An Investigation of Reprogramming for Cross-Language Adaptation in Speaker Verification Systems 2026-09-05T11:57:03Z Language mismatch is among the most common and challenging domain mismatches in deploying speaker verification (SV) systems. Adversarial reprogramming has shown promising results in cross-language adaptation for SV. The reprogramming is implemented by padding learnable parameters on the two sides of input speech signals. In this paper, we investigate the relationship between the number of padded parameters and the performance of the reprogrammed models. Sufficient experiments are conducted with different scales of SV models and datasets. The results demonstrate that reprogramming consistently improves the performance of cross-language SV, while the improvement is saturated or even degraded when using larger padding lengths. The performance is mainly determined by the capacity of the original SV models instead of the number of padded parameters. The SV models with larger scales have higher upper bounds in performance and can endure longer padding without performance degradation. 2024-11-18T07:47:21Z Accepted to the 14th International Symposium on Chinese Spoken Language Processing (ISCSLP 2024) Proceedings of the 2024 IEEE 14th International Symposium on Chinese Spoken Language Processing (ISCSLP), pp. 388-392, 2024 Jingyu Li Aemon Yat Fei Chiu Tan Lee 10.1109/ISCSLP63861.2024.10800573 http://arxiv.org/abs/2501.05310v4 A Large-Scale Probing Analysis of Speaker-Specific Attributes in Self-Supervised Speech Representations 2026-09-05T11:50:51Z Enhancing explainability in speech self-supervised learning (SSL) is important for understanding and effectively utilising speech SSL representations. This study conducts a large-scale layer-wise probing analysis of 11 speech SSL models, examining speaker identity together with acoustic, prosodic, and paralinguistic attributes. The results confirm a general hierarchy wherein initial layers encode fundamental acoustics and middle layers synthesise abstract traits. The consensus that final layers purely abstract linguistic content is challenged, as larger models unexpectedly recover speaker-discriminative information in their deep layers. Furthermore, the intermediate representations of speech SSL models are found to provide stronger probing performance for several attributes than specialised speaker embeddings. These insights give a systematic empirical characterisation of speaker-related information in SSL models, providing guidelines for selecting representations for downstream tasks. 2025-01-09T15:26:33Z Accepted to the 15th International Symposium on Chinese Spoken Language Processing (ISCSLP 2026) Aemon Yat Fei Chiu Kei Ching Fung Roger Tsz Yeung Li Jingyu Li Tan Lee http://arxiv.org/abs/2104.00239v3 Positive Sample Propagation along the Audio-Visual Event Line 2026-09-05T11:30:39Z Visual and audio signals often coexist in natural environments, forming audio-visual events (AVEs). Given a video, we aim to localize video segments containing an AVE and identify its category. In order to learn discriminative features for a classifier, it is pivotal to identify the helpful (or positive) audio-visual segment pairs while filtering out the irrelevant ones, regardless whether they are synchronized or not. To this end, we propose a new positive sample propagation (PSP) module to discover and exploit the closely related audio-visual pairs by evaluating the relationship within every possible pair. It can be done by constructing an all-pair similarity map between each audio and visual segment, and only aggregating the features from the pairs with high similarity scores. To encourage the network to extract high correlated features for positive samples, a new audio-visual pair similarity loss is proposed. We also propose a new weighting branch to better exploit the temporal correlations in weakly supervised setting. We perform extensive experiments on the public AVE dataset and achieve new state-of-the-art accuracy in both fully and weakly supervised settings, thus verifying the effectiveness of our method. 2021-04-01T03:53:57Z Accepted to CVPR 2021. Code is available at https://github.com/jasongief/PSP_CVPR_2021 Jinxing Zhou Liang Zheng Yiran Zhong Shijie Hao Meng Wang http://arxiv.org/abs/2608.27176v2 When Text Misleads: Inconsistent-Aware Reasoning for Audio-Grounded Dialogue 2026-09-05T08:31:34Z Understanding spoken dialogue requires joint reasoning over lexical content and paralinguistic acoustic signals such as emotion and conversational intent. However, existing evaluations often allow shortcuts based on transcripts or single-modality solutions, obscuring whether models genuinely ground predictions in speech. We formalize this failure mode as cross-modal disagreement, where transcripts suggest plausible but incorrect surface interpretations while acoustic cues such as prosody or speaking style support different answers. We develop a scalable framework that identifies text-biased surface interpretations and converts disagreement regions into conflict QA examples. We also include consistent cases where transcript-based and speech-grounded interpretations agree, enabling evaluation beyond adversarial audio dependence. This results in ContraTalk, a controlled benchmark containing 501 questions across five discourse dimensions: interaction behavior, emotion state, dialogue act, social stance, and conversational intent. We further develop an agentic-style reasoning framework that converts speech into an Audio Twin, a text-readable representation of localized acoustic cues that exposes acoustic evidence to the reasoning model. Experiments show that strong text-only LLMs exceed 90% accuracy in consistent cases but drop to 33-48% in conflict cases. Direct AudioLLMs provide only partial grounding, still selecting the transcript-biased trap in roughly 30-40% of conflict cases. Our Audio Twin framework improves conflict-case accuracy while reducing trap selection, but its consistent-case behavior remains backbone-dependent. These results identify transcript-based shortcuts as an important failure mode in spoken dialogue understanding and show that explicit acoustic evidence aggregation provides a more controllable interface for diagnosing and improving speech-grounded reasoning. 2026-08-27T14:26:46Z 24 pages, 4 figures Yen-Ju Lu Yuzhe Wang Yaohan Guan Xiluo He Jiarui Hai Mingrui Liang Kaavya Chaparala Thomas Thebaud Laureano Moro-Velazquez Najim Dehak Jesus Villalba