https://arxiv.org/api/z2iquAwRaikxRk5dBEfSiL354tE2026-09-11T20:43:53Z223624515http://arxiv.org/abs/2608.24558v3Array-Agnostic Ambisonics Encoding via Diffusion Posterior Sampling2026-09-08T15:48:00ZSpatial audio enhances user immersion by reproducing 3D sound fields, with Ambisonics being a widely adopted representation. While Ambisonics is theoretically independent of the recording setup, practical microphone arrays introduce hardware-dependent encoding artifacts. Moreover, existing data-driven solutions lack flexibility, as they are typically restricted to fixed array geometries. To overcome these limitations, we propose ADEPS, a generative framework that explicitly embeds the physical acquisition model into the inference process. By leveraging this formulation, ADEPS effectively compensates for array-specific distortions while enabling zero-shot encoding across arbitrary array topologies. We train the underlying generative prior in an unsupervised manner solely on target Ambisonic representations. Extensive evaluations across diverse simulated and real microphone arrays demonstrate that ADEPS consistently outperforms both traditional linear and parametric baselines in spatial fidelity and spectral quality.2026-08-25T13:43:31ZAmit MilsteinNir ShlezingerBoaz Rafaelyhttp://arxiv.org/abs/2609.08795v1Interpreting Dolphin Vocal Sequences via Multiple Sequence Alignment2026-09-08T14:23:10ZDolphin communication understanding is essential for uncovering the linguistic complexity and social structures of wild pods. We adapt the ClustalW bioinformatics algorithm to analyze continuous acoustic data, treating vocalizations as high-dimensional spectral feature vectors. By replacing discrete scoring with a continuous Gaussian kernel similarity measure, our framework generates Multiple Sequence Alignment (MSA) visualizations that reveal shared structural patterns. These alignments highlight temporal motifs such as synchronized burst pulses in aggressive contexts that are difficult to detect through standard spectrogram inspection.2026-09-08T14:23:10ZDaniel KohlsdorfDenise HerzingThad Starnerhttp://arxiv.org/abs/2608.10878v3X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction2026-09-08T12:28:34ZAccurate and responsive turn-taking is essential for spoken dialogue systems, which must distinguish in real time between user interruptions, backchannels that should be ignored, and the completion of an utterance. Prior modular approaches typically optimize turn state prediction at the utterance or fixed-chunk level, creating a mismatch with the continuous turn state estimate, and often depend on an auxiliary ASR model, which limits responsiveness and increases overall system complexity. Therefore, we present X2-Turn, a frame-synchronous turn state prediction method via delayed-stream modeling. Specifically, building on the pretrained Voxtral Realtime model, we introduce a frame-synchronous turn state head that operates in parallel with the ASR head on shared streaming representations, jointly predicting ASR tokens and fine-grained turn states at the frame level. Experiments on bilingual EasyTurn and Full-Duplex-Bench demonstrate that the proposed method achieves an effective trade-off between turn state accuracy and decision latency.2026-08-11T12:54:52ZKaiqi FuRime WenAltman LinShawn QinRoy GanHao WangQian Wanghttp://arxiv.org/abs/2609.08587v1Rescuing Performance from the Demo: Co-Designing Drum Gesture Mappings with a Percussionist2026-09-08T11:23:33ZAugmenting instruments with sensors and neural network mappings is a well-explored digital musical instrument design approach. While augmentations can create new expressive opportunities, they also exert aesthetic influence and can constrain musicians' gestural language, which, if left unchecked, can lead to technological capture. To examine this, we conducted a study with a professional percussionist, co-developing a gesture mapping toolkit and recording a ten-track album. Drawing on the concept of productive dissonance, our study aimed to hold the musician's aesthetic in tension with technological constraints. This, along with a practice-based reflective approach, supported the development of a continuous gesture recognition method for percussive mapping and surfaced insights into the design process. We identify knowing-when as a form of tacit knowledge that supported productive dissonance, and raise an open question: absent a musician's broader social context, how do we know whether a technology's influence is genuinely supporting their practice?2026-09-08T11:23:33ZAccepted for publication at AIMC 2026Jordie ShierTeresa PelinskiCharalampos SaitisAndrew RobertsonAndrew McPhersonhttp://arxiv.org/abs/2605.12036v2Towards Fine-Grained Multi-Dimensional Speech Understanding: Data Pipeline, Benchmark, and Model2026-09-08T10:58:33ZWhile speech Large Language Models (LLMs) excel at conventional tasks like basic speech recognition, they lack fine-grained, multi-dimensional perception. This deficiency is evident in their struggle to disentangle complex features like micro-acoustic cues, acoustic scenes, and paralinguistic signals. This resulting incomplete comprehension of real-world speech fundamentally bottlenecks the development of perceptive and empathetic next-generation speech systems. At its core, this persistent perceptual limitation primarily stems from three interacting factors: scarce high-quality expressive data, absent fine-grained modeling for multi-dimensional attributes, and reliance on restricted coverage, coarse-grained benchmarks. We address these challenges through three pillars: First, our robust data curation pipeline resolves complex acoustic environments and long-audio timestamp alignment challenges to extract a high-quality spontaneous speech corpus from audiovisual sources. Second, we construct FMSU-Bench, a pioneering benchmark covering 14 speech attribute dimensions to rigorously assess the fine-grained, multi-dimensional speech understanding capabilities of current models. Third, empowered by our curated corpus, we introduce FM-Speech. Driven by a decoupled attribute modeling and progressive curriculum fine-tuning framework, it substantially elevates fine-grained, multi-dimensional acoustic perception. Extensive evaluations on FMSU-Bench reveal that current speech LLMs still require significant improvement in multi-dimensional, fine-grained understanding. In contrast, FM-Speech substantially outperforms current open-source models, establishing a robust paradigm for real-world speech understanding.2026-05-12T12:19:33ZGuojian LiZhixian ZhaoZhennan LinJingbin HuQirui ZhanYuang CaoPengyuan XieChuan XieJie LiuQiang ZhangZhonghua FuLei Xiehttp://arxiv.org/abs/2609.08542v1Spatial Audio Coding Through Relative Room Impulse Response Estimation2026-09-08T10:26:01ZImmersive virtual listening relies on spatial audio technologies such as Higher-Order Ambisonics (HOA), which represent sound scenes as multichannel signals. As the desired spatial resolution increases, so does the number of channels, making efficient compression essential for transmission over bandwidth-limited networks. Moreover, to facilitate deployment by network operators, the target bitrate for immersive audio coding should ideally remain close to the 25 kbps currently allocated to VoLTE audio services. State-of-the-art parametric codecs, such as the recently standardized Immersive Voice and Audio Services (IVAS) codec, achieve compression by transmitting spatial metadata together with a reduced number of transport channels. However, recent studies have shown that IVAS performance degrades on reverberant content, particularly at low bitrates, a limitation that suggests its inability to accurately model room acoustics. In this paper, we propose a novel HOA coding scheme based on the explicit and blind estimation of the Relative Spatial Room Impulse Response (ReSRIR), using a beamformed version of the HOA signal as a reference signal. By exploiting the structure and sparsity of the estimated ReSRIR, we derive an efficient parametric representation for immersive audio coding. Experimental evaluations show that the proposed method achieves higher compression than IVAS in the single-transport-channel regime, while maintaining comparable to slightly better quality.2026-09-08T10:26:01ZSubmitted to International Networked Immersive Audio 2026 (satellite Event of IEEE IS2 2026)Nour BouayedAdrien LlaveJérôme DanielPascal Scalarthttp://arxiv.org/abs/2609.08429v1Semantic Refinement of Universal Audio Representations through Audio-Description Alignment2026-09-08T08:34:40ZUniversal audio representations must preserve acoustic detail while making high-level concepts accessible across speech, music, environmental sound, and downstream models of different capacities. We study semantic refinement of an acoustically pretrained encoder by adding audio-description alignment to a foundation of BEST-RQ, reconstruction, and CTC. We compare matched control, shuffled-description, and correctly paired trajectories to distinguish correct correspondence from an extra contrastive objective. Each endpoint is frozen and evaluated with a temporal-mean linear probe and a sequence-aware LLM readout, testing whether the refined information is directly accessible and remains useful to a stronger model. Across three paired seeds, correct alignment improves domain-balanced classification by 4.66 points with the linear probe and 2.59 points with the sequence-aware LLM, with positive changes in every domain. Correct pairing accounts for 87% of the linear-probe gain, while the LLM shows its clearest correspondence-specific benefit in captioning. Dense acoustic objectives provide complementary gains under both readouts. A separate 24-layer continuation remains competitive with leading public encoders under the shared evaluator, supporting the recipe beyond the controlled study.2026-09-08T08:34:40ZLejun MinJunyu DaiRuichen ZhengXinyue FanYang XiangHuaichen ZhangXingchen SongYufei ShiHan ZhaoXiangang Lihttp://arxiv.org/abs/2609.08422v1Beyond Localisation Accuracy: Sensorimotor Effects of HRTF Individualisation2026-09-08T08:30:17ZEveryday listening requires the brain to integrate cues from the body, environment, other senses, and movement, continuously translating auditory information into action. Yet HRTF individualisation is still commonly assessed through localisation accuracy, which may not fully capture its effects on this sensorimotor process. Here, we investigate whether these effects can instead be revealed through behaviour in a more ecologically valid listening task. We used an aurally guided visual search paradigm in which listeners located a visual target using a co-located virtual sound while moving freely, comparing individualised and non-individualised HRTFs under anechoic and reverberant conditions. Performance was assessed through response times and measures of movement organisation. In anechoic conditions, individualised HRTFs produced faster responses than non-individualised HRTFs, with an average reduction of approximately 200ms and the clearest benefit for front-back source locations. This advantage was expressed primarily in movement initiation, whereas overall movement extent was only weakly affected. Under reverberant conditions, HRTF-dependent differences disappeared. These results suggest that HRTF individualisation can influence how listeners plan and initiate orienting actions even when differences in conventional localisation outcomes are limited. Assessing sensorimotor behaviour alongside localisation performance may therefore provide a more sensitive and ecologically relevant account of the perceptual benefits of HRTF individualisation.2026-09-08T08:30:17Z18 pages, 5 figuresFulvio MissoniKatarina C. PooleTim Murray-BrowneAndrea CanessaLorenzo Picinalihttp://arxiv.org/abs/2608.23759v2The ISCSLP 2026 Real-World Audio-Visual Speech Enhancement Challenge2026-09-08T08:01:03ZAudio-visual speech enhancement (AVSE) uses visual-speech cues from a target speaker to recover that speaker's speech from noisy or overlapping speech. Many widely used protocols construct mixed signals from separately recorded audio sources and assume reliable video, leaving their performance under natural overlap and visual failure insufficiently characterized. The Real-World AVSE Challenge evaluates two related settings. Track~1 comprises two scenarios: real-world mixtures recorded with two speakers speaking simultaneously, without a corresponding clean reference signal, and synthetic remixes obtained by manually mixing the separately recorded speech of two speakers, with a clean reference signal available; Track~2 reuses audio but pairs it with a degraded target video and contains additional 3-m far-field recordings. The speakers in the development and test sets are disjoint. Evaluation metrics include clean-waveform fidelity, learned quality estimates, transcription accuracy, and speaker identification. In the remix task on the development set, the baseline model achieved an SI-SDR of $-4.069$~dB and an STOI of $0.388$ on Track~1, and an SI-SDR of $-2.851$~dB and an STOI of $0.470$ on Track~2. We release the AV-ConvTasNet checkpoints, the offline evaluator, and the official baseline results on the development and test sets.2026-08-24T18:52:23ZThe First Real-World Audio-Visual Speech Enhancement (AVSE) ChallengeKai LiWenze RenJunjie LiCheng YuPeijun YangChien-yu HuangHaibin WuSzu-Wei FuWen-Chin HuangHsin-Min WangXiaolin HuMing LiDeLiang WangYu Tsaohttp://arxiv.org/abs/2609.08147v1ConversationalVoice: Full-Duplex Speech Data from Real Conversations through Source-Faithful Reconstruction and Conversation-Grounded Expansion2026-09-08T02:28:43ZFull-duplex speech models require training data that preserves turn-taking, overlap, interruption, and backchannel behavior, yet these signals are entangled across speakers in noisy real-world recordings. We present Conversational Voice, a pipeline that converts real two-speaker excerpts into three complementary training-data artifacts. (1) Separation recovers speaker-specific tracks with stable speaker assignments, a canonical transcript, and naturally observed interaction timing. (2) Reconstruction generates speech in matched voices from a fixed source transcript, reconstructs the source turn order, pauses, and overlaps, and adds word-level alignment and delivery instructions. (3) Expansion generates new dialogue constrained by the source context, speakers, and observed interaction pattern. Automatic speaker-verification metrics remain strong across stages, with same-speaker similarity of 0.983-0.991 and positive discrimination margins of 0.199-0.209. Predicted speech quality (NISQA MOS) is 3.56 for separation, 4.41 for reconstruction, and 4.61 for expansion. A Gemini-based automatic evaluator assigns expansion mean scores of 4.94/5 for contextual coherence and 4.80/5 for dialogue naturalness. Expansion and reconstruction exhibit broadly similar interaction profiles; expansion's turn, overlap-event, backchannel, and interruption rates are 4.6%, 8.0%, 13.2%, and 16.0% lower, respectively. We evaluate data properties only; downstream gains in full-duplex model training remain for future work.2026-09-08T02:28:43Z12 pages, 3 figures, 1 table, preprintRichard Yucheng HeBaodong CaoChen XuYihang LiuTairan Chenhttp://arxiv.org/abs/2609.03203v2VoxReason: Listener-Free Evaluation of Source-Grounded Speech Planning Before Synthesis2026-09-07T22:11:56ZExpressive speech systems have to decide how an utterance is delivered before any waveform is rendered. In dialogue agents, narration, and role-conditioned TTS, that planning step sets affect, pitch, energy, rate, pause, emphasis, and stance, yet standard audio metrics rarely show whether those choices were actually licensed by the source record. This leaves a practical evaluation gap: a system may sound plausible while relying on a memorized script instead of the cue that governs delivery. VoxReason casts this pre-synthesis step as a listener-free task for source-grounded speech planning. Systems output a source-cited speaking plan, and a deterministic verifier checks citation legality, slot agreement, unsupported state, schema validity, and one-cue counterfactual locality. On 1,440 checked source-label cases, shortcut controls show why slot accuracy alone is unsafe: a key-lookup oracle reaches 1.000 plan-slot accuracy on seen keys, while an emotion prior still reaches 0.958 slot accuracy on source-key-disjoint cases without citing intensity or identity. In a separate 100-case learned source-key-disjoint comparison, a 7B locality SFT+CF repair improves plan-slot accuracy/locality from 0.684/0.141 to 0.919/1.000, and removing source records lowers citation-required grounded score by 0.488. The resulting benchmark isolates whether planned delivery is warranted by the record before waveform evaluation or listener studies are used.2026-09-02T22:35:02ZMengzhe Genghttp://arxiv.org/abs/2505.20961v2Efficient and Microphone-Fault-Tolerant 3D Sound Source Localization2026-09-07T21:36:57ZSound source localization (SSL) is a critical technology for determining the position of sound sources in complex environments. However, existing methods face challenges such as high computational costs and precise calibration requirements, limiting their deployment in dynamic or resource-constrained environments. This paper introduces a novel 3D SSL framework, which uses sparse cross-attention, pretraining, and adaptive signal coherence metrics, to achieve accurate and computationally efficient localization with fewer input microphones. The framework also supports operational microphones at unknown positions: their recordings remain available and are used to estimate both source and microphone positions. Preliminary experiments demonstrate its scalability for multi-source localization without requiring additional hardware. This work advances SSL by balancing the model's performance and efficiency and improving its robustness for real-world scenarios.2025-05-27T09:56:16ZAccepted by Interspeech 2025 ConferenceYiyuan YangShitong XuNiki TrigoniAndrew Markhamhttp://arxiv.org/abs/2607.06259v2Goodbye Equal Error Rate, Hello Local Information Disclosure: Evaluating Voice Anonymisation against 1-to-N Linkage Threats2026-09-07T15:59:24ZVoice anonymisation aims to protect speaker identity. Currently, its empirical privacy evaluation heavily relies on the Equal Error Rate (EER). Originally designed for biometric verification, EER aggregates scores globally, implicitly assuming an attacker is only trying to verify if two specific voice samples match (a 1-to-1 comparison). This introduces a threat model mismatch with real-world database linkage attacks, where an attacker searches across a fixed set of N enrolled identities (a 1-to-N closed-set search), allowing global averages to obscure localised privacy failures. While recent 1-to-N metrics address this aggregation issue, they abstract away the magnitude of the biometric evidence. In this paper, we propose a modular, information-theoretic evaluation framework explicitly designed for the 1-to-N linkage threat model. Within this framework, our core metric, Local Information Disclosure (LID), provides a principled estimate of the identity information disclosed by a single trial utterance in bits by mapping raw similarity scores to an estimated posterior distribution over the enrolled identities. Demonstrating this framework on VoicePrivacy 2024 Challenge systems reveals how global metrics can obscure privacy vulnerabilities. Even for top-performing systems where standard evaluations report near-perfect EERs (48%), our metric exposes that attackers gain a statistical advantage in at least 63% of trials with maximum observed disclosures reaching 1 bit (effectively doubling the attacker's confidence in the target speaker after observing a single trial utterance). We conclude that shifting toward explainable metrics and appropriate threat models is a practical step toward identifying worst-case vulnerabilities and aligning with strict privacy regulations.2026-07-07T13:25:54ZAccepted for publication at the Symposium on Security and Privacy in Speech Communication (SPSC 2026)Dāvis ŠternsKonstantinos DrossosNatasha FernandesTom BäckströmCatuscia Palamidessihttp://arxiv.org/abs/2602.19574v2CTC-TTS: LLM-based dual-streaming text-to-speech with CTC alignment2026-09-07T14:38:58ZLarge-language-model (LLM)-based text-to-speech (TTS) systems can generate natural speech, but most are not designed for low-latency dual-streaming synthesis. High-quality dual-streaming TTS depends on accurate text--speech alignment and well-designed training sequences that balance synthesis quality and latency. Prior work often relies on GMM-HMM based forced-alignment toolkits (e.g., MFA), which are pipeline-heavy and less flexible than neural aligners; fixed-ratio interleaving of text and speech tokens struggles to capture text--speech alignment regularities. We propose CTC-TTS, which replaces MFA with a CTC based aligner and introduces a bi-word based interleaving strategy. Two variants are designed: CTC-TTS-L (token concatenation along the sequence length) for higher quality and CTC-TTS-F (embedding stacking along the feature dimension) for lower latency. Experiments show that CTC-TTS outperforms fixed-ratio interleaving and MFA-based baselines on streaming synthesis and zero-shot tasks.2026-02-23T07:44:14ZFix two typos in Figure 2 in INTERSPEECH 2026 versionHanwen LiuSaierdaer YusuyinHao HuangZhijian Ouhttp://arxiv.org/abs/2609.07399v1Open-Set Vessel Re-Identification from Underwater Ship-Radiated Noise with a Raw-Waveform Selective-Kernel Acoustic Neural Network (SKANN) and a Cross-Passage Evaluation Protocol2026-09-07T12:13:43ZUnderwater acoustic target recognition has converged on closed-set classification by vessel type, a task that does not answer whether a monitoring system has heard this hull before. We formalise open-set, cross-passage vessel re-identification on public hydrophone data and specify a protocol that removes the two easiest routes to a high score: hull-disjoint splits keyed to MMSI/IMO, galleries and queries from disjoint passages of each hull, source-pure galleries, and an audio-adjudicated transit-deduplication gate. We describe SKANN, a raw-waveform encoder whose front end is a four-scale bank of learned filters fused by selective-kernel attention, trained with an angular-margin objective and an augmentation regime that perturbs recording chain, ambient noise and multipath while preserving the narrowband lines that carry identity. On a 40-hull IARA gallery (96 queries, 98 passage candidates), cross-passage rank-1 is 0.25 for the embedding and 0.26 for an automated narrowband-tonal comparator; the two are statistically indistinguishable at the top of the ranking, the embedding orders the rest of the list more reliably (AUC 0.82 vs 0.76), and their score fusion reaches rank-1 0.35 -- the only contrast that attains nominal significance, presented as evidence of partial complementarity, not as a recommendation. Transit deduplication alone removes a 16-21 point apparent rank-1 advantage, larger than any between-method difference. Two further findings delimit what public data can support: ShipsEar cannot separate hull identity from recording channel under an identity protocol, and cross-network fine-tuning helps vessels seen during fine-tuning but is a null result on unseen ones. The results support analyst triage over a ranked shortlist, not identification. Checkpoint, validation embeddings, transit map and per-query outputs are released under CC-BY-4.0 (doi:10.5281/zenodo.22160138).2026-09-07T12:13:43Z15 pages, 5 figures. An Indian provisional patent application (202611107132, filed 6 September 2026) covers aspects of the method described hereSunil Tyagi