https://arxiv.org/api/JJT1eOaCXbMxVGji5wkJsL48qyI 2026-09-11T17:47:08Z 31591 0 15 http://arxiv.org/abs/2606.23835v3 ABACUS: Adapting Unified Foundation Model for Bridging Image Count Understanding and Generation 2026-09-10T17:09:47Z We present ABACUS, a unified vision-language model that jointly addresses object counting, crowd counting, referring-expression counting, and count-faithful image generation within a single 3B-parameter model. ABACUS introduces three contributions: density-aware adaptive zooming paired with an objectness map from multi-head self-attention decomposition to spatially ground count predictions; a boundary-aware count policy trained via GRPO with nested local, boundary, and global rewards to eliminate over- and undercounting at crop boundaries; and a cycle-consistent GRPO strategy in which the frozen understanding branch scores generated candidates on count-deviation and aesthetic quality, closing the understanding-generation synergy gap without any external critic or annotation. ABACUS achieves state-of-the-art results across seven benchmarks spanning object counting (FSC-147, CARPK), crowd counting (ShanghaiTech A/B), referring-expression counting (REC-8K), count-faithful generation (CoCoCount, T2I-CompBench, GenEval), and count reasoning (CountQA), surpassing both task-specific specialists and larger generalist models. Project page is at https://mondalanindya.github.io/pages/ABACUS. 2026-06-22T18:16:31Z ACM Transactions on Graphics/ SIGGRAPH ASIA 2026, webpage: https://mondalanindya.github.io/pages/ABACUS Anindya Mondal Sauradip Nag Anjan Dutta http://arxiv.org/abs/2609.11506v1 UBone3D: Physics-Rectified Conditional Flow Matching for Anatomical 3D Shape Completion from Ultrasound 2026-09-10T13:13:29Z Three-dimensional ultrasound (US) is a safe, radiation-free complementary modality to CT and X-rays for longitudinal monitoring, yet its segmentation-derived partial point clouds are extremely artifact-laden. Consequently, it is challenging to recover a clean and complete anatomical structure from such US point clouds. In this paper, we present UBone3D, a novel framework based on physics-rectified conditional flow matching (CFM) that performs point cloud completion directly from partial US observations. UBone3D models deterministic physics artifacts (e.g., surface thickening, streaking, dropouts) via a simulated physics proxy, and introduces test-time physics rectification to steer the shape completion. At inference, the completion is jointly steered by two decoupled forces: (1) anatomical plausibility enforced by a CT-trained generative shape prior, BoneFM, and (2) physics consistency enforced by USimNet in the ultrasound formation space. Extensive experiments on simulated and in-vivo data demonstrate significant improvements in reconstruction accuracy and anatomical fidelity over existing baselines. 2026-09-10T13:13:29Z Accepted to ECCV 2026. Camera-ready Author Version Weiying Chen Yuchong Gao Siyuan Li Marek Reformat Rui Zheng Edmond Lou http://arxiv.org/abs/2602.10420v4 Prediction--Loss Alignment for Sampler--Robust Flow Matching Training 2026-09-10T09:51:53Z Recent work has popularized a practical recipe in diffusion and flow matching: predict the clean signal $x$, convert it to a velocity, and train through a velocity-space loss. The conversion contains a singular endpoint amplification and therefore appears prone to unstable optimization, yet recent systems obtain strong empirical results with this recipe. We investigate this tension through the integrability of the pre-optimizer stochastic-gradient second moment. Under stated initialization conditions, the moment diverges under Uniform sampling; boundary-suppressing sampling can restore integrability under an additional upper-growth condition. We then show that prediction--loss alignment eliminates this conversion-induced source of non-integrability. Under a uniform moment bound, alignment yields a finite second moment for every timestep density, including Uniform sampling. Controlled experiments across continuous and binary settings reproduce the predicted sampler-dependent instability and show that aligned objectives remain trainable across the tested samplers. These results reconcile pointwise amplification with sampler-dependent empirical success and support alignment as a principled route to more robust flow-matching training. 2026-02-11T02:02:30Z 24 pages, 9 tables, 10 figures. This version corrects errors in the experimental evaluation and revises the affected results and conclusions. It supersedes earlier versions; readers should refer to the corrected results presented here Jiadong Hong Lei Liu Xinyu Bian Wenjie Wang Zhaoyang Zhang http://arxiv.org/abs/2609.11310v1 Your Model Already Knows Don't Teach It, Learn to Ask It: Soft Prompting for Few-Shot Adaptation of Vision-Language Models 2026-09-10T09:38:11Z We address few-shot object detection with vision-language models (VLMs) in out-of-domain settings such as aerial, industrial, and medical imagery, using only ten annotated images for supervision. Existing adaptation methods are discrete prompt optimization and LoRA fine-tuning. We revisit a third option: soft prompting, where a small number of continuous prompt tokens are optimized while the pretrained backbone remains frozen. We identify two key design choices. First, placing prompt tokens at the cross-modal boundary between visual and text tokens outperforms other placements (10.0 vs. 8.4 mAP). Second, initializing prompts from the empty space token outperforms semantic and random initialization. With these choices, one to three learned tokens (7,168 parameters on average) match the best LoRA configuration on Roboflow20-VL (14.2 mAP, 10-shot) while training over 20,000x fewer parameters. Soft prompting remains harder to optimize, exhibiting higher variance across random seeds. Unlike LoRA, however, it causes no forgetting: the LoRA rank matching our accuracy reduces NaturalBench VQA accuracy by 35% relative, rising to 56% at the largest rank, whereas soft prompting leaves pretrained performance unchanged. The learned tokens behave like prompts rather than weights. They transfer to a newer model without retraining (+0.8 mAP on Qwen3.5-9B) and can be verbalized into readable prompts competitive with prompt-search methods (matching DetPO and outperforming GEPA). The approach also extends beyond detection. On RoboCasa manipulation tasks, the frozen $π_{0.5}$ vision-language-action policy benefits from soft prompting, matching the LoRA baseline on two of three tasks when tokens are placed at the gradient bottleneck. These results suggest modern VLMs already encode much of what is needed for specialized domains; the challenge is learning how to ask. 2026-09-10T09:38:11Z Gautam Rajendrakumar Gare Siyi Li Hewei Wang Cesar Daniel Hernandez Wei Zhao Wolfgang M. Pauli John Galeotti Deva Ramanan http://arxiv.org/abs/2609.10988v1 Exponential Pixelating Integral transform with dual fractal features for enhanced chest X-ray abnormality detection 2026-09-10T02:10:23Z The heightened prevalence of respiratory disorders, particularly exacerbated by a significant upswing in fatalities due to the novel coronavirus, underscores the critical need for early detection and timely intervention. This imperative is paramount, possessing the potential to profoundly impact and safeguard numerous lives. Medically, chest radiography stands out as an essential and economically viable medical imaging approach for diagnosing and assessing the severity of diverse Respiratory Disorders. However, their detection in Chest X-Rays is a cumbersome task even for well-trained radiologists owing to low contrast issues, overlapping of the tissue structures, subjective variability, and the presence of noise. To address these issues, a novel analytical model termed Exponential Pixelating Integral is introduced for the automatic detection of infections in Chest X-Rays in this work. Initially, the presented Exponential Pixelating Integral enhances the pixel intensities to overcome the low-contrast issues that are then polar-transformed followed by their representation using the locally invariant Mandelbrot and Julia fractal geometries for effective distinction of structural features. The collated features labeled Exponential Pixelating Integral with dually characterized fractal features are then classified by the non-parametric multivariate adaptive regression splines to establish an ensemble model between each pair of classes for effective diagnosis of diverse diseases. Rigorous analysis of the proposed classification framework on large medical benchmarked datasets showcases its superiority over its peers by registering a higher classification accuracy and F1 scores ranging from 98.46 to 99.45% and 96.53-98.10% respectively, making it a precise and interpretable automated system for diagnosing respiratory disorders. 2026-09-10T02:10:23Z Preprint of the article published in the ELSEVIER journal Computers in Biology and Medicine (CIBM), Vol. 182, November 2024. Final version available at DOI: https://doi.org/10.1016/j.compbiomed.2024.109093 Computers in Biology and Medicine (CIBM), Vol. 182, November 2024, 109093 Naveenraj Kamalakannan Sri Ram Macharla M Kanimozhi M S Sudhakar 10.1016/j.compbiomed.2024.109093 http://arxiv.org/abs/2609.10914v1 Seamless Whole Slide Label-Free Virtual Staining 2026-09-09T23:59:55Z Label-free virtual staining offers a compelling, non-destructive alternative to standard histopathology; however, its clinical adoption is hindered by the computational bottlenecks inherent to processing gigapixel Whole Slide Images (WSIs). Current deep learning approaches require patch-based inference to avoid memory constraints, which disrupts global tissue continuity and introduces tiling artifacts--displaying visible seams and color shifts. To address this, we introduce the Consistency Memory Bank (COMB), a novel label-free virtual staining framework that enforces spatial and channel consistency across tiles without memory bottlenecks. COMB decouples context storage from computation, utilizing a dynamic retrieval mechanism to fetch feature representations from adjacent tiles. This enables a retrieval-based context integration strategy that adopts local padding to resolve spatial discontinuities and neighbor-aware channel attention to stabilize statistical drift. Further optimized with a sliding window schedule to ensure minimal memory overhead, our method demonstrates superior performance over state-of-the-art baselines, achieving significant improvements in both perceptual fidelity and tiling consistency, while suggesting its downstream utility in tumor segmentation. Code is available at https://github.com/dou0000/COMB. 2026-09-09T23:59:55Z Accepted to MICCAI 2026 Dou Hoon Kwark Kianoush Falahkheirkhah Ji-hun Oh Shirui Luo Volodymyr Kindratenko Rohit Bhargava http://arxiv.org/abs/2609.10858v1 Annotating anatomy and pathology in the National Lung Screening Trial computed tomography images 2026-09-09T21:52:36Z Large-scale public medical imaging datasets contribute critically to translational research. When accompanied by rich clinical and multi-omics data, they can stimulate exploratory research and enable secondary analyses. Expert annotations of such imaging collections can support the development of new image analysis tools. Continuous enrichment of images with image-derived data makes them more usable for researchers without expertise in image analysis or access to large-scale computational resources. The National Lung Screening Trial (NLST) released a rich longitudinal dataset that includes Computed Tomography (CT) images for over 26,000 patients. We introduce three Digital Imaging and Communications in Medicine (DICOM) formatted datasets, complementing NLST CT images, shared as analysis results in the National Cancer Institute Imaging Data Commons (IDC). Two of those (IDC NLSTSeg and IDC NLSTSybil) contain DICOM-harmonized annotations and extracted measurements (for 581 and 601 NLST patients, respectively) shared earlier using research formats (Sybil and NLSTseg). The third one (TotalSegmentator-CT-Segmentations) contains volumetric segmentations generated using TotalSegmentator and radiomics features for each segment for 26,194 NLST patients. 2026-09-09T21:52:36Z Deepa Krishnaswamy Vamsi Thiriveedhi Suraj Pai David Clunie Igor Octaviano Christopher P. Bridge Steve Pieper Ron Kikinis Andrey Fedorov http://arxiv.org/abs/2609.10825v1 Scale-Aware 3D Deep Learning for Robust Brain Metastasis Detection in Multimodal MRI 2026-09-09T20:53:10Z Detecting brain metastases in magnetic resonance imaging (MRI) remains challenging because lesions vary widely in size and appearance, with very small metastases occupying only a minute fraction of a three-dimensional input. We investigate whether combining different spatial fields of view (FOVs) improves lesion detection in multimodal MRI and present a scale-aware 3D deep-learning framework. The method uses independently trained $96^3$ and $64^3$ 3D U-Nets whose whole-volume probability maps are combined by weighted late fusion. This design allows us to study the effect of spatial context separately from image resolution and modality choice. On a 97-patient development cohort, cross-FOV fusion improved lesion-level precision and F1 while substantially reducing false positives relative to the individual models. A same-FOV ensemble control showed that these gains were not explained solely by averaging independently trained networks, supporting a contribution from complementary spatial context. An exploratory cross-FOV agreement filter reduced false positives but did not improve overall F1. These results support cross-FOV probability fusion as a simple and computationally practical strategy for improving the precision-false-positive trade-off in 3D brain-metastasis detection. 2026-09-09T20:53:10Z 12 pages, 3 figures, 3 tables Proceedings of the International Conference on Medical Image Computing and Computer-Assisted Intervention (MICCAI 2026), Strasbourg, Sep 2026 Sylvain Jaume Hongming Wang Simon K. Warfield http://arxiv.org/abs/2503.17970v2 PathoHR: Breast Cancer Survival Prediction on High-Resolution Pathological Images 2026-09-09T19:18:39Z Breast cancer survival prediction in computational pathology presents a remarkable challenge due to tumor heterogeneity. For instance, different regions of the same tumor in the pathology image can show distinct morphological and molecular characteristics. This makes it difficult to extract representative features from whole slide images (WSIs) that truly reflect the tumor's aggressive potential and likely survival outcomes. In this paper, we present PathoHR, a novel pipeline for accurate breast cancer survival prediction that enhances any size of pathological images to enable more effective feature learning. Our approach entails (1) the incorporation of a plug-and-play high-resolution Vision Transformer (ViT) to enhance patch-wise WSI representation, enabling more detailed and comprehensive feature extraction, (2) the systematic evaluation of multiple advanced similarity metrics for comparing WSI-extracted features, optimizing the representation learning process to better capture tumor characteristics, (3) the demonstration that smaller image patches enhanced follow the proposed pipeline can achieve equivalent or superior prediction accuracy compared to raw larger patches, while significantly reducing computational overhead. Experimental findings valid that PathoHR provides the potential way of integrating enhanced image resolution with optimized feature learning to advance computational pathology, offering a promising direction for more accurate and efficient breast cancer survival prediction. Code will be available at https://github.com/AIGeeksGroup/PathoHR. 2025-03-23T07:37:24Z First author comment: We are withdrawing this manuscript to address several unresolved limitations and complete the necessary internal review and approval procedures before further dissemination (confirmed by the corresponding author) Yang Luo Shiru Wang Jun Liu Jiaxuan Xiao Rundong Xue Zeyu Zhang Hao Zhang Yu Lu Yang Zhao Yutong Xie http://arxiv.org/abs/2412.19990v3 SegKAN: High-Resolution Medical Image Segmentation with Long-Distance Dependencies 2026-09-09T19:17:15Z Hepatic vessels in computed tomography scans often suffer from image fragmentation and noise interference, making it difficult to maintain vessel integrity and posing significant challenges for vessel segmentation. To address this issue, we propose an innovative model: SegKAN. First, we improve the conventional embedding module by adopting a novel convolutional network structure for image embedding, which smooths out image noise and prevents issues such as gradient explosion in subsequent stages. Next, we transform the spatial relationships between Patch blocks into temporal relationships to solve the problem of capturing positional relationships between Patch blocks in traditional Vision Transformer models. We conducted experiments on a Hepatic vessel dataset, and compared to the existing state-of-the-art model, the Dice score improved by 1.78%. These results demonstrate that the proposed new structure effectively enhances the segmentation performance of high-resolution extended objects. Code will be available at https://github.com/goblin327/SegKAN 2024-12-28T03:27:21Z First author comment: Due to unresolved limitations in this study, we believe that the conclusions are not yet sufficiently supported. We therefore withdraw this preprint and advise readers not to cite it (confirmed by the corresponding author) Shengbo Tan Rundong Xue Shipeng Luo Zeyu Zhang Xinran Wang Lei Zhang Daji Ergu Zhang Yi Yang Zhao Ying Cai http://arxiv.org/abs/2609.10733v1 XPos3R: Cross-Modal Transformer for Intraoperative 2D/3D Registration 2026-09-09T18:29:26Z Intraoperative 2D/3D registration, which aligns live X-ray images with preoperative volumes, is essential for image-guided interventions. Previous regression-based methods suffer from limited generalization, thus requiring time-consuming patient-specific retraining. Inspired by recent geometry foundation models such as DUSt3R, we propose XPos3R, a generalizable pose regression method that eliminates preoperative preparation. Unlike existing geometry models designed for homogeneous inputs, XPos3R extends this paradigm to multi-modal inputs, namely 2D X-rays and 3D volumes. Specifically, we introduce an asymmetric encoder-decoder architecture that improves cross-modal feature alignment while maintaining computational efficiency. To scale training under limited medical data, we adopt an anatomy-specific data curation strategy and construct million-scale synthetic datasets. Evaluated on real-world benchmarks, a single pretrained XPos3R surpasses patient-specific methods in both accuracy and robustness. With test-time optimization completed in seconds, it further reduces the 3D error to <4 mm and the reprojection error to <1 mm. The strong generalization, accuracy, and efficiency of XPos3R highlight its clinical potential, while its asymmetric framework may inspire broader cross-modal vision geometry tasks. 2026-09-09T18:29:26Z Accepted to ECCV 2026 Shiyan Su Ruyi Zha Hongdong Li Xuelian Cheng Zongyuan Ge http://arxiv.org/abs/2608.31048v2 OmniRAS: Standardizing Foundation Model Training and Evaluation in Robot-Assisted Surgery 2026-09-09T15:42:25Z Few foundation models exist for robot-assisted surgery, partly because large robotic-surgery video corpora are difficult to assemble and existing models are evaluated mostly on laparoscopic benchmarks. Further, most existing models are evaluated on a small set of public benchmarks, mostly focused on laparoscopic surgery. We present OmniRAS, a family of 1B- and 2B-parameter V-JEPA-2.1 encoders for robot-assisted surgery, and detail their training. First, we release two densely annotated robotic-cholecystectomy datasets: OmniRAS-PR and a multi-label YT-Chole tool-verb-target task, the first triplet-style annotation for robotic cholecystectomy, together with splits, probe protocols, and an inter-rater study validating the shared phase ontology. Second, we document continued pretraining at up to 256 compute nodes with global batch 6,144 over 19 sources totaling approximately 2,650 hours of surgical video, 51% robotic, and analyze compute and data composition. Third, we evaluate against raw V-JEPA-2.1 and specialized surgical models on six tasks spanning triplet, phase, and step recognition, action segmentation, and detection, under frozen-encoder and final-four-block fine-tuning regimes. Across three seeds, this yields 254 downstream runs, including 109 with partial backbone fine-tuning. The best OmniRAS models achieve the strongest adapted results across all task families, while frozen differences are smaller. 2026-08-31T16:24:26Z 18 pages, 11 figures, 4 tables. Under review at IEEE Transactions on Robotics (T-RO) Leonardo Borgioli Neil Getty Wenli Xiu Jessica Cassiani Alvaro Ducas Carlos Agustin Orda Hira Waris Fangfang Xia Rick Stevens Pier Cristoforo Giulianotti Milos Zefran http://arxiv.org/abs/2505.03380v2 Synergistic Vision-Language Reinforcement Enables Scalable On-Demand Analysis across Diverse Clinical Tasks 2026-09-09T15:03:16Z Accurate delineation of tumors and surrounding organs-at-risk is essential for radiotherapy, surgery and treatment response assessment, yet remains time-consuming and expertise-intensive. Existing artificial intelligence systems often require manual spatial prompts or task-specific retraining, while generic class labels provide limited semantic grounding for heterogeneous disease targets. Here we present SyRe, a promptable segmentation foundation model based on Synergistic vision-language Reinforcement. SyRe strengthens bidirectional interaction between visual and linguistic representations to improve semantically grounded spatial understanding. To support large-scale training, we introduce the Color Region Description strategy and construct SyReData, comprising 20 million image-mask-description triplets across 9 modalities and 229 segmentation tasks. Training with diversified prompt forms further enables open-ended prompting, invalid-prompt rejection and flexible switching between single- and multi-target analysis. SyRe achieves accurate text-prompted segmentation across diverse clinical scenarios, with particularly strong performance on disease-related targets. Across 28 unseen external datasets, including 20 cancer types and multinational in-house cohorts, SyRe generalizes robustly under real-world distribution shifts. SyRe-generated masks also preserve clinically relevant quantitative information in pathology and yield radiomics features that stratify survival and improve prognostic modeling across five retrospective CT and MRI tumor cohorts. Finally, clinician-in-the-loop refinement enables efficient case-level correction when greater precision is required. These results establish SyRe as a generalizable foundation for scalable quantitative oncology and clinician-guided segmentation refinement. 2025-05-06T10:00:08Z Haonan Wang Jiaji Mao Lehan Wang Qixiang Zhang Marawan Elbatel Yi Qin Huijun Hu Baoxun Li Wenhui Deng Weifeng Qin Hongrui Li Jialin Liang Jun Shen Xiaomeng Li http://arxiv.org/abs/2609.10112v1 Storage-Scalable Progressive Semantic Communication via Knowledge-Base Reuse 2026-09-09T12:46:19Z Existing knowledge-base-assisted semantic communication schemes commonly adopt either single knowledge-base quantization (SKBQ) or multi-knowledge-base residual quantization (MKBQ). SKBQ incurs limited storage overhead but has restricted quantization capacity, whereas MKBQ supports progressive refinement by assigning an independent knowledge base (KB) to each stage, causing the KB storage to grow linearly with the transmission depth. To address this problem, we propose storage-scalable knowledge-base reuse quantization (SSKBQ), which reuses a compact set of KBs across multiple residual refinement stages and thereby decouples the number of transmission stages from the number of maintained KBs. A stage-aware residual supervision mechanism is further introduced to regularize intermediate quantized representations and encourage progressive refinement. Experimental results demonstrate that KB reuse provides an effective solution to the storage scalability problem while maintaining competitive progressive reconstruction performance. 2026-09-09T12:46:19Z Submitted manuscript Heng Zhu Ye Liu Kun Zhu Feifei Song http://arxiv.org/abs/2608.18915v2 Simple, Safe, and Overlooked: Reclaiming Sustainable Domain Generalization with Statistical Color Matching 2026-09-09T12:22:42Z Hardware shifts, color variations, and changing patient characteristics between development and deployment routinely break trained medical image classifiers. Existing remedies fall short: standard color jittering provides insufficient diversity, while deep generative style transfer algorithms hallucinate features, destroy clinically relevant structures, and waste massive compute resources. To address this, we revisit classical statistical color matching and repurpose it as Colorist, a highly efficient data augmentation strategy that applies global mean-standard deviation matching directly in the RGB color space. We demonstrate that this training-free, fully interpretable approach safely generates structurally intact domain variations, outperforming deep generative models in structural fidelity and color alignment. Across out-of-distribution histopathology, peripheral blood, dermatology, and retinal datasets, it improves balanced accuracy by up to +9% over state-of-the-art domain generalization regularizers and by +13% over an unaugmented baseline. Moreover, by avoiding neural networks in the augmentation loop, Colorist preserves anatomical structure, minimizes carbon footprint, and integrates seamlessly into standard dataloaders. Together, these findings establish statistical matching as a safe, interpretable, yet overlooked alternative to deep architectures for clinical robustness. Source code is available at https://github.com/sdoerrich97/colorist. 2026-08-19T13:47:51Z Accepted to DEMI @ MICCAI 2026 (4th Workshop in Data Engineering in Medical Imaging) Sebastian Doerrich Francesco Di Salvo Shyam Nandan Rai Marco Lents Christian Ledig