https://arxiv.org/api/JJT1eOaCXbMxVGji5wkJsL48qyI2026-09-11T17:47:08Z31591015http://arxiv.org/abs/2606.23835v3ABACUS: Adapting Unified Foundation Model for Bridging Image Count Understanding and Generation2026-09-10T17:09:47ZWe present ABACUS, a unified vision-language model that jointly addresses object counting, crowd counting, referring-expression counting, and count-faithful image generation within a single 3B-parameter model. ABACUS introduces three contributions: density-aware adaptive zooming paired with an objectness map from multi-head self-attention decomposition to spatially ground count predictions; a boundary-aware count policy trained via GRPO with nested local, boundary, and global rewards to eliminate over- and undercounting at crop boundaries; and a cycle-consistent GRPO strategy in which the frozen understanding branch scores generated candidates on count-deviation and aesthetic quality, closing the understanding-generation synergy gap without any external critic or annotation. ABACUS achieves state-of-the-art results across seven benchmarks spanning object counting (FSC-147, CARPK), crowd counting (ShanghaiTech A/B), referring-expression counting (REC-8K), count-faithful generation (CoCoCount, T2I-CompBench, GenEval), and count reasoning (CountQA), surpassing both task-specific specialists and larger generalist models. Project page is at https://mondalanindya.github.io/pages/ABACUS.2026-06-22T18:16:31ZACM Transactions on Graphics/ SIGGRAPH ASIA 2026, webpage: https://mondalanindya.github.io/pages/ABACUSAnindya MondalSauradip NagAnjan Duttahttp://arxiv.org/abs/2609.11506v1UBone3D: Physics-Rectified Conditional Flow Matching for Anatomical 3D Shape Completion from Ultrasound2026-09-10T13:13:29ZThree-dimensional ultrasound (US) is a safe, radiation-free complementary modality to CT and X-rays for longitudinal monitoring, yet its segmentation-derived partial point clouds are extremely artifact-laden. Consequently, it is challenging to recover a clean and complete anatomical structure from such US point clouds. In this paper, we present UBone3D, a novel framework based on physics-rectified conditional flow matching (CFM) that performs point cloud completion directly from partial US observations. UBone3D models deterministic physics artifacts (e.g., surface thickening, streaking, dropouts) via a simulated physics proxy, and introduces test-time physics rectification to steer the shape completion. At inference, the completion is jointly steered by two decoupled forces: (1) anatomical plausibility enforced by a CT-trained generative shape prior, BoneFM, and (2) physics consistency enforced by USimNet in the ultrasound formation space. Extensive experiments on simulated and in-vivo data demonstrate significant improvements in reconstruction accuracy and anatomical fidelity over existing baselines.2026-09-10T13:13:29ZAccepted to ECCV 2026. Camera-ready Author VersionWeiying ChenYuchong GaoSiyuan LiMarek ReformatRui ZhengEdmond Louhttp://arxiv.org/abs/2602.10420v4Prediction--Loss Alignment for Sampler--Robust Flow Matching Training2026-09-10T09:51:53ZRecent work has popularized a practical recipe in diffusion and flow matching: predict the clean signal $x$, convert it to a velocity, and train through a velocity-space loss. The conversion contains a singular endpoint amplification and therefore appears prone to unstable optimization, yet recent systems obtain strong empirical results with this recipe. We investigate this tension through the integrability of the pre-optimizer stochastic-gradient second moment. Under stated initialization conditions, the moment diverges under Uniform sampling; boundary-suppressing sampling can restore integrability under an additional upper-growth condition. We then show that prediction--loss alignment eliminates this conversion-induced source of non-integrability. Under a uniform moment bound, alignment yields a finite second moment for every timestep density, including Uniform sampling. Controlled experiments across continuous and binary settings reproduce the predicted sampler-dependent instability and show that aligned objectives remain trainable across the tested samplers. These results reconcile pointwise amplification with sampler-dependent empirical success and support alignment as a principled route to more robust flow-matching training.2026-02-11T02:02:30Z24 pages, 9 tables, 10 figures. This version corrects errors in the experimental evaluation and revises the affected results and conclusions. It supersedes earlier versions; readers should refer to the corrected results presented hereJiadong HongLei LiuXinyu BianWenjie WangZhaoyang Zhanghttp://arxiv.org/abs/2609.11310v1Your Model Already Knows Don't Teach It, Learn to Ask It: Soft Prompting for Few-Shot Adaptation of Vision-Language Models2026-09-10T09:38:11ZWe address few-shot object detection with vision-language models (VLMs) in out-of-domain settings such as aerial, industrial, and medical imagery, using only ten annotated images for supervision. Existing adaptation methods are discrete prompt optimization and LoRA fine-tuning. We revisit a third option: soft prompting, where a small number of continuous prompt tokens are optimized while the pretrained backbone remains frozen.
We identify two key design choices. First, placing prompt tokens at the cross-modal boundary between visual and text tokens outperforms other placements (10.0 vs. 8.4 mAP). Second, initializing prompts from the empty space token outperforms semantic and random initialization.
With these choices, one to three learned tokens (7,168 parameters on average) match the best LoRA configuration on Roboflow20-VL (14.2 mAP, 10-shot) while training over 20,000x fewer parameters. Soft prompting remains harder to optimize, exhibiting higher variance across random seeds. Unlike LoRA, however, it causes no forgetting: the LoRA rank matching our accuracy reduces NaturalBench VQA accuracy by 35% relative, rising to 56% at the largest rank, whereas soft prompting leaves pretrained performance unchanged.
The learned tokens behave like prompts rather than weights. They transfer to a newer model without retraining (+0.8 mAP on Qwen3.5-9B) and can be verbalized into readable prompts competitive with prompt-search methods (matching DetPO and outperforming GEPA).
The approach also extends beyond detection. On RoboCasa manipulation tasks, the frozen $π_{0.5}$ vision-language-action policy benefits from soft prompting, matching the LoRA baseline on two of three tasks when tokens are placed at the gradient bottleneck. These results suggest modern VLMs already encode much of what is needed for specialized domains; the challenge is learning how to ask.2026-09-10T09:38:11ZGautam Rajendrakumar GareSiyi LiHewei WangCesar Daniel HernandezWei ZhaoWolfgang M. PauliJohn GaleottiDeva Ramananhttp://arxiv.org/abs/2609.10988v1Exponential Pixelating Integral transform with dual fractal features for enhanced chest X-ray abnormality detection2026-09-10T02:10:23ZThe heightened prevalence of respiratory disorders, particularly exacerbated by a significant upswing in fatalities due to the novel coronavirus, underscores the critical need for early detection and timely intervention. This imperative is paramount, possessing the potential to profoundly impact and safeguard numerous lives. Medically, chest radiography stands out as an essential and economically viable medical imaging approach for diagnosing and assessing the severity of diverse Respiratory Disorders. However, their detection in Chest X-Rays is a cumbersome task even for well-trained radiologists owing to low contrast issues, overlapping of the tissue structures, subjective variability, and the presence of noise. To address these issues, a novel analytical model termed Exponential Pixelating Integral is introduced for the automatic detection of infections in Chest X-Rays in this work. Initially, the presented Exponential Pixelating Integral enhances the pixel intensities to overcome the low-contrast issues that are then polar-transformed followed by their representation using the locally invariant Mandelbrot and Julia fractal geometries for effective distinction of structural features. The collated features labeled Exponential Pixelating Integral with dually characterized fractal features are then classified by the non-parametric multivariate adaptive regression splines to establish an ensemble model between each pair of classes for effective diagnosis of diverse diseases. Rigorous analysis of the proposed classification framework on large medical benchmarked datasets showcases its superiority over its peers by registering a higher classification accuracy and F1 scores ranging from 98.46 to 99.45% and 96.53-98.10% respectively, making it a precise and interpretable automated system for diagnosing respiratory disorders.2026-09-10T02:10:23ZPreprint of the article published in the ELSEVIER journal Computers in Biology and Medicine (CIBM), Vol. 182, November 2024. Final version available at DOI: https://doi.org/10.1016/j.compbiomed.2024.109093Computers in Biology and Medicine (CIBM), Vol. 182, November 2024, 109093Naveenraj KamalakannanSri Ram MacharlaM KanimozhiM S Sudhakar10.1016/j.compbiomed.2024.109093http://arxiv.org/abs/2609.10914v1Seamless Whole Slide Label-Free Virtual Staining2026-09-09T23:59:55ZLabel-free virtual staining offers a compelling, non-destructive alternative to standard histopathology; however, its clinical adoption is hindered by the computational bottlenecks inherent to processing gigapixel Whole Slide Images (WSIs). Current deep learning approaches require patch-based inference to avoid memory constraints, which disrupts global tissue continuity and introduces tiling artifacts--displaying visible seams and color shifts. To address this, we introduce the Consistency Memory Bank (COMB), a novel label-free virtual staining framework that enforces spatial and channel consistency across tiles without memory bottlenecks. COMB decouples context storage from computation, utilizing a dynamic retrieval mechanism to fetch feature representations from adjacent tiles. This enables a retrieval-based context integration strategy that adopts local padding to resolve spatial discontinuities and neighbor-aware channel attention to stabilize statistical drift. Further optimized with a sliding window schedule to ensure minimal memory overhead, our method demonstrates superior performance over state-of-the-art baselines, achieving significant improvements in both perceptual fidelity and tiling consistency, while suggesting its downstream utility in tumor segmentation. Code is available at https://github.com/dou0000/COMB.2026-09-09T23:59:55ZAccepted to MICCAI 2026Dou Hoon KwarkKianoush FalahkheirkhahJi-hun OhShirui LuoVolodymyr KindratenkoRohit Bhargavahttp://arxiv.org/abs/2609.10858v1Annotating anatomy and pathology in the National Lung Screening Trial computed tomography images2026-09-09T21:52:36ZLarge-scale public medical imaging datasets contribute critically to translational research. When accompanied by rich clinical and multi-omics data, they can stimulate exploratory research and enable secondary analyses. Expert annotations of such imaging collections can support the development of new image analysis tools. Continuous enrichment of images with image-derived data makes them more usable for researchers without expertise in image analysis or access to large-scale computational resources. The National Lung Screening Trial (NLST) released a rich longitudinal dataset that includes Computed Tomography (CT) images for over 26,000 patients. We introduce three Digital Imaging and Communications in Medicine (DICOM) formatted datasets, complementing NLST CT images, shared as analysis results in the National Cancer Institute Imaging Data Commons (IDC). Two of those (IDC NLSTSeg and IDC NLSTSybil) contain DICOM-harmonized annotations and extracted measurements (for 581 and 601 NLST patients, respectively) shared earlier using research formats (Sybil and NLSTseg). The third one (TotalSegmentator-CT-Segmentations) contains volumetric segmentations generated using TotalSegmentator and radiomics features for each segment for 26,194 NLST patients.2026-09-09T21:52:36ZDeepa KrishnaswamyVamsi ThiriveedhiSuraj PaiDavid ClunieIgor OctavianoChristopher P. BridgeSteve PieperRon KikinisAndrey Fedorovhttp://arxiv.org/abs/2609.10825v1Scale-Aware 3D Deep Learning for Robust Brain Metastasis Detection in Multimodal MRI2026-09-09T20:53:10ZDetecting brain metastases in magnetic resonance imaging (MRI) remains challenging because lesions vary widely in size and appearance, with very small metastases occupying only a minute fraction of a three-dimensional input. We investigate whether combining different spatial fields of view (FOVs) improves lesion detection in multimodal MRI and present a scale-aware 3D deep-learning framework. The method uses independently trained $96^3$ and $64^3$ 3D U-Nets whose whole-volume probability maps are combined by weighted late fusion. This design allows us to study the effect of spatial context separately from image resolution and modality choice. On a 97-patient development cohort, cross-FOV fusion improved lesion-level precision and F1 while substantially reducing false positives relative to the individual models. A same-FOV ensemble control showed that these gains were not explained solely by averaging independently trained networks, supporting a contribution from complementary spatial context. An exploratory cross-FOV agreement filter reduced false positives but did not improve overall F1. These results support cross-FOV probability fusion as a simple and computationally practical strategy for improving the precision-false-positive trade-off in 3D brain-metastasis detection.2026-09-09T20:53:10Z12 pages, 3 figures, 3 tablesProceedings of the International Conference on Medical Image Computing and Computer-Assisted Intervention (MICCAI 2026), Strasbourg, Sep 2026Sylvain JaumeHongming WangSimon K. Warfieldhttp://arxiv.org/abs/2503.17970v2PathoHR: Breast Cancer Survival Prediction on High-Resolution Pathological Images2026-09-09T19:18:39ZBreast cancer survival prediction in computational pathology presents a remarkable challenge due to tumor heterogeneity. For instance, different regions of the same tumor in the pathology image can show distinct morphological and molecular characteristics. This makes it difficult to extract representative features from whole slide images (WSIs) that truly reflect the tumor's aggressive potential and likely survival outcomes. In this paper, we present PathoHR, a novel pipeline for accurate breast cancer survival prediction that enhances any size of pathological images to enable more effective feature learning. Our approach entails (1) the incorporation of a plug-and-play high-resolution Vision Transformer (ViT) to enhance patch-wise WSI representation, enabling more detailed and comprehensive feature extraction, (2) the systematic evaluation of multiple advanced similarity metrics for comparing WSI-extracted features, optimizing the representation learning process to better capture tumor characteristics, (3) the demonstration that smaller image patches enhanced follow the proposed pipeline can achieve equivalent or superior prediction accuracy compared to raw larger patches, while significantly reducing computational overhead. Experimental findings valid that PathoHR provides the potential way of integrating enhanced image resolution with optimized feature learning to advance computational pathology, offering a promising direction for more accurate and efficient breast cancer survival prediction. Code will be available at https://github.com/AIGeeksGroup/PathoHR.2025-03-23T07:37:24ZFirst author comment: We are withdrawing this manuscript to address several unresolved limitations and complete the necessary internal review and approval procedures before further dissemination (confirmed by the corresponding author)Yang LuoShiru WangJun LiuJiaxuan XiaoRundong XueZeyu ZhangHao ZhangYu LuYang ZhaoYutong Xiehttp://arxiv.org/abs/2412.19990v3SegKAN: High-Resolution Medical Image Segmentation with Long-Distance Dependencies2026-09-09T19:17:15ZHepatic vessels in computed tomography scans often suffer from image fragmentation and noise interference, making it difficult to maintain vessel integrity and posing significant challenges for vessel segmentation. To address this issue, we propose an innovative model: SegKAN. First, we improve the conventional embedding module by adopting a novel convolutional network structure for image embedding, which smooths out image noise and prevents issues such as gradient explosion in subsequent stages. Next, we transform the spatial relationships between Patch blocks into temporal relationships to solve the problem of capturing positional relationships between Patch blocks in traditional Vision Transformer models. We conducted experiments on a Hepatic vessel dataset, and compared to the existing state-of-the-art model, the Dice score improved by 1.78%. These results demonstrate that the proposed new structure effectively enhances the segmentation performance of high-resolution extended objects. Code will be available at https://github.com/goblin327/SegKAN2024-12-28T03:27:21ZFirst author comment: Due to unresolved limitations in this study, we believe that the conclusions are not yet sufficiently supported. We therefore withdraw this preprint and advise readers not to cite it (confirmed by the corresponding author)Shengbo TanRundong XueShipeng LuoZeyu ZhangXinran WangLei ZhangDaji ErguZhang YiYang ZhaoYing Caihttp://arxiv.org/abs/2609.10733v1XPos3R: Cross-Modal Transformer for Intraoperative 2D/3D Registration2026-09-09T18:29:26ZIntraoperative 2D/3D registration, which aligns live X-ray images with preoperative volumes, is essential for image-guided interventions. Previous regression-based methods suffer from limited generalization, thus requiring time-consuming patient-specific retraining. Inspired by recent geometry foundation models such as DUSt3R, we propose XPos3R, a generalizable pose regression method that eliminates preoperative preparation. Unlike existing geometry models designed for homogeneous inputs, XPos3R extends this paradigm to multi-modal inputs, namely 2D X-rays and 3D volumes. Specifically, we introduce an asymmetric encoder-decoder architecture that improves cross-modal feature alignment while maintaining computational efficiency. To scale training under limited medical data, we adopt an anatomy-specific data curation strategy and construct million-scale synthetic datasets. Evaluated on real-world benchmarks, a single pretrained XPos3R surpasses patient-specific methods in both accuracy and robustness. With test-time optimization completed in seconds, it further reduces the 3D error to <4 mm and the reprojection error to <1 mm. The strong generalization, accuracy, and efficiency of XPos3R highlight its clinical potential, while its asymmetric framework may inspire broader cross-modal vision geometry tasks.2026-09-09T18:29:26ZAccepted to ECCV 2026Shiyan SuRuyi ZhaHongdong LiXuelian ChengZongyuan Gehttp://arxiv.org/abs/2608.31048v2OmniRAS: Standardizing Foundation Model Training and Evaluation in Robot-Assisted Surgery2026-09-09T15:42:25ZFew foundation models exist for robot-assisted surgery, partly because large robotic-surgery video corpora are difficult to assemble and existing models are evaluated mostly on laparoscopic benchmarks. Further, most existing models are evaluated on a small set of public benchmarks, mostly focused on laparoscopic surgery. We present OmniRAS, a family of 1B- and 2B-parameter V-JEPA-2.1 encoders for robot-assisted surgery, and detail their training. First, we release two densely annotated robotic-cholecystectomy datasets: OmniRAS-PR and a multi-label YT-Chole tool-verb-target task, the first triplet-style annotation for robotic cholecystectomy, together with splits, probe protocols, and an inter-rater study validating the shared phase ontology. Second, we document continued pretraining at up to 256 compute nodes with global batch 6,144 over 19 sources totaling approximately 2,650 hours of surgical video, 51% robotic, and analyze compute and data composition. Third, we evaluate against raw V-JEPA-2.1 and specialized surgical models on six tasks spanning triplet, phase, and step recognition, action segmentation, and detection, under frozen-encoder and final-four-block fine-tuning regimes. Across three seeds, this yields 254 downstream runs, including 109 with partial backbone fine-tuning. The best OmniRAS models achieve the strongest adapted results across all task families, while frozen differences are smaller.2026-08-31T16:24:26Z18 pages, 11 figures, 4 tables. Under review at IEEE Transactions on Robotics (T-RO)Leonardo BorgioliNeil GettyWenli XiuJessica CassianiAlvaro DucasCarlos Agustin OrdaHira WarisFangfang XiaRick StevensPier Cristoforo GiulianottiMilos Zefranhttp://arxiv.org/abs/2505.03380v2Synergistic Vision-Language Reinforcement Enables Scalable On-Demand Analysis across Diverse Clinical Tasks2026-09-09T15:03:16ZAccurate delineation of tumors and surrounding organs-at-risk is essential for radiotherapy, surgery and treatment response assessment, yet remains time-consuming and expertise-intensive. Existing artificial intelligence systems often require manual spatial prompts or task-specific retraining, while generic class labels provide limited semantic grounding for heterogeneous disease targets. Here we present SyRe, a promptable segmentation foundation model based on Synergistic vision-language Reinforcement. SyRe strengthens bidirectional interaction between visual and linguistic representations to improve semantically grounded spatial understanding. To support large-scale training, we introduce the Color Region Description strategy and construct SyReData, comprising 20 million image-mask-description triplets across 9 modalities and 229 segmentation tasks. Training with diversified prompt forms further enables open-ended prompting, invalid-prompt rejection and flexible switching between single- and multi-target analysis. SyRe achieves accurate text-prompted segmentation across diverse clinical scenarios, with particularly strong performance on disease-related targets. Across 28 unseen external datasets, including 20 cancer types and multinational in-house cohorts, SyRe generalizes robustly under real-world distribution shifts. SyRe-generated masks also preserve clinically relevant quantitative information in pathology and yield radiomics features that stratify survival and improve prognostic modeling across five retrospective CT and MRI tumor cohorts. Finally, clinician-in-the-loop refinement enables efficient case-level correction when greater precision is required. These results establish SyRe as a generalizable foundation for scalable quantitative oncology and clinician-guided segmentation refinement.2025-05-06T10:00:08ZHaonan WangJiaji MaoLehan WangQixiang ZhangMarawan ElbatelYi QinHuijun HuBaoxun LiWenhui DengWeifeng QinHongrui LiJialin LiangJun ShenXiaomeng Lihttp://arxiv.org/abs/2609.10112v1Storage-Scalable Progressive Semantic Communication via Knowledge-Base Reuse2026-09-09T12:46:19ZExisting knowledge-base-assisted semantic communication schemes commonly adopt either single knowledge-base quantization (SKBQ) or multi-knowledge-base residual quantization (MKBQ). SKBQ incurs limited storage overhead but has restricted quantization capacity, whereas MKBQ supports progressive refinement by assigning an independent knowledge base (KB) to each stage, causing the KB storage to grow linearly with the transmission depth. To address this problem, we propose storage-scalable knowledge-base reuse quantization (SSKBQ), which reuses a compact set of KBs across multiple residual refinement stages and thereby decouples the number of transmission stages from the number of maintained KBs. A stage-aware residual supervision mechanism is further introduced to regularize intermediate quantized representations and encourage progressive refinement. Experimental results demonstrate that KB reuse provides an effective solution to the storage scalability problem while maintaining competitive progressive reconstruction performance.2026-09-09T12:46:19ZSubmitted manuscriptHeng ZhuYe LiuKun ZhuFeifei Songhttp://arxiv.org/abs/2608.18915v2Simple, Safe, and Overlooked: Reclaiming Sustainable Domain Generalization with Statistical Color Matching2026-09-09T12:22:42ZHardware shifts, color variations, and changing patient characteristics between development and deployment routinely break trained medical image classifiers. Existing remedies fall short: standard color jittering provides insufficient diversity, while deep generative style transfer algorithms hallucinate features, destroy clinically relevant structures, and waste massive compute resources. To address this, we revisit classical statistical color matching and repurpose it as Colorist, a highly efficient data augmentation strategy that applies global mean-standard deviation matching directly in the RGB color space. We demonstrate that this training-free, fully interpretable approach safely generates structurally intact domain variations, outperforming deep generative models in structural fidelity and color alignment. Across out-of-distribution histopathology, peripheral blood, dermatology, and retinal datasets, it improves balanced accuracy by up to +9% over state-of-the-art domain generalization regularizers and by +13% over an unaugmented baseline. Moreover, by avoiding neural networks in the augmentation loop, Colorist preserves anatomical structure, minimizes carbon footprint, and integrates seamlessly into standard dataloaders. Together, these findings establish statistical matching as a safe, interpretable, yet overlooked alternative to deep architectures for clinical robustness. Source code is available at https://github.com/sdoerrich97/colorist.2026-08-19T13:47:51ZAccepted to DEMI @ MICCAI 2026 (4th Workshop in Data Engineering in Medical Imaging)Sebastian DoerrichFrancesco Di SalvoShyam Nandan RaiMarco LentsChristian Ledig