~ similar to 2607.22380· 20 results
Jiaju Han, Ma Yaqi, Yahui Chai, Xuemeng Sun +7 more
This paper introduces MonoIR-RS, a large-scale infrared remote-sensing vision-language dataset and benchmark for understanding infrared imagery.
FLORO is a multimodal geospatial foundation model that learns transferable remote sensing representations from a small, diverse corpus, achieving strong performance across various sensor types and res…
Pingping Liu, Aohua Li, Yubing Lu, Jin Kuang +2 more
The paper proposes RPCASSM, a novel state space model leveraging Robust PCA (RPCA) to accurately detect and segment infrared small targets by separately modeling background and target information base…
This paper proposes STRMSR, a through-plane super-resolution framework for clinical cardiac MRI using HR reference views and intermediate SR results as memory.
The paper introduces Multifidelity Proper Orthogonal Decomposition (MFPOD), a method that significantly reduces the computational cost of dimension reduction by intelligently combining data from cheap…
This paper demonstrates that fusing multi-viewpoint data from multiple satellites significantly enhances the accuracy of space object detection in congested LEO constellations, establishing multi-view…
First implementation of a 3D Gaussian renderer on an Intelligence Processing Unit (IPU) with 1,472 independent tiles using only on-chip SRAM.
Jiahao He, Yihua Shao, Zhengkai Zhao, Pan Gao +5 more
The paper presents GrainGS, a dynamic Gaussian framework for scene reconstruction that balances fine-grained motion modeling, structural stability, and compact representation.
The paper demonstrates that for FFT-based radar imaging on Apple Silicon, the limiting factor for half-precision (FP16) is dynamic range, not mantissa precision, and proposes a block-floating-point (B…
The paper proposes a fast and lightweight novel view synthesis method using a differentiable Multiplane Image (MPI) representation, achieving significant speed and size improvements over state-of-the-…
Xinjue Wang, Xiuheng Wang, Yejun Zhang, Sergiy A. Vorobyov +2 more
The paper investigates whether using fine-grained, tensorized adapters (CP components) instead of standard LoRA ranks improves the accuracy-budget trade-off in PEFT, finding that while they fill budge…
The paper introduces an adaptive feature-optimized vision front end that intelligently selects and budgets visual features for 3D reconstruction, significantly improving reconstruction quality and com…
The paper demonstrates a coordinated, cross-modal spoofing attack that successfully deceives state-of-the-art multi-sensor fusion systems in autonomous vehicles by making multiple sensors agree on a f…
Hao Kong, Di Liu, Shuo Huai, Xiangzhong Luo +4 more
This paper proposes a dynamic image cropping framework and compound shrinking strategy to reduce computational overhead of CNNs on edge AI, achieving higher accuracy and lower computational cost.