Invisible Shortcuts: Why Vision Encoders Know Your Camera
arXiv:2608.05424v1 Announce Type: new Abstract: Deep vision models exploit shortcuts, relying on cues that correlate with supervision signals. Prior work has focused on visible biases, such…
arXiv:2608.05424v1 Announce Type: new Abstract: Deep vision models exploit shortcuts, relying on cues that correlate with supervision signals. Prior work has focused on visible biases, such…
arXiv:2608.03571v2 Announce Type: replace Abstract: Recent works train agents by constructing large-scale multimodal environment pools. However, we find that simply increasing the number of multimodal environments…
arXiv:2608.05393v1 Announce Type: new Abstract: Prompt tuning, a leading parameter-efficient adaptation paradigm in NLP, has recently been extended to computer vision. Visual prompt tuning (VPT) adapts…
arXiv:2607.04884v2 Announce Type: replace Abstract: We present HunyuanOCR-1.5, a lightweight end-to-end OCR-specialized vision-language model. HunyuanOCR unifies document parsing, text spotting, information extraction, text-image translation, and multi-image…
arXiv:2608.05389v1 Announce Type: new Abstract: Background: Accurate glioma subregion delineation is important for radiotherapy planning and longitudinal monitoring, but manual contour correction is time-consuming. Models such…
arXiv:2604.04444v2 Announce Type: replace Abstract: Open-vocabulary object detection (OVOD) enables models to detect any object category, including unseen ones. Benefiting from large-scale pre-training, existing OVOD methods…
arXiv:2608.05356v1 Announce Type: new Abstract: High-definition 3D LiDAR maps are important for autonomous driving and smart-city services, which require reliable detection of object-level changes in multi-temporal…
arXiv:2506.14243v4 Announce Type: replace Abstract: LiDAR-based place recognition is critical for long-term autonomous driving without GPS. Existing handcrafted feature methods face dual limitations. First, descriptor instability…
arXiv:2608.05341v1 Announce Type: new Abstract: Vision-Language Models (VLMs) for radiology report generation are typically trained on retrospective clinical reports, which suffer from omission noise: clinically present…
arXiv:2608.06037v1 Announce Type: cross Abstract: Relational inductive biases are essential for capturing structural dependencies among data. This study investigates a dual-level relational framework for image classification,…
arXiv:2608.05213v1 Announce Type: new Abstract: The style of a painting is not monolithic: color, texture, and structure may come from different sources. Existing reference-guided methods transfer…
arXiv:2608.05226v1 Announce Type: new Abstract: Neuron counting and segmentation in microscopy images of neuronal cultures is a routine and time-consuming task in neuroscience research, traditionally performed…
arXiv:2608.05237v1 Announce Type: new Abstract: Current few-step autoregressive video diffusion models depend on previous fully denoised clean frames as context for all denoising steps of the…
arXiv:2608.05260v1 Announce Type: new Abstract: Contrastive vision-language models such as CLIP and BLIP are typically trained on short image captions, limiting their ability to retrieve images…
arXiv:2608.05258v1 Announce Type: new Abstract: Gradient-weighted Class Activation Mapping (Grad-CAM) is widely used to visualize model decisions, but it was originally formulated for convolutional neural networks,…
arXiv:2511.14907v2 Announce Type: replace Abstract: Computational pathology holds substantial promise for improving diagnosis and guiding treatment decisions. Recent pathology foundation models enable the extraction of rich…
arXiv:2608.05333v1 Announce Type: new Abstract: In-context learning (ICL) adapts medical image segmentation models to unseen structures and modalities without retraining by conditioning on a task-specific support…
arXiv:2606.09368v2 Announce Type: replace Abstract: Scene Graphs (SGs) provide structured representations of visual scenes by modeling objects and their pairwise relationships. Despite recent progress, existing datasets…
arXiv:2608.05210v1 Announce Type: new Abstract: Picture books and comics have long been used to disseminate hateful narratives because they are easily understood even by children, as…
arXiv:2608.05145v2 Announce Type: replace Abstract: While modern 3D reconstruction excels at modeling object geometry and appearance, it largely ignores the rich acoustic cues revealed through physical…
arXiv:2608.05600v1 Announce Type: cross Abstract: Flow-based generative models are typically sampled by solving a deterministic ordinary differential equation (ODE), whereas online reinforcement learning requires stochastic rollouts…
arXiv:2507.14022v2 Announce Type: replace Abstract: This study proposes the Cognitive Pairwise Comparison Classification Model Selection (CPC-CMS) framework for document-level sentiment analysis. The CPC, based on expert…
arXiv:2608.05166v1 Announce Type: new Abstract: We present an evaluation of cognitive bias expression in state-of-the-art instruction-tuned LLMs under realistic multi-turn interaction settings. Our work introduces a…
arXiv:2608.06352v1 Announce Type: cross Abstract: Training terminal agents requires executable and verifiable tasks that are not merely solvable, but appropriately challenging for learning. Executable validation establishes…