Disentangling 3D Modeling from Spatial Reasoning
arXiv:2608.05242v1 Announce Type: new Abstract: In this work, we explore an alternative paradigm for spatial reasoning by explicitly disentangling 3D perception from reasoning, rather than jointly…
arXiv:2608.05242v1 Announce Type: new Abstract: In this work, we explore an alternative paradigm for spatial reasoning by explicitly disentangling 3D perception from reasoning, rather than jointly…
arXiv:2608.05249v1 Announce Type: new Abstract: Real-world multimodal instructions often bundle multiple requirements with unequal importance, yet most multimodal training data still reduce instruction following to answering…
arXiv:2608.05243v1 Announce Type: new Abstract: Factorized generative models commonly regularize a latent style variable z_s by matching its marginal distribution to a fixed Gaussian prior and…
arXiv:2608.05253v1 Announce Type: new Abstract: Quantized orthogonal fine-tuning (qoft) enables parameter-efficient adaptation of low-bit language models by learning structured activation rotations before frozen quantized weights. However,…
arXiv:2608.05250v1 Announce Type: new Abstract: Multi-task supervised fine-tuning (SFT) often casts a heterogeneous data mixture as a single optimization problem, even though different tasks may reach…
arXiv:2608.06366v1 Announce Type: cross Abstract: Electronic health record (EHR) feature engineering is a major bottleneck in clinical research and AI, accounting for 39-45% of data scientists'…
arXiv:2608.06130v1 Announce Type: cross Abstract: AI agents performing cryptographic operations (signing Git commits, authenticating API calls, issuing certificates) currently store private keys in software-accessible locations: plaintext…
arXiv:2605.04970v3 Announce Type: replace Abstract: Modern LLMs show mastery over an ever-growing range of skills, as well as the ability to compose them flexibly. However, extending…
arXiv:2601.03895v2 Announce Type: replace Abstract: Group Relative Policy Optimization (GRPO) has emerged as a popular algorithm for reinforcement learning with large language models (LLMs). However, GRPO…
arXiv:2512.08991v3 Announce Type: replace-cross Abstract: End-to-end image controllers that map raw camera frames directly to control actions are increasingly deployed in safety-critical systems. However, formally verifying…
arXiv:2605.23971v2 Announce Type: replace-cross Abstract: This work presents a physics-guided machine-learning framework for carbon monoxide concentration inference from experimentally measured resistance transients of a mixed-phase SnO-SnO$_2$…
arXiv:2608.05450v1 Announce Type: new Abstract: Pixel-space diffusion models avoid the reconstruction ceiling of latent diffusion models by generating directly in image space. However, their substantially higher…
arXiv:2605.20247v2 Announce Type: replace-cross Abstract: Catastrophic forgetting remains a major obstacle to continual learning in large language models (LLMs) and vision--language models (VLMs). Although Mixture-of-Experts (MoE)…
arXiv:2608.05424v1 Announce Type: new Abstract: Deep vision models exploit shortcuts, relying on cues that correlate with supervision signals. Prior work has focused on visible biases, such…
arXiv:2608.03571v2 Announce Type: replace Abstract: Recent works train agents by constructing large-scale multimodal environment pools. However, we find that simply increasing the number of multimodal environments…
arXiv:2608.05393v1 Announce Type: new Abstract: Prompt tuning, a leading parameter-efficient adaptation paradigm in NLP, has recently been extended to computer vision. Visual prompt tuning (VPT) adapts…
arXiv:2607.04884v2 Announce Type: replace Abstract: We present HunyuanOCR-1.5, a lightweight end-to-end OCR-specialized vision-language model. HunyuanOCR unifies document parsing, text spotting, information extraction, text-image translation, and multi-image…
arXiv:2608.05389v1 Announce Type: new Abstract: Background: Accurate glioma subregion delineation is important for radiotherapy planning and longitudinal monitoring, but manual contour correction is time-consuming. Models such…
arXiv:2604.04444v2 Announce Type: replace Abstract: Open-vocabulary object detection (OVOD) enables models to detect any object category, including unseen ones. Benefiting from large-scale pre-training, existing OVOD methods…
arXiv:2608.05356v1 Announce Type: new Abstract: High-definition 3D LiDAR maps are important for autonomous driving and smart-city services, which require reliable detection of object-level changes in multi-temporal…
arXiv:2506.14243v4 Announce Type: replace Abstract: LiDAR-based place recognition is critical for long-term autonomous driving without GPS. Existing handcrafted feature methods face dual limitations. First, descriptor instability…
arXiv:2608.05341v1 Announce Type: new Abstract: Vision-Language Models (VLMs) for radiology report generation are typically trained on retrospective clinical reports, which suffer from omission noise: clinically present…
arXiv:2608.06037v1 Announce Type: cross Abstract: Relational inductive biases are essential for capturing structural dependencies among data. This study investigates a dual-level relational framework for image classification,…
arXiv:2608.05213v1 Announce Type: new Abstract: The style of a painting is not monolithic: color, texture, and structure may come from different sources. Existing reference-guided methods transfer…