Adversarial Attacks on Deep OCR Systems
arXiv:2608.07636v1 Announce Type: cross Abstract: Deep-OCR (DeepSeek-OCR) advances document recognition by treating the visual modality as an optical compression medium, enabling long-context OCR at low token…
arXiv:2608.07636v1 Announce Type: cross Abstract: Deep-OCR (DeepSeek-OCR) advances document recognition by treating the visual modality as an optical compression medium, enabling long-context OCR at low token…
arXiv:2608.07562v1 Announce Type: new Abstract: Urban flooding poses an escalating threat to transportation infrastructure, yet no operational system provides real-time, street-level flood-depth estimates at centimeter resolution.…
arXiv:2608.07543v1 Announce Type: new Abstract: Background and Study Aims: Accurate optical diagnosis of colorectal polyps guides resection strategy and surveillance, with multimodal large language models (MLLMs)…
arXiv:2608.07547v1 Announce Type: new Abstract: Indoor scene layout generation is a challenging task in interior design. Existing methods often oversimplify the task by reducing room conditions…
arXiv:2608.07549v1 Announce Type: new Abstract: Triangle meshes provide explicit and accurate surface geometry, yet their irregular topology connectivity makes 3D mesh tokenization a geometric sampling problem:…
arXiv:2608.07554v1 Announce Type: new Abstract: The development of vision-based wildfire detection systems for unmanned aerial vehicles is constrained by the limited availability of diverse real-world training…
arXiv:2608.07550v1 Announce Type: new Abstract: Vision-language models return structured chest-radiograph findings through interfaces exposing no confidence score, so a receiving institution cannot read off how far…
arXiv:2608.07561v1 Announce Type: new Abstract: Chronic kidney disease (CKD) is a silent disease. Its progression may not significantly hamper a person's daily routine. Human kidney function…
arXiv:2608.07559v1 Announce Type: new Abstract: In 3D reconstruction, mirrors introduce significant challenges by creating distorted and fragmented spaces, resulting in inaccurate and unreliable 3D models. As…
arXiv:2506.00633v3 Announce Type: replace Abstract: Generating semantically controllable 3D CT volumes from radiology reports requires more than a rich text encoder, it requires vision-language alignment grounded…
arXiv:2608.08965v1 Announce Type: cross Abstract: Underwater images often suffer from diverse and coexisting degradations, including color distortion, scattering haze, texture attenuation, and uneven illumination. These degradations…
arXiv:2607.10744v4 Announce Type: replace Abstract: Benefiting from the powerful priors embedded in large-scale pre-training data and the emerging commonsense reasoning ability, large language models (LLMs) have…
arXiv:2603.07066v2 Announce Type: replace Abstract: Generative diffusion models are increasingly used for medical imaging data augmentation, but text prompting cannot produce causal training data. Re-prompting rerolls…
arXiv:2608.09529v1 Announce Type: new Abstract: As an important subfield of cross-modal generation, synthesizing static visual content in the form of images from audio, namely audio-to-image (A2I)…
arXiv:2506.01353v3 Announce Type: replace-cross Abstract: The integration of brain-computer interfaces (BCIs), in particular electroencephalography (EEG), with artificial intelligence (AI) has shown tremendous promise in decoding human…
arXiv:2608.09637v1 Announce Type: new Abstract: Diffusion models have enabled high-quality video generation in recent years, but the high cost of iterative sampling hinders their practical deployment.…
arXiv:2602.02704v2 Announce Type: replace Abstract: Reasoning over ultra-long documents requires synthesizing sparse evidence scattered across distant segments under strict memory constraints. While streaming agents enable scalable…
arXiv:2608.07891v1 Announce Type: new Abstract: Self-introductions are common in legislative committee testimonies. Successfully detecting them and extracting the speaker's name is enormously helpful in the task…
arXiv:2510.22028v4 Announce Type: replace Abstract: Quality Estimation (QE) metrics are vital in machine translation for reference-free evaluation and increasingly serve as selection criteria in data filtering…
arXiv:2608.07862v1 Announce Type: new Abstract: Existing safety evaluation datasets for large language models (LLMs) predominantly focus on English and Western contexts, often overlooking the linguistic diversity…
arXiv:2608.09861v1 Announce Type: cross Abstract: Audio-visual interaction is the standard for patient-physician consultations, enabling natural communication and effective assessment of illness through non-verbal cues. While text-based…
arXiv:2608.07852v1 Announce Type: new Abstract: How a language model internally represents who is speaking, the Assistant, an assigned roleplay persona, or a narrated story character, remains…
arXiv:2608.07527v1 Announce Type: new Abstract: Long-document understanding requires models to find and combine evidence across many pages, layouts, tables, figures, and charts. Existing retrieval-augmented systems usually…
arXiv:2608.07812v1 Announce Type: new Abstract: A host of recent studies have evaluated the cognitive and developmental alignment of Foundation Models (FMs). These investigations include evaluations of…