Recognizing Co-Speech Gestures in-the-Wild
arXiv:2605.31589v2 Announce Type: replace Abstract: While humans naturally gesture during speech, only a sparse subset of these co-speech gestures are visually depictive and semantically linked to…
arXiv:2605.31589v2 Announce Type: replace Abstract: While humans naturally gesture during speech, only a sparse subset of these co-speech gestures are visually depictive and semantically linked to…
arXiv:2608.06407v1 Announce Type: new Abstract: Automated Sign Language Recognition for under-represented languages remains a largely unsolved problem. Central African Sign Language (CASL) exemplifies this gap: the…
arXiv:2608.06467v1 Announce Type: new Abstract: Facial expression recognition (FER) in videos is challenging because models must identify subtle, temporally evolving affective states that vary across individuals.…
arXiv:2608.06490v1 Announce Type: new Abstract: We present InsertFuse, a unified framework for multi-category reference-guided image insertion. Its key idea is to decouple category-specific expertise learning from…
arXiv:2608.06599v1 Announce Type: new Abstract: Mandibular reconstruction restores facial continuity and oral function after segmental resection. Patient-specific cutting guides transfer a computed tomography (CT)-based plan to…
arXiv:2608.06580v1 Announce Type: new Abstract: Face Recognition (FR) systems in surveillance settings often encounter Low Resolution (LR) faces, those whose face region falls below the standard…
arXiv:2608.06613v1 Announce Type: new Abstract: Self-supervised 3D medical foundation models are increasingly used as general-purpose feature extractors, yet their sensitivity to MRI artifacts remains poorly understood.…
arXiv:2608.06612v1 Announce Type: new Abstract: The plethora of readily available geospatial data offers exciting opportunities to learn high quality representations of the planet, but the sheer…
arXiv:2608.01035v2 Announce Type: replace-cross Abstract: Vision-Language-Action (VLA) models have emerged as a prominent paradigm for end-to-end autonomous driving; however, their efficient deployment is severely constrained by…
arXiv:2608.05137v2 Announce Type: replace Abstract: Understanding 3D scenes is fundamental to embodied intelligence, requiring joint reasoning over heterogeneous information from multiple modalities, including visual and geometric…
arXiv:2608.07193v1 Announce Type: cross Abstract: Visual-token pruning can substantially reduce the inference cost of multimodal large language models (MLLMs), yet existing methods largely rely on fixed,…
arXiv:2608.06934v1 Announce Type: cross Abstract: Visual perception of walkability varies substantially across individuals, reflecting differences in personal characteristics, experiences, and preferences. Existing studies, however, often reduce…
arXiv:2512.11534v2 Announce Type: replace Abstract: Key frame selection is essentially a set-level optimization problem: the quality of the selected subset depends on the interactions among frames,…
arXiv:2505.04397v3 Announce Type: replace Abstract: Modern vision networks are dominated by additive local transformations, whereas explicit multiplicative local interactions remain underexplored. Product units offer a direct…
arXiv:2604.11240v2 Announce Type: replace Abstract: Token pruning has emerged as an effective approach to reduce the substantial computational overhead of Large Vision-Language Models (LVLMs) by discarding…
arXiv:2608.06607v1 Announce Type: new Abstract: Most document-extraction systems use a single model for all documents. This is simple but can be costly for easy cases and…
arXiv:2607.27853v2 Announce Type: replace Abstract: Powered by advances in LLMs and autonomous agents, deep research has become one of the most widely adopted agentic products. However,…
arXiv:2608.06589v1 Announce Type: new Abstract: While large language model outputs are frequently analysed as a collective super variety termed "AI language," this chapter argues that this…
arXiv:2606.15059v2 Announce Type: replace Abstract: Simultaneous speech-to-speech translation (SimulS2ST) enables real-time cross-lingual communication, but existing evaluation has focused largely on short or pre-segmented speech rather than…
arXiv:2608.06549v1 Announce Type: new Abstract: LLMs are increasingly being applied to tasks involving institutional and political texts, but existing benchmarks evaluate them on isolated documents or…
arXiv:2604.06756v2 Announce Type: replace Abstract: Large language models (LLMs) has been widely adopted as a scalable surrogate for human evaluation, yet such judges remain imperfect and…
arXiv:2608.06539v1 Announce Type: new Abstract: False-presupposition QA (FPQA) tests LLMs on their ability to identify false presuppositions in questions and abstain or correct them rather than…
arXiv:2608.07449v1 Announce Type: cross Abstract: LLM agents increasingly adapt to recurring tasks by accumulating procedural knowledge in skills. These skills are lightweight, reusable textual artifacts that…
arXiv:2608.06532v1 Announce Type: new Abstract: LVLMs are increasingly used to read financial charts, tables, and documents, where a single misread figure can move a decision and…