Not All LLM Reasoning is Visible in the Chain-of-Thought
arXiv:2607.22925v1 Announce Type: new Abstract: A key question for AI safety is whether a language model expresses all of its reasoning in its output tokens. We…
arXiv:2607.22925v1 Announce Type: new Abstract: A key question for AI safety is whether a language model expresses all of its reasoning in its output tokens. We…
arXiv:2408.03910v3 Announce Type: replace-cross Abstract: Large Language Models (LLMs) excel in stand-alone code tasks like HumanEval and MBPP, but struggle with handling entire code repositories. This…
arXiv:2607.22709v1 Announce Type: new Abstract: The proliferation of internet memes has introduced new complexities to automated content moderation, particularly in detecting misogyny. Memes often rely on…
arXiv:2607.22763v1 Announce Type: new Abstract: Leaf veins exhibit remarkable diversity in architecture and patterning, yet existing gene--environment association studies have primarily quantified leaf venation using a…
arXiv:2607.22777v1 Announce Type: new Abstract: Protein language models learn transferable sequence representations. However, because they primarily model contextual dependencies along amino-acid sequences, their training objectives do…
arXiv:2607.24472v1 Announce Type: cross Abstract: We develop a general framework of identification and estimation for automatic debiased machine learning (DML) where the parameter of interest $theta_0$…
arXiv:2607.23723v1 Announce Type: cross Abstract: We study the problem of recovering latent inner products from a random geometric graph with anisotropic Gaussian latent points. More precisely,…
arXiv:2607.24507v1 Announce Type: cross Abstract: Existing methods mainly adapt pretrained autoregressive (AR) language models to masked diffusion, whereas we directly adapt them to uniform-noise diffusion, where…
arXiv:2607.22565v1 Announce Type: new Abstract: With the widespread deployment of edge-side AI inference, edge platforms are increasingly required to support latency-sensitive, highly concurrent, and reliability-critical applications.…
arXiv:2607.22859v1 Announce Type: new Abstract: Mathematical Word Problems (MWPs) are an important benchmark for evaluating natural language understanding and quantitative reasoning. Despite recent progress in high…
arXiv:2607.24579v1 Announce Type: cross Abstract: When computing sub/super-level-set persistent homology (PH), the effect of noise may introduce millions of (short-lived) topological generators, presenting an obstacle to…
arXiv:2603.08521v3 Announce Type: replace Abstract: Understanding dynamic 3D environments in a spatially continuous and temporally consistent manner is fundamental for robotics and autonomous driving. While recent…
arXiv:2607.22766v1 Announce Type: new Abstract: The alignment of Large Language Models (LLMs) is increasingly bottlenecked by data quality. As datasets scale, massive preference and instruction-tuning corpora…
arXiv:2605.19911v2 Announce Type: replace-cross Abstract: Photonic neuromorphic computing offers a promising route to overcoming the limitations of conventional Von Neumann architectures by exploiting the high bandwidth,…
arXiv:2606.02101v3 Announce Type: replace Abstract: This paper proposes a method of creating synthetic data (SD) that will have two important advantages for the user compared to…
arXiv:2607.23304v1 Announce Type: new Abstract: Modern predictive systems are expected to adapt their behavior to the specific situation they are facing. A clinical model should not…
arXiv:2510.23942v2 Announce Type: replace Abstract: We describe a theory and implementation of an intuitionistic decentralized framework for causal discovery using judo calculus, which is formally defined…
arXiv:2607.23838v1 Announce Type: cross Abstract: Retrieval-Augmented Generation (RAG) lets a large language model answer questions using documents retrieved from an external knowledge base at query time.…
arXiv:2607.22884v1 Announce Type: new Abstract: We propose CHiPS, a lightweight character-level authorship attribution method for Romanian texts. All reported experiments are closed-set: the true author is…
arXiv:2607.22824v1 Announce Type: cross Abstract: Comparing CT reconstruction methods fairly is labor-intensive and largely manual, and many benchmarks use idealized data. We ask whether a large…
arXiv:2607.22712v1 Announce Type: new Abstract: Single-cell light microscopy images have become an important data source for characterizing cell phenotypes, but their complexity and heterogeneity pose challenges…
arXiv:2607.22748v1 Announce Type: new Abstract: Modern neural networks primarily adapt through parameter modification within predefined computational structures. While recent methods introduce modularity, conditional computation, and parameter-efficient…
arXiv:2607.24537v1 Announce Type: cross Abstract: In the Big Data era, the scalability of clustering algorithms constitutes a key challenge. Traditional density-based methods (e.g., DBSCAN) offer robustness…
arXiv:2607.23679v1 Announce Type: cross Abstract: Recent years have witnessed increasing interests in tackling heteroscedastic noise in bandits and reinforcement learning. In these works, the cumulative variance…