The Advantage of Fine-Grained Training
arXiv:2509.05130v2 Announce Type: replace Abstract: In classification problems, models are trained to predict a class label based on the input data features. However, class labels are…
arXiv:2509.05130v2 Announce Type: replace Abstract: In classification problems, models are trained to predict a class label based on the input data features. However, class labels are…
arXiv:2607.20656v3 Announce Type: replace Abstract: Effective decision-making in complex and changing environments requires balancing short-term and long-term consequences. In reinforcement learning (RL), this trade-off is typically…
arXiv:2606.00675v2 Announce Type: replace Abstract: Water research in Brazil largely overlooks the widespread damming of small streams for agricultural uses including watering cattle, farm-scale hydropower, irrigation,…
arXiv:2607.26336v1 Announce Type: new Abstract: In model-based reinforcement learning, world models exist as internal simulators, but their training often conflates statistical correlations with causal mechanisms. This…
arXiv:2607.26247v1 Announce Type: new Abstract: Low-rank adaptation (LoRA) fine-tunes large pretrained models at a fraction of the cost of full fine-tuning, but its performance depends strongly…
arXiv:2502.10605v4 Announce Type: replace-cross Abstract: Problem definition: Estimating causal effects of interventions is central to policy and operations, but outcome data are often missing or costly…
arXiv:2607.27036v1 Announce Type: cross Abstract: Video diffusion-based world models enable long autoregressive video generation for robotics, autonomous driving and simulation tasks, yet sliding-window autoregressive inference suffers…
arXiv:2607.26339v1 Announce Type: new Abstract: Retrieval-Augmented Generation (RAG) systems ground large language models (LLMs) in external corpora, but this reliance exposes them to corpus poisoning: maliciously…
arXiv:2607.26192v1 Announce Type: new Abstract: Input-dependent controller coefficients are often treated as evidence of dynamic inference or computational savings. This interpretation conflates three properties: coefficient variation,…
arXiv:2607.26246v1 Announce Type: new Abstract: On-policy distillation (OPD), which aligns a student with the teacher's token-level distribution on the student's own rollouts, is an effective paradigm…