Width-Robust Learnability in Mean-Field Bayesian Neural Networks
arXiv:2607.05735v1 Announce Type: new Abstract: Infinite-width limits are a standard way to reason about neural networks, but it is not automatic that the limiting learner has…
arXiv:2607.05735v1 Announce Type: new Abstract: Infinite-width limits are a standard way to reason about neural networks, but it is not automatic that the limiting learner has…
arXiv:2607.05694v1 Announce Type: new Abstract: Logit-based watermarking is a widely used mechanism for identifying LLM generated content, yet its effectiveness is governed by a fundamental trade-off…
arXiv:2607.04360v1 Announce Type: new Abstract: Conditional generative models have emerged as powerful tools for sampling from target conditional distributions, driving substantial advances across a wide range…
arXiv:2606.10559v2 Announce Type: replace-cross Abstract: If the denominator in a tamed stochastic gradient Langevin update uses the current stochastic-gradient draw, the conditional mean can be biased…
arXiv:2410.13800v4 Announce Type: replace Abstract: Physically motivated stochastic dynamics are widely used to sample from high-dimensional distributions. However, such samplers often get trapped in metastable states,…
arXiv:2506.15199v4 Announce Type: replace-cross Abstract: While there are many applications of ML to scientific problems that look promising, visuals can be deceiving. Using numerical analysis techniques,…
arXiv:2510.22298v2 Announce Type: replace Abstract: Uncovering the causal mechanisms of complex real-world systems remains a significant challenge, as these systems often entail high data collection costs…
arXiv:2602.02908v2 Announce Type: replace-cross Abstract: Diffusion models trained on different, non-overlapping subsets of a dataset often produce strikingly similar outputs when given the same noise seed.…
arXiv:2607.04442v1 Announce Type: new Abstract: Diffusion models (DMs) are a state-of-the-art generative method to approximately sample from an unknown distribution. Their training and evaluation primarily rely…
arXiv:2606.16730v2 Announce Type: replace Abstract: We re-interpret Transformer pretraining as a fast-slow, singularly perturbed flow along depth, with untied weights as its non-autonomous feature. The linearised…