Just 10 more hours to qwen 3.8 27b…
What the first thing you're going to ask it? submitted by /u/RandumbRedditor1000 [link] [comments]
Every story across every category, newest first. Each card links to the original publisher; daily-brief posts open as editorial pages.
What the first thing you're going to ask it? submitted by /u/RandumbRedditor1000 [link] [comments]
Ok, maybe it was more like 7 hourslet's see. Zhipu doesn't have anything fancy, neither scale nor architecture. Pure brutal training competence. How far further could…
Semi-supervised semantic segmentation has long turned on one question, which pseudo-labels to trust, and a generation of selection rules, dynamic thresholds, per-class curricula, soft confidence weights,…
new X update makes me feel like DeepSeek, maybe that's why I empathize so muchJMB 🧙♂️: pretty good for a blind guy tho. i had to…
Large-scale manipulation data is essential for robot learning, yet collecting robot demonstrations remains expensive and difficult to scale. Meanwhile, abundant egocentric human manipulation videos provide rich…
Try Grok 4.6 image & video understanding, it’s a major upgrade!Yun-Ta Tsai: Grok 4.6 multimodal is a step change from Grok 4.5. It’s one of the…
Guys, look here please.RedNote's AI lab says they have a new RL regimen (TEMPO; beta version) that rewards world exploration allowing a 16B active MoE to…
LmaoMakes no sense though, it doesn’t use fewer steps or tokens than Flash. How is this data generated?(Most likely still Pareto-mogging)Zain: whats insane is that DeepSeek-v4-pro…
RT @zephyr_z9: The magic of AFD (Attention FFN Disagg)It can handle the big boi Sol too (over 3T parameters)This is the first commerciall…
No single large language model (LLM) is optimal across all queries and budget constraints, making model routing essential for cost-effective deployment. Existing routers adopt diverse formulations…
We present DreamX-Phi 1.0, an action-conditioned video world model for robotic manipulation that, given an observed frame, a language instruction, and a prescribed action sequence comprising…
Spatial intelligence is becoming a foundation for embodied agents, robotic planning, and multimodal assistants. To improve the spatial reasoning ability of VLM agents, existing work has…
Closed models used to be the default. Not anymore.NYC devs: come chat about it in person.Together AI × OpenRouter are hosting a fireside chat on the…
arXiv:2608.12640v1 Announce Type: cross Abstract: Causal discovery aims to uncover the underlying causal relationships given data generated from a system. The goal, however, is not merely…
arXiv:2608.12623v1 Announce Type: cross Abstract: Language model classifiers with explanations are used for moderation, routing, topic triage, and low-resource annotation. We study black-box auditing when the…
arXiv:2605.07060v3 Announce Type: replace-cross Abstract: Physics-informed neural networks (PINNs) provide a mesh-free framework for solving PDE-constrained inverse problems, but their extension to Bayesian inversion still faces…
arXiv:2608.12589v1 Announce Type: cross Abstract: Factor-MIDAS regressions forecast a low-frequency target by extracting common factors from a large panel of high-frequency predictors via principal component analysis…
arXiv:2608.12470v1 Announce Type: cross Abstract: Large astrophysical simulation campaigns often generate training data by sampling parameters across a Uniform prior box. Due to the proposal's sharp…
arXiv:2504.18455v2 Announce Type: replace Abstract: We study distributed multiview representation learning, a problem in which $K$ clients each observe a distinct but possibly statistically correlated view.…
arXiv:2608.12973v1 Announce Type: new Abstract: In this paper, we study how to perform statistical inference for quantile temporal difference learning (QTD) in distributional reinforcement learning. Assuming…
arXiv:2608.13171v1 Announce Type: new Abstract: To avoid missing important variables and their connections in networks, more and more variables are included in network analysis. Here we…
arXiv:2608.13418v1 Announce Type: new Abstract: Given a dataset where a portion of the samples are contaminated, our goal is to recover the underlying clean population distribution.…
arXiv:2608.13201v1 Announce Type: new Abstract: We develop the statistical and algorithmic theory of inverse optimal transport (IOT) under the feature-parameterized cost C_theta(i,j) = -theta^T phi(i,j). The…
arXiv:2604.24196v4 Announce Type: replace Abstract: A drifting model is a one-step generator trained by moving each sample along a field of kernel-weighted attraction toward data samples…