These startups are chasing the next big thing in LLMs
MIT Technology Review’s What’s Next series looks across industries, trends, and technologies to give you a first look at the future. You can read the rest…
Every story across every category, newest first. Each card links to the original publisher; daily-brief posts open as editorial pages.
MIT Technology Review’s What’s Next series looks across industries, trends, and technologies to give you a first look at the future. You can read the rest…
Security firm PromptArmor shows how hidden instructions in a PDF can hijack Atlassian's AI agent Rovo, silently forwarding sensitive data from Jira and Confluence to an…
Code and data available at github.com/KieronKretschmar/latent-awarenessTL;DRWe take two eval-gaming model organisms (Hua et al.'s (2025) organism and RogueQwen) and apply direct preference optimization (DPO) to their…
ggml-webgpu : refactor several wgsl files and simplify flash_attn wgsl. (#26134) Website: https://llama.app macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS…
Test data from public benchmarks inevitably leaks into pretraining corpora, inflating evaluation scores once memorized. Contamination mitigation evaluation intervenes in the decoding process to suppress memorization…
Spec-dec has been a thing for a while, in fact, it's wasn't an idea that was born for LLM inference. E.g. Uber's https://github.com/uber/submitqueue applied it to…
Recent works train agents by constructing large-scale multimodal environment pools. However, we find that simply increasing the number of multimodal environments does not always benefit. We…
Supervised Fine-Tuning (SFT) and Reinforcement Learning (RL) exhibit fundamentally different behaviors in enhancing multi-task reasoning for large language models (LLMs). Our preliminary experiments revealed a phenomenon:…
«The Iran War seems to be a game of mutually assured humiliation»Special Providence Poster (GDP Stan): @teortaxesTex The Iran War seems to be a game of…
MiniMax H3 is an open-weight, general-purpose multimodal video generation model that works across text, images, video, and audio. In ComfyUI, you can use H3 for text-to-video,…
One potential risk of developing general-purpose robots is that they could greatly reduce the friction required to establish a totalitarian regime. If robots became physically capable…
RT himanshuI was chatting with this person from a Neolab (valued around ~2B) and they are already looking for design partners and enterprises to partner on…
Grok Imagine is focused on professional usefulness, fun for consumers and overall ease-of-useX Freeze: Grok 4.5 just ranked #1 on Design Arena for Daily Usage -…
comments like this on the aie channel miss the point.- we are building a community and an industry that is bigger than any one person can…
YesBeff (e/acc): 🗣️: "So your goal is to colonize the Moon, use it to build the AI Dyson Swarm, climb the Kardashev scale and usher a…
Whoops! Enabled Sol as a plan model for a V4-Flash project in omp, and it INSTANTLY blew through $19 and my remaining OR credits. The $0.31…
Overnight in AI: Google breaks up DeepMind, an AI autonomously hacked an Australian gym, and Anthropic defaults Claude Code to auto mode.
bro…Hunter Weiss: DARPA Lift Contest Winners 1st Avidrone $1.25M 2nd MTech $750K3rd Xtreme Aerial $500K
Article URL: https://www.docker.com/products/docker-sandboxes/ Comments URL: https://news.ycombinator.com/item?id=49239751 Points: 27 # Comments: 11
ByteDance’s Seed team has introduced SeedRealtime, a native audio-visual full-duplex LLM. The model fuses audio, video and text in a single unified architecture. It interacts in…
**Frontier API vulnerability** revealed exposure of hidden reasoning traces including sensitive data like **62 unique API keys** and **33 passwords**, raising privacy and operational-security concerns. Discussions…
**Meta** re-enters the open-weight frontier with the release of **Muse Glimmer**, a **30B dense**, multimodal, agent-focused model under **Apache 2.0**, optimized for always-on local agents and…
Humans don't have "less alien values" than AI, WeiHumans have human values. That is what non-alien values are.Thanks for being consistent, though. Yes, the logical endpoint…
This is the advice I wish I had when I started trying to become an AI safety research engineer.The LandscapeStart by working out which issues you…