Zuck will "share more on open source" soon
submitted by /u/realmvp77 [link] [comments]
submitted by /u/realmvp77 [link] [comments]
Most coverage of Microsoft's SkillOpt centers on its 52/52 result. The more consequential finding is in Section 4.3: the exported best_skill.md keeps working in environments it…
An AI model from Meta also hacked another company during testing Stop me if you've heard this one before: An AI model from the parent company…
xAI's Grokipedia, an online encyclopedia with AI-generated articles that Elon Musk once promised would be a "massive improvement" over Wikipedia, apparently hasn't been updated since April…
One of the things the pelican benchmark is still useful for is visually representing (to a tiny extent) the improvements in a single model familyHere's Meta…
Run Claude Code sessions on your own compute
We ran 900 DeepSWE rollouts on DeepSeek-V4 Flash and GPT-5.6 Luna. Luna leads pass@1 by 14 points; DeepSeek delivers 4.8x the solves per dollar.
Large language models (LLMs) with integrated search tools show strong promise in open-domain question answering (QA), yet they often struggle to produce complete answer set to…
Agent Plugins 1.0.0 is a new, vendor-neutral directory specification—backed by Google, Amazon, Microsoft, and others—for packaging Agent Skills and MCP servers into a single portable unit.…
New OpenAI Signals data shows how people use ChatGPT worldwide, with country-level insights on adoption, usage trends, and evolving behavior.
The quality of open-weight language models has dramatically improved in recent years. Sharing weights greatly facilitates model adoption by enabling their use across diverse hardware and…
Millennium and Anthropic are building a digital risk analyst with Claude
Introducing Muse Code and Muse Spark 1.2 Yet more evidence that the most important characteristic of any model these days is long-sequence agentic tool calling. Meta…
Just had to create an "accidental-cyberattacks" tag on my blogWe're up to four now: the original OpenAI+Hugging Face one, Anthropic's me-too attacks, then two new ones…
Third-party cyber evaluations involving OpenAI models And another one. I had to create a accidental-cyberattacks tag to keep track of them all! This post from OpenAI…
Is it your prediction that Anthropic ARR will not be 100B or higher by end of this year?If so, how will you update your worldview if…
Google DeepMind lost both Demis Hassabis and Jeff Dean in a single day; AI agents from Anthropic and OpenAI went rogue during UK safety tests; Meta…
Incident Report: unsanctioned agent behaviour during cyber testing It happened again. This time it was the UK government's AI Security Institute who accidentally attacked other companies…
Prime Agent is an open-source coding and research agent for general and long-running work. A self-improving RLM harness for coding and long-running autonomous tasks. Designed to…
I mean, Deepseek V4 Flash is an absolutely fantastic model, even though I can't run it on my machine it's so fascinating to see how it…
RT SauersUPDATE: a challenger emergesMTS: SITUATION DETECTED: A Meta model hacked into another company's systems during cybersecurity testing, per The Information.
Thanks to everyone who submitted an essay to the contest (LW mirror). Links to all the entrants are below.I’m reviewing them now and aim to announce…
Dr. Alex Turner (@TurnTrout) is an AI safety researcher with pioneering work in activation steering and power-seeking theory. He recently resigned from Google DeepMind over the…
RT S.E. Robinson, Jr.If you have paid attention, you would see Elon says this often. Fashion has become stagnant and is due for an update. Maybe…