Third-party cyber evaluations involving OpenAI models
Article URL: https://openai.com/index/third-party-cyber-evaluations-involving-openai-models/ Comments URL: https://news.ycombinator.com/item?id=49175248 Points: 5 # Comments: 0
Article URL: https://openai.com/index/third-party-cyber-evaluations-involving-openai-models/ Comments URL: https://news.ycombinator.com/item?id=49175248 Points: 5 # Comments: 0
Article URL: https://www.troyhunt.com/thanks-fedex-this-is-why-we-keep-getting-phished/ Comments URL: https://news.ycombinator.com/item?id=49175192 Points: 91 # Comments: 11
More info: https://github.com/lechmazur/debate https://preview.redd.it/5rr60v33bfhh1.png?width=3600&format=png&auto=webp&s=bbe9ff2e3c2e599e887a4ec0da7bcc26f71ad4fb This benchmark measures how well LLMs hold an argument under adversarial, multi-turn opposition across a wide range of topics. It rewards broad…
The UK’s @AISecurityInst (AISI) has published a report on their recent cybersecurity evaluation of Anthropic’s Claude Mythos 5 and OpenAI’s GPT-5.6 Sol. The models attempted to…
SpaceX has ramped up its purchases of Tesla Megapacks for its xAI data centers.
We're detailing two new incidents that occurred during external cyber evaluations conducted by independent evaluation partners.We outline what happened, how the activity was contained, and how…
Driven by demand for AI capacity, AMD's data center revenue more than doubled year-over-year in its latest earnings report, reaching $6.7 billion. That's up from $5.8…
SpaceX has committed to using Nvidia GPUs exclusively because they are the best
Article URL: https://twitter.com/i/status/2084739205071343837 Comments URL: https://news.ycombinator.com/item?id=49174900 Points: 41 # Comments: 15
SpaceX's AI revenue grew more than three times to $2.6 billion from the year before, mostly because of deals that the company made to provide compute…
Linkpost for my Substack piece, lightly adapted. Donors are holding AI wealth at enormous risk. Their giving is often extremely risk-averse.“What is the most common mistake…
Governor who touted Texas as AI “epicenter” pauses data center grid connections.
sampler : remove "full-context windows" from history-based samplers (#26524) Resolve -1 to 1024 instead of ctx-len for samplers Because of backend-sampling we initialize samplers before the…
I’ve seen a lot of people whose reviewers went silent after initial reviews, but I am also noting abnormally quiet authors. I ultimately withdrew my paper,…
This is a very late post about a project that I did a few months ago as part of the application to Neel Nanda's MATS 10.0…
RT RinaGrok 4.5 just turned Blender into a conversation.Instead of manually importing assets, fixing rigs, placing armies, adjusting cameras, and debugging objects that face the wrong…
I’ve been running a fairly opinionated evaluation loop on ~35B A3B/MoE-class models for coding over the past few months. Not synthetic benchmarks: actual dev workflows, iterative…
Article URL: https://www.sec.gov/Archives/edgar/data/1795071/000179507126000002/xslFormDX01/primary_doc.xml Comments URL: https://news.ycombinator.com/item?id=49174407 Points: 16 # Comments: 5
Article URL: https://electrek.co/2026/08/04/waymo-co-ceo-camera-only-self-driving-tesla/ Comments URL: https://news.ycombinator.com/item?id=49174369 Points: 12 # Comments: 4
The same Starmind V1 satellite design (minus solar & radiator) will be deployed on the ground in our data centers. It’s a major improvement in data…
RT NVIDIAAI compute is going to orbit. 🚀@SpaceX’s Starmind AI1 satellite compute payload is powered by NVIDIA Vera Rubin NVL72, bringing AI factory compute closer to…
AI compute is going to orbit. 🚀@SpaceX’s Starmind AI1 satellite compute payload is powered by NVIDIA Vera Rubin NVL72, bringing AI factory compute closer to the…
A new SaferAI report finds Z.ai's open-weight GLM-5.2 approaches frontier AI capabilities while lacking key safety mitigations, renewing concerns that powerful open models could outpace governance…
We analyzed Kimi K3 and GPT-5.6 Sol on DeepSWE. A Kimi-first cascade with test-suite verification outperformed Sol alone at a lower cost per completed task.