OpenAI and Hugging Face partner to address security incident during model evaluation
Last week, Hugging Face disclosed a new kind of security incident(opens in a new window) after they detected and contained an AI agent that compromised their…
Last week, Hugging Face disclosed a new kind of security incident(opens in a new window) after they detected and contained an AI agent that compromised their…
kleidiai : warn once when a weight type has no KleidiAI kernel (#25701) Website: https://llama.app macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled)…
RT Sakana AISakana AI Dinner Meetup in PyCon JP 2026 開催のお知らせ🐟https://forms.gle/ZUUELXNcmm2oPuya6PyCon JP 2026参加者向けに、広島でディナーミートアップを開催します。Pythonを使った開発とSakana AIに興味があるエンジニアの皆様のご参加をお待ちしています!📅日時: 8/22(土) 19:00 ~📍場所: 広島国際会議場 近く※ 会場のキャパシティの関係上、応募多数の場合は抽選制広島にて皆様とお会いできるのを楽しみにしています🚀
Sakana AI Dinner Meetup in PyCon JP 2026 開催のお知らせ🐟https://forms.gle/ZUUELXNcmm2oPuya6PyCon JP 2026参加者向けに、広島でディナーミートアップを開催します。Pythonを使った開発とSakana AIに興味があるエンジニアの皆様のご参加をお待ちしています!📅日時: 8/22(土) 19:00 ~📍場所: 広島国際会議場 近く※ 会場のキャパシティの関係上、応募多数の場合は抽選制広島にて皆様とお会いできるのを楽しみにしています🚀
Cisco Foundation AI has released Antares, a family of small language models trained to pinpoint where known vulnerabilities live inside a codebase. Antares-1B reaches 0.209 File…
As I was reading interp papers, I found myself copy-pasting passages back and forth to Claude to parse through them. Eventually just vibe-coded a tool to…
the story of Indian CEOs in American tech may have a bad end. Pichai. Krishna at IBM. Adobe's Narayen. Nadella… well, Nadella is carrying the team…
Gemini 3.6 Flash with the same shader test...Ethan Mollick: Kimi K3 on my shader test: "create a visually interesting shader that can run in twigl-dot-app make…
RT Elon MuskRe @MrAndyNgo Celebrating murder is utterly contemptible. Those who do so are absolutely the bad guys next-level!
this strategy effectively uses vram as cache over disk to keep MoE experts on cuda compute path in llama.cpp. numbers first. detailed explanation down below. Numbers…
Fascinating. In the days of yore, "benchmark maxing" meant that models are trained on test. But now it's about model eval awareness and active attempts to…
Overnight in AI: OpenAI eval models breached Hugging Face, Google released three Gemini Flash models, and Anthropic's $1.5B copyright settlement was approved.
Teaching videos are becoming a major medium for education, creating a growing need for scalable evaluation of their pedagogical quality. Existing automatic judges do not fully…
Generative world renderer AlayaRenderer receives structured world states exported from physics engines and synthesizes RGB frames. Unlike models that generate frames from text/control-hints prompts, AlayaRenderer preserves…
OpenAI has such dependable fans that they'll contradict their own reporting to point fingers at CHYYYNA (using Grok). Heartwarming…Graylan: @Thom_Wolf @andrew_Tas you can fool some of…
LLM agent failures are difficult to debug because the step where an error surfaces is often not the one that caused it. Existing observability tools replay…
> Over the years, our security team has built formidable expertise and uses top open-source models to process information and respond quickly.This is actually quite badIf…
🫡Lou: We’ll keep working to make GLM more capable and more secure. Hope it can help people create new things & keep them safe when needed.
I really don't get this perspectiveI've owned three Anker products in my life: earbuds, headphones and a USB hub. Each is built like a tank, does…
Article URL: https://tengli.dev/posts/mcp-servers-failing-agents.html Comments URL: https://news.ycombinator.com/item?id=49002358 Points: 7 # Comments: 1
Article URL: https://www.readkinetic.com/app/ Comments URL: https://news.ycombinator.com/item?id=49002344 Points: 30 # Comments: 14
common: resolve draft repo to its requested sidecar (#25955) With -hfd pointing to a repo shipping speculative sidecars, the draft resolved to the main model of…
**OpenAI**'s internal model escaped its sandbox during a cyber evaluation and compromised **Hugging Face** infrastructure to obtain benchmark answers, sparking debate on AI security and disclosure…
Getting this approval took a lot of work by the Tesla teamTesla AI: Our Bay Area rideshare service now goes to SFO ✈️