Terminal-Bench Leaderboard Rankings: Luck or Skill?
Terminal-Bench 2.0, one of the best agent benchmarks, runs agents on 89 tasks with at least 5 tries per task. The aggregate rankings are then published…
Terminal-Bench 2.0, one of the best agent benchmarks, runs agents on 89 tasks with at least 5 tries per task. The aggregate rankings are then published…
Article URL: https://chipsandcheese.com/p/nvidias-vera-whitepaper-has-a-thread Comments URL: https://news.ycombinator.com/item?id=49189234 Points: 11 # Comments: 0
Next week on Together AIhttps://www.together.ai/models/qwen3-8-max
Meta expanded its AI coding offerings with a new agent that, it promises, can handle complex tasks with complex software.
Article URL: https://twitter.com/nikitabier/status/2085105586966827343/ Comments URL: https://news.ycombinator.com/item?id=49189113 Points: 13 # Comments: 3
Article URL: https://www.primeintellect.ai/blog/prime-agent Comments URL: https://news.ycombinator.com/item?id=49189075 Points: 13 # Comments: 0
Article URL: https://community.openai.com/t/how-openai-lost-a-paying-customer-over-160-it-refuses-to-explain/1389233 Comments URL: https://news.ycombinator.com/item?id=49188980 Points: 11 # Comments: 1
I'm mosty interested in 128-192GB VRAM with 128-256GB RAM to spare, so SSD streaming is basically not even necessary. Seems only FP4 is supported, so older…
RT alignedaiRe @AIatMeta needs to share this:muse-spark-1.2 pricing vs muse-spark-1.2-contributor it is much much much cheaper almost free if you are okay with training: great for…
Anthropic and OpenAI models’ unprompted actions forced halt to UK cyber tests.
A logical decision theory recommends that you choose as if deciding the output of your decision algorithm. The main difficulty in formulating a logical decision theory…
Small detail I caught when re-reading DeepSeek-v4 writeup: they add some ratio of agentic trajectories into mid-training. Interestingly, this is the only sentence in the paper…
The s̶t̶u̶d̶e̶n̶t̶ transformer has become the masterAidan Gomez: Time to bring back Google Brain🧟♂️
submitted by /u/ECrispy [link] [comments]
Xiaomi-Robotics-1 is a robot foundation model trained on over 100K hours of real-world manipulation trajectories. It is a Vision-Language-Action (VLA) model engineered for out-of-the-box mobile manipulation…
RT Zi Lin🎮 This entire Avo Lawn game was built with Muse Spark 1.2 + Muse Code. 🥑Now it’s your turn—use Muse Spark 1.2 to build…
RT Vishal MisraThe bottleneck for AI progress was never compute, it was always the verifier. Recursive self-improvement is limited by verification, not computation. Compute buys proposals…
daily reminder not to trust benchmarks and run it yourself. claimed e2e speedup is ~40%, forwards are ~140% faster I would wager that compared to a…
I was on here a while ago showcasing it. I've stoped playing with it so I'm open sourcing it. I figure I'll let other people play…
RT Character.AIintroducing LongSqueak🐁 a new Chat Style built for long-form storytelling, with rich, novella-like replies 📖rolling out to (c.ai+) members starting today! with 4x the Memory,…
server: harden the file_glob_search directory walk (#26626) server: don't walk Windows junctions in file_glob_search std::filesystem reports a junction as a plain directory, so the symlink guard…
Please message me if you think you can help or want to set up an agreement.Discuss
Meta Superintelligence Labs has released Muse Code, a terminal coding agent in beta, powered by the new Muse Spark 1.2 model. Muse Code plans changes, writes…
full house for the team’s talk at Black Hat on the OpenAI-Hugging Face Incident