US Govt to individually approve who gets GPT 5.6
Article URL: https://old.reddit.com/r/LocalLLaMA/comments/1ufo0un/us_govt_to_individually_approve_who_gets_gpt_56/ Comments URL: https://news.ycombinator.com/item?id=48683021 Points: 10 # Comments: 4
Article URL: https://old.reddit.com/r/LocalLLaMA/comments/1ufo0un/us_govt_to_individually_approve_who_gets_gpt_56/ Comments URL: https://news.ycombinator.com/item?id=48683021 Points: 10 # Comments: 4
I made a quick guide for myself while wanting to try the new models, so I share it with you. It's pretty basic, but it may…
RT Sakana AIRe CoffeeBenchでは、6体のエージェントがメールや取引で相互作用し、各社が利益の最大化を目指します。LLMエージェントが経営を担う社会が来たときに、協調や競争、ときに不正はどう現れるのか。CoffeeBenchは、それを観察するための実験場でもあります。
RT Sakana AISakanaAIは、有限責任あずさ監査法人と共同で、LLMエージェントの長期的な経営能力を評価する新しいベンチマーク「CoffeeBench」を公開しました。ブログ:https://sakana.ai/coffee-bench/現実の経済では、消費者へ直接売るビジネスだけでなく、企業同士が継続的に取引するビジネスも重要です。CoffeeBench は、農家・焙煎店・小売店の計6社が参加するコーヒー業界のサプライチェーンをシミュレーションし、各社をLLMエージェントが運営。90日間にわたって価格交渉・発注・在庫管理などを行い、純利益の最大化を目指します。最新のLLMを同じ環境で競わせると、経営成績は大きく分かれ
Despite their widespread use, the role of reward models in shaping reinforcement learning is poorly understood. Reward models offer a tempting promise: they automatically estimate response…
Speculative decoding (SD) accelerates autoregressive Large Language Models (LLMs) by drafting multiple tokens and verifying them in parallel, but it faces a scaling limitation: increasing the…
Scientific reasoning models for biology combine language models with foundation models trained on multimodal biological data, including DNA, RNA, and proteins. These models are built through…
**OpenAI** previewed **GPT-5.6** with three variants: **Sol** (flagship), **Terra** (mid-tier), and **Luna** (lower-cost), launching under a restricted rollout mandated by the U.S. government, limiting access to…
I split qwen 27b and Gemma 4 26b (moe) across a 5080, and 2x 5060ti. I noticed setting split mode to tensor mode will cause looping…
Article URL: https://akrites.org/letter/ Comments URL: https://news.ycombinator.com/item?id=48682737 Points: 5 # Comments: 0
ChatGPT edu is expiring on our university, no word on renewal, and I have a bunch of impatient grad students and a vibe-coding supervisor who all…
Very impressive, genuinely frontier /goal behaviorcedric: I gave GLM-5.2 one goal:Turn a 2D dungeon crawler it built into a 2.5D isometric Diablo-style game with the visual…
This is very good. I'd say the most impressive attempt at a non-invasive brain recording technology I've seen so far. The end goal is high-resolution ultrasound-based…
Sharing a project I have been working on called Third Eye. It does visual geolocation. Given a video, it figures out where it was filmed using…
Down to $4100 for a new Unitree humanoid modelOne consumer GPU's worthPoe Zhao: Unitree cut its R1 humanoid robot to RMB 29,900 ($4,100). Immediate availability. No…
I want to bring your attention to JetSpec because it looks strictly smarter and stronger than previous speculative decoding and block diffusion approaches (yes, again).Avg 1000…
Curious to hear how long you guys are waiting for a long context session(100k tokens+) to resume in coding agents running locally. submitted by /u/sayamss [link]…
I officially petition my followers and anyone interested who speaks Chinese fluently and might have the chops for this job: Please, try to join DeepSeek. This…
Modern generative world models render increasingly realistic action-controllable futures, yet they frequently hallucinate: rollouts remain visually fluent while drifting from the ground-truth dynamics. We hypothesize that…
A classical intuition holds that verifying a solution is easier than producing one. For today's coding agents, this intuition is being inverted: as foundation models develop…
Vance is just correct. Nixon was a great man of history, a thinker, and his "corruption" was tame and sounds quaint now. Similar power level inflation…
RT vLLMExcited to see @cohere open-source how they use AI coding agents to maintain their vLLM fork. 🙌Keeping a long-lived fork in sync is tricky. Their…
Clueless fools: "AI subscriptions are subsidized! or no, new limits, free lunch is over!"Reality: Arnaud and his likes are *subsidizing Anthropic*I repeat, AI is a very…
> The executives agreed to join on the condition that Sjoerdsma would refrain from raising human rights issues or arms sales to TaiwanMan, the Dutch are…