Re Read more on ParallelKernelBench: https://www.together.ai/blog/parallelkernelbench
Re Read more on ParallelKernelBench: https://www.together.ai/blog/parallelkernelbench
Re Read more on ParallelKernelBench: https://www.together.ai/blog/parallelkernelbench
Multi-GPU kernels are the real test for coding models. Today at @aiDotEngineer, @simran_s_arora shared ParallelKernelBench, an open-source benchmark for evaluating whether LLMs can write fast CUDA…
Re 9/ ParallelKernelBench: Benchmarking LLMs on Multi-GPU Kernel GenerationPaper: https://www.alphaxiv.org/abs/2606.parallel-kernel-bench
Re 6/ When RL Meets Adaptive Speculative Training: A Unified Training-Serving System (Aurora)Paper: https://arxiv.org/abs/2602.06932
Re 7/ Untied Ulysses: Memory-Efficient Context Parallelism via Headwise ChunkingPaper: https://arxiv.org/abs/2602.21196
Re 8/ Opportunistic Expert Activation: Batch-Aware Expert Routing for Faster Decode Without Retraining (OEA)Paper: https://arxiv.org/abs/2511.02237
Re 3/ Learning to Discover at Test Time (TTT-Discover)Paper: https://arxiv.org/abs/2601.16175
Re 5/ V1: Unifying Generation and Self-Verification for Parallel ReasonersPaper: https://arxiv.org/abs/2603.04304
Re 2/ ThunderAgent: A Simple, Fast and Program-Aware Agentic Inference SystemPaper: https://arxiv.org/abs/2602.13692
Re 4/ Escaping the Verifier: Learning to Reason via Demonstrations (RARO)Paper: https://arxiv.org/abs/2511.21667