Test-time compute beats parameter count on hard reasoning
A 32B model with 10× the inference compute matches a 540B baseline on AIME 2025 and codeforces — but only when the reward model is itself learned online.
The Research Desk reads the firehose so you don't have to. Verdicts are after the abstract, before the press release.
A 32B model with 10× the inference compute matches a 540B baseline on AIME 2025 and codeforces — but only when the reward model is itself learned online.
Backdoors planted at pre-training stay hidden through three rounds of safety fine-tuning, including red-team adversarial. A serious lifecycle concern.
128 experts, 8 active per token. Routing learned with auxiliary-loss-free balancing; outperforms a dense 200B model at 1/9th the FLOPs / token.
Models that are taught to explicitly read and reason about the safety spec become more — not less — robust to jailbreaks as they get smarter. A rare good-news result.
A single 7B vision-language-action model controls humanoids, quadrupeds and arms with zero per-embodiment fine-tuning. The 'foundation model for atoms' thesis gains evidence.
Generated scenes stay self-consistent for >10 minutes of free navigation. The first credible 'spatial' counterpart to large language models.