Research worth knowing
Papers and technical reports, benchmarks and what they still measure, training methods and architectures, datasets and corpora, and what alignment and interpretability work found.
On the register
- Activation probes as monitors watching
- Agent and work-task benchmarks watching
- Agentic misalignment evaluations watching
- Automated research agents watching
- Benchmark re-issues watching
- Chain-of-thought monitorability watching
- Coding benchmark saturation watching
- Diffusion language models watching
- Evaluation awareness and sandbagging watching
- Genome models at base resolution watching
- Hybrid and linear attention in shipped models watching
- Interactive benchmarks watching
- Learned weather models in operational use watching
- Machine-checked mathematics watching
- Muon and second-order optimisers watching
- On-policy distillation watching
- Open video datasets watching
- Openly licensed pretraining corpora watching
- Operators disclosing agents acting outside sanction watching
- Peer review under volume watching
- Residual stream redesigns watching
- Retroactive opt-out in training data watching
- Reward hacking and emergent misalignment watching
- Safeguards for open-weight release watching
- Sparse attention for million-token context watching
- Sparse autoencoders and transcoders watching
- Synthetic data in pretraining corpora watching
- Task time horizons watching
- Test-time compute and overthinking watching
- Unsolved-problem benchmarks watching
- Verifier-gated supervision watching
Everything filed here
- 2026-09-30 OpenAI publishes safety-case recommendations for frontier reinforcement learning runs · Research worth knowing · 1 source
- 2026-09-29 AISI says GPT-6 Astra completed unsanctioned attacks in simulations · Research worth knowing · 1 source
- 2026-09-28 ThreatDown says the CARBONATO botnet installs an unmodified Hermes Agent and targets AI API keys first. · Research worth knowing · 1 source
- 2026-09-28 ProgramDistill, a benchmark built from 26 web apps, scores GPT-6 Astra at 49.2 percent and Claude Opus 5 at 28.8. · Research worth knowing · 1 source
- 2026-09-27 DeepSeek describes DSec, a sandbox platform for agentic training handling 3 million sandboxes a day · Research worth knowing · 1 source
- 2026-09-27 Paper describes PrePssmCas, a classifier that fuses language-model and PSSM features for Cas proteins · Research worth knowing · 1 source
- 2026-09-26 Study finds most tested coding-agent harnesses let agents delete their own execution traces · Research worth knowing · 1 source
- 2026-09-26 Paper finds prompt injection can shift Jev's typed decisions, though rarely to the attacker's target · Research worth knowing · 1 source
- 2026-09-24 Anthropic says Claude found a new bacteriophage enzyme system · Research worth knowing · 1 source
- 2026-09-22 Google Research describes a generative-UI framework for classroom simulations · Research worth knowing · 2 sources