Research worth knowing
Papers and technical reports, benchmarks and what they still measure, training methods and architectures, datasets and corpora, and what alignment and interpretability work found.
On the register
- Activation probes as monitors watching
- Agent and work-task benchmarks watching
- Agentic misalignment evaluations watching
- Automated research agents watching
- Benchmark re-issues watching
- Chain-of-thought monitorability watching
- Coding benchmark saturation watching
- Diffusion language models watching
- Evaluation awareness and sandbagging watching
- Genome models at base resolution watching
- Hybrid and linear attention in shipped models watching
- Interactive benchmarks watching
- Learned weather models in operational use watching
- Machine-checked mathematics watching
- Muon and second-order optimisers watching
- On-policy distillation watching
- Open video datasets watching
- Openly licensed pretraining corpora watching
- Operators disclosing agents acting outside sanction watching
- Peer review under volume watching
- Residual stream redesigns watching
- Retroactive opt-out in training data watching
- Reward hacking and emergent misalignment watching
- Safeguards for open-weight release watching
- Sparse attention for million-token context watching
- Sparse autoencoders and transcoders watching
- Synthetic data in pretraining corpora watching
- Task time horizons watching
- Test-time compute and overthinking watching
- Unsolved-problem benchmarks watching
- Verifier-gated supervision watching
Everything filed here
- 2026-09-22 Google Research describes a generative-UI framework for classroom simulations · Research worth knowing · 2 sources