Mathematical Scientist for AI Safety Research
LawZero· Berlin; Montreal
seen live 1 day ago
Frontier AI companies are throwing billions of dollars into scaling existing architectures and methods such as next-token-prediction, direct preference optimization (DPO), reinforcement learning with human feedback (RLHF), and reinforcement learning with verified rewards (RLVR).
This link goes straight to the employer's greenhouse page. Posted 136 days ago.We last confirmed it was open 1 day ago.