I build systems that decide whether AI is safe enough for millions of users

LLM evaluation and AI safety with a background in causal inference. Formerly Trust, Safety, & Quality research lead at a Series B AI startup.