AI capabilities have scaled. Near cutting-edge, open-weight models lower the barrier for state and non-state actors to build advanced AI systems for adversarial purposes. As cost pushes more companies toward open-weight models, the deployment surface for models that pose enterprise and societal risk will greatly expand. Anthropic has documented fitness-seeking risk, where models perform misaligned actions judged to be optimal for the requested goal. Detecting and countering these risks requires production-grade AI safety infrastructure, not just research. And the AI safety community is undersupplied in people who can ship that infrastructure at scale. That gap motivated my transition into the space, and what I'm now building toward.
I'm building Clearwood with co-founder Michael Klear (CTO), focused on white-box monitoring and control for enterprise deployments of agentic models. We were recently funded by a BlueDot Impact rapid grant. Before transitioning into AI safety, I spent 6 years leading teams delivering critical consumer-facing ML capabilities at scale, earned a PhD in statistics, and did academic research in economics and computational biology. Before cofounding Clearwood, I received fellowships from Coefficient Giving and Simplex AI Safety to fund study and new research in technical AI safety.
Research
I explored increasing atomicity of BatchTopK SAEs, and finding belief geometries outside of tightly defined numeric use cases -- things like modular arithmetic or geographic distance -- in the representation space of transformers trained on natural language. The MetaSAE repo is meant for public reuse. The repo for finding belief geometries is research code.
Earlier in my transition, I did exploratory work on information storage versus annotated latent computation in a thinking model's chain of thought, describing activation distributions in decoder-only transformers, and investigating classification through activation distributions.
