lab notes

About

This is a log of small, self-contained AI safety and security experiments, mostly run on a laptop.

I’m Lily Sijia Li, and I work on making AI safe in high-stakes environments. I spent 5 years at Cisco and later Microsoft, securing those same environments against human attackers in financial services, healthcare, and legal.

But what do we do when the attacker isn’t a human, but the AI itself? That’s the question I now research at the University of Oxford. Models can fail by accident, be compromised with a backdoor, or be misaligned. In each case, the system inside the high-stakes environment is the thing we need to watch.

I currently work on two research projects, both about getting AI to know when not to act. One looks at how LLM agents behave when they call tools, and how to stop them when they behave anomalously. The other is on reward modelling for clinical generation, where the model has to know when to answer and when to abstain.

This blog lab notes is where I publish small experiments around the same questions. Each post is an experiment, and what I found.


Find me on GitHub or LinkedIn.