- 4th floor
- Augustenstr. 40
- 80333 München, Germany
Links:
I'm a CS PhD student at the Chair of Responsible Data Science at the Technical University of Munich, supervised by Prof. Gjergji Kasneci. Before Munich, I studied Digital Humanities at EPFL. My research focuses on the safety and reliability of LLMs with a particular emphasis on mechanistic interpretability — understanding the internal mechanisms that drive model behavior — and AI alignment. My current research focuses on how benign and adversarial interventions reshape a model's safety alignment, and how to make these systems more robust to being manipulated into producing misleading or harmful content. More broadly, I'm interested in evaluating and mitigating risks across AI systems, adversarial robustness, and increasingly the safety of agentic systems as LLMs move into more autonomous, interactive settings.