Hallo there! I'm Yiderigun Borjigin.
I am a doctoral researcher in AI safety at Saarland University, supervised by Prof. Dr. Roland Aydin, working on behavioral evaluation of LLMs, chain-of-thought faithfulness, and mechanistic interpretability.
My research asks when a model's output is driven by evidence and when it is driven by something else in its context — a number planted in a retrieved document, a confidently phrased wrong reasoning trace, an artifact of how data was serialized into text. I build controlled benchmarks that isolate one causal factor, run them across open-weight and frontier models, and trace the behavior back to internal state via residual-stream patching. The finding that recurs is that high task accuracy does not imply robustness.
