Hi there! I'm Yiderigun Borjigin.
I am a doctoral researcher at Saarland University, supervised by Prof. Dr. Roland Aydin. I study the reliability and alignment of large language models through behavioral evaluation and mechanistic interpretability.
My work asks how context shapes a model’s answers: when it follows relevant evidence, and when it is swayed by a suggested number, a confident but incorrect explanation, or the way data is presented. I build controlled evaluations to study these effects and use interventions on model activations to investigate their mechanisms. My work also includes pretraining-time alignment and the reliability of language models in scientific applications.
Looking ahead, I’m interested in AI alignment and agentic systems, especially how interactions among multiple agents can give rise to social dynamics and safety risks.
