Curriculum vitae
Download PDFResearch experience
Doctoral Researcher, AI Safety and Scientific ML
Saarland University · supervised by Prof. Dr. Roland Aydin
Lead author on LLM safety research spanning behavioral evaluation, chain-of-thought faithfulness, and mechanistic analysis of how context steers model answers.
Research Associate
Helmholtz-Zentrum Hereon, Institute of Material Systems Modeling, Geesthacht
Built AnchorBench, and ran a mechanistic study of chain-of-thought faithfulness using activation patching and logit-lens attribution.
Master Thesis Researcher
BMW Group, Research and Innovation Center (FIZ), Munich
Deep learning models for multivariate time-series prediction of vehicle thermal behavior, with transfer learning and domain adaptation.
Teaching Assistant, Fundamentals of Artificial Intelligence
Technical University of Munich · built exercises on constraint satisfaction
Education
PhD in Artificial Intelligence (in progress)
Saarland University, Saarbrücken
M.Sc. Robotics, Cognition, Intelligence
Technical University of Munich · ML, deep learning, computer vision
B.Eng. Automotive Engineering
Jilin University, Changchun, China
Research methods and infrastructure
Evaluation infrastructure
Benchmark and harness design in Python, built end to end; vLLM on multi-GPU H100 nodes, HuggingFace Transformers, OpenRouter and provider APIs (OpenAI, Anthropic, Google, xAI); batched greedy and sampled decoding, deterministic answer extraction, parse-free logit readouts, seed-controlled data generation, checksummed result releases.
Interpretability
Residual-stream activation patching, donor-state and layer-sweep interventions, logit-lens attribution, next-token logit margins, teacher-forced span scoring.
Model adaptation
PyTorch, LoRA / QLoRA fine-tuning (PEFT), supervised fine-tuning pipelines, transfer learning, domain adaptation.
Experimental statistics
Cluster bootstrap CIs, Wilcoxon signed-rank tests, Benjamini–Hochberg and Holm correction, TOST equivalence testing, pre-declared thresholds and validity audits, ablation and sensitivity design.
Benchmarks and models evaluated
MMLU-Redux, MMLU-Pro, GPQA-Diamond, GSM8K, MATH-500, AIME, CruxEval; Llama 3.1–3.3, Qwen2.5 / Qwen3, Gemma 2–3, OLMo 2, DeepSeek-R1-Distill, gpt-oss, GPT-5.x, Claude, Gemini, Grok.
