Behavioral evaluation of LLMs
Controlled benchmarks that hold the evidence fixed and vary one causal factor at a time, run across open-weight and frontier models, to measure when context rather than evidence decides the answer.

Doctoral researcher in AI safety · Saarland University
Controlled benchmarks that hold the evidence fixed and vary one causal factor at a time, run across open-weight and frontier models, to measure when context rather than evidence decides the answer.
Residual-stream activation patching, donor-state and layer-sweep interventions, and logit-lens attribution to locate where a reasoning trace becomes causally load-bearing.
Reliability metrics for the interface between numerical data and language models, including inverse parameter estimation under zero-shot, in-context, and LoRA/QLoRA regimes.
14,400 released prompts with gold answers and anchor values; seed-controlled deterministic generation, one shared parser across all conditions (median parse rate 99.9%), and SHA-256 checksums linking every published number to the raw generations behind it.