Open sea under a clear sky, a motorboat leaving a long wake near the horizon

Research and artifacts

I study LLM reliability and alignment: how context and training shape model behavior, and how controlled experiments can reveal why models fail.

Behavioral evaluation of LLMs

Controlled benchmarks that hold the evidence fixed and vary one causal factor at a time, run across open-weight and frontier models, to measure when context rather than evidence decides the answer.

Reasoning behavior and mechanistic interpretability

How incorrect reasoning in the context influences a model’s answer, and what activation interventions reveal about that influence. This connects behavioral evaluation with questions about chain-of-thought faithfulness.

Pretraining-time alignment

Collaborative work on Synthetic Persona Pretraining, studying how shaping a model’s persona during pretraining affects its later alignment and robustness.

AI for science: text interfaces for PDE modeling

Reliability metrics for the interface between numerical data and language models, including inverse parameter estimation under zero-shot, in-context, and LoRA/QLoRA regimes.

Future interests

Alignment and multi-agent systems

I want to explore the development and safety of LLM-based agentic systems, particularly systems in which multiple agents interact. I’m curious about whether phenomena resembling those in human societies emerge, how they arise, and what risks they create. These are questions I’m developing into a research direction.

Open-source artifacts

AnchorBench — evaluation harness and public dataset

2026

14,400 released prompts with gold answers and anchor values; seed-controlled deterministic generation, one shared parser across all conditions (median parse rate 99.9%), and SHA-256 checksums linking every published number to the raw generations behind it.

github.com/Ydrg9989/AnchorBench huggingface.co/datasets/Yiderigun/AnchorBench