Li Bearden / Writing
Writing
Working notes
Notes on LLM evaluation methodology, eval validity, and reproducibility — and whether published benchmark numbers survive being re-run. Posts live on Substack; this is the index.
All posts
01
AUG 25, 2026
Sycophancy Meta-Evals: Metric-Dependent Reliability
the results of hunting for model instability
↗
02
AUG 23, 2026
Sycophancy Meta-Evals: Conflation vs. Cancellation
why sycophancy averages erase behavioral structure
↗
03
AUG 22, 2026
Sycophancy Meta-Evals: Caving vs. Self-Correction
why sycophancy benchmarks disagree on what counts as a failure
↗
04
AUG 21, 2026
Making Networking Possible Outside an AI Safety Hub
How a regional EA conference helped me overcome geographic isolation and connect with community
↗
05
AUG 20, 2026
Thinking in Public
why I'm writing every day about AI safety evals
↗