Li Bearden / Writing
Writing
Working notes
Notes on LLM evaluation methodology, eval validity, and reproducibility — and whether published benchmark numbers survive being re-run. Posts live on Substack; this is the index.
All posts
01
AUG 25, 2026
Working like you’re being shot at, maladaptive ambition, and other things to avoid
why sustainability in AI Safety research stays hard despite all of the kumbayas
↗
02
AUG 25, 2026
Sycophancy Meta-Evals: Metric-Dependent Reliability
the results of hunting for model instability
↗
03
AUG 23, 2026
Sycophancy Meta-Evals: Conflation vs. Cancellation
why sycophancy averages erase behavioral structure
↗
04
AUG 22, 2026
Sycophancy Meta-Evals: Caving vs. Self-Correction
why sycophancy benchmarks disagree on what counts as a failure
↗
05
AUG 21, 2026
Making Networking Possible Outside an AI Safety Hub
How a regional EA conference helped me overcome geographic isolation and connect with community
↗
06
AUG 20, 2026
Thinking in Public
why I'm writing every day about AI safety evals
↗