Li Bearden

Li Bearden / About

About & CV

Elliot “Li” Bearden

I'm an AI safety research engineer focused on LLM evaluation methodology. My research is on run-level reproducibility of sycophancy benchmarks — whether the single-run numbers these benchmarks report survive being re-run, with PARROT as a positive control and SycEval as the open falsification test. It started with an identity-priming effect that collapsed into a false positive under replication; chasing down why is how reproducibility became the actual question. MSCS completed July 2026.

Five years deploying speech and language models at Deepgram gave me a front-row seat to the gap between benchmark performance and deployment behavior — the gap my research now targets directly. I built and owned the evaluation and custom-training infrastructure, including the pipeline that cut custom model delivery from 14 days to 1.

Earlier civic-tech work as a Civic Digital Fellow at the NIH National Library of Medicine (through Coding it Forward) shapes how I think about institutional dynamics in safety review — how measurement infrastructure gets designed, contested, and used inside organizations.

Li Bearden

Where I work

Engaging with AI-safety communities across Southeast Asia

Currently based in Chiang Mai, Thailand — remote-first and a US citizen, working close to the region's emerging AI-safety hubs. Open to short in-person sprints, including the EU and UK, for a few weeks at a time.

Chiang Mai · home base AI-safety hubs Southeast Asia · scope

Path so far

Placed on a real time axis — experience on the spine, degrees braced across the years they ran

Today

Skills

Research methods

LLM evaluation designcontrolled prompt-variant experimentssycophancy & honesty evaluationeval validity / reproducibilitymeasurement for hard-to-measure behaviors

Tools

PythonRustSQLPyTorchHuggingFaceOpenAI EvalsW&BvLLMRayDockerKubernetesAWS