Independent researcher working on evaluation and measurement for ML systems. Mostly I try to break claims — other people's, and my own. Several of the results below are the ones that killed my original hypothesis.
Everything here follows the same rule: the plan is frozen before the run, negative results are reported next to the positive ones, and every number reproduces from a single command.
Are LLM judges actually reliable? — a line of work:
-
🧑⚖️ llm-judge-independence — six judges from four labs and three countries carry about 1.9 independent votes, not six (mean pairwise error correlation ρ̄ ≈ 0.42, stable across four robustness arms). Majority voting and appeal rounds assume errors are independent. They aren't.
-
🧪 llm-jury — six LLM judges on RAGTruth hallucination detection. The headline finding turned out to be about the benchmark: of ~30 cases where every judge disagreed with the gold label, 11 were provable annotation errors, and all 11 of the detector's "false alarms" I checked were real hallucinations the annotators had missed. One month, ~$12 of API calls.
Other work:
- 📈 driftbet — attributes what caused a data stream to drift (model, noise, the world, or bad labels) without labels, with anytime-valid statistical guarantees.
- 🔍 whest-teardown — reproduced a published estimator; a plain baseline beat it ~9×, and I found an accounting hole in its scoring.
- 🔭 A pre-registered null result on real Kepler photometry and a cross-model study of LLM belief revision — same method, different fields.
Currently: looking for work on LLM evaluation, benchmark quality, and model-quality measurement. Open to remote and contract.
