Dear editors,
I have a question before I start work. I've read #127, so I'll skip the parts already settled there. I understand that partial replication is fine, and that writing original code from the paper's description (consulting the original code only to fill gaps) is what makes it a replication rather than a reproduction. My question is only about prior coverage.
The original article is Binz et al., "A foundation model to predict and capture human cognition," 644, 1002–1009 (2025). DOI: 10.1038/s41586-025-09215-4
What I would replicate: The paper's central claim is that Centaur, a language model fine-tuned on the Psych-101 dataset of 160 psychology experiments, predicts held-out human behaviour better than domain-specific cognitive models. All published analyses, and all published critiques, use the 70B version. The authors also released an 8B version, noting on its model card that it runs on Google Colab's free GPUs. Nobody appears to have checked whether the original findings, or the subsequent criticisms, hold for that smaller model, which is the only version accessible to researchers without significant compute. (To my knowledge)
I would write my own evaluation code from the paper's description and report whether the 8B model reproduces the headline result.
My question: Centaur has attracted several published critiques and re-analyses (Liu & Ding in National Science Open; "Large Language Models Do Not Simulate Human Psychology"; others). I'm not aware of any independent reimplementation of the original evaluation, but I don't want to start work if you'd consider the original research already replicated.
Would this be in scope, or is it already covered?
If the answer is no, I'd be grateful to know plainly.
Thank you for your time,
Anish Sugan, Independent
ORCID: 0009-0003-2773-5242
Dear editors,
I have a question before I start work. I've read #127, so I'll skip the parts already settled there. I understand that partial replication is fine, and that writing original code from the paper's description (consulting the original code only to fill gaps) is what makes it a replication rather than a reproduction. My question is only about prior coverage.
The original article is Binz et al., "A foundation model to predict and capture human cognition," 644, 1002–1009 (2025). DOI: 10.1038/s41586-025-09215-4
What I would replicate: The paper's central claim is that Centaur, a language model fine-tuned on the Psych-101 dataset of 160 psychology experiments, predicts held-out human behaviour better than domain-specific cognitive models. All published analyses, and all published critiques, use the 70B version. The authors also released an 8B version, noting on its model card that it runs on Google Colab's free GPUs. Nobody appears to have checked whether the original findings, or the subsequent criticisms, hold for that smaller model, which is the only version accessible to researchers without significant compute. (To my knowledge)
I would write my own evaluation code from the paper's description and report whether the 8B model reproduces the headline result.
My question: Centaur has attracted several published critiques and re-analyses (Liu & Ding in National Science Open; "Large Language Models Do Not Simulate Human Psychology"; others). I'm not aware of any independent reimplementation of the original evaluation, but I don't want to start work if you'd consider the original research already replicated.
Would this be in scope, or is it already covered?
If the answer is no, I'd be grateful to know plainly.
Thank you for your time,
Anish Sugan, Independent
ORCID: 0009-0003-2773-5242