Evaluating Large Language Models in Theory of Mind Tasks
Has this study been replicated?
The atlas records 1 replication of this study. Recorded outcomes: 1 failed.
Replications
- Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks, Ullman (2023). Outcome recorded: failed.
we show that small variations that maintain the principles of ToM turn the results on their head.
View paper
Other studies in the atlas
Failed replications · Successful replications · All browse pages