Kerem Oktar

Postdoctoral Researcher,
Meta FAIR

PhD in Psychology, Princeton

Social Intelligence in Minds and Machines

I build computational models of how intelligent systems form beliefs, make decisions, and navigate social interactions, then use those insights to evaluate and improve AI.

At Meta FAIR, I am extending this work from evaluation toward post-training and multi-agent systems.

Selected Work

Muse Spark Safety & Preparedness Report Meta · 2026 · Sycophancy evaluation contributor

The report evaluates Meta's LLM, Muse Spark, across safety and behavioral dimensions; I contributed to the sycophancy benchmark, which measures inappropriate agreement and pushback.

Report · arXiv
Sycophancy rate plotted against excessive anti-sycophancy rate for Muse Spark and other frontier models.
How Beliefs Persist Amid Controversy: The Paths to Persistence Model Psychological Review · 2025 · Oktar & Lombrozo

We unite theories across the social sciences into four concrete reasons people resist disagreement—distrusting dissenters, treating issues as subjective, anticipating costs of changing, or lacking the cognitive resources to update—and test them across five preregistered studies.

Paper
Figure 10 from Paths to Persistence: nested model comparisons showing the predictive value of the model's paths and interactions.

Full list: CV · Google Scholar.

Recent Research

HorizonBench Can models track a user as their preferences change over months? DetailsClose
Preprint · 2026 · Li, Paranjape, Oktar et al.

HorizonBench generates six-month conversations from structured mental-state graphs, with ground-truth provenance for every preference change; across 25 frontier models, most score at or below chance, with updated-state tracking the central bottleneck. Paper · Code · Data

Figure 1 from HorizonBench: the pipeline for constructing conversations with evolving user preferences and counterfactual responses.
Brittlebench Do benchmark conclusions survive harmless changes in wording? DetailsClose
Preprint · 2026 · Romanou, Ibrahim, Ross, Oktar et al.

Brittlebench introduces a variance-decomposition framework that separates task difficulty from prompt sensitivity; semantics-preserving perturbations change model rankings in 63% of cases and can explain up to half of performance variance. Paper

Figure 1 from Brittlebench: the meta-evaluation framework and performance variability under prompt perturbations.
Intuitive Theories of Truth How do people decide what can be true, is true, and should be asserted? DetailsClose
Trends in Cognitive Sciences · In press · Equal contribution · Oktar, Handley-Miner et al.

Intuitive Theories of Truth organizes everyday truth judgments around aptness, judgment, and assertion, explaining how people can disagree about truth even when they share the same evidence. Paper

Figure 1 from Intuitive Theories of Truth: two people apply different standards of aptness, judgment, and assertion to the same statement.

About

I was born and raised in Istanbul (🧿); studied economics and cognitive science at Pomona College, CA (☀️); completed my PhD in Psychology at Princeton, NJ (❄️); and now research social cognition in AI systems at Meta FAIR in Seattle, WA (☔).

My research has won CogSci’s Marr Prize for best paper and SPP’s Poster Prize; Princeton’s Center for Human Values supported my PhD through a Rockefeller Prize Fellowship. Beyond research, I co-designed and co-taught Psychology of Justice at Edna Mahan, a women’s prison in New Jersey, in 2024. I have also found and tamed the Oscar Mayer Wienermobile.

Contact

Feel free to contact me at oktar[dot]research[at]gmail.com with regards to research / collaboration / mentorship /… - I love talking about science.

Here is an anonymous feedback form.

Writing

On the Reasonable Effectiveness of Judgment: Why human judgment matters to mathematics and model training.

Why False Premises Birth Truths: An intuition for material implication.

More writing: Power analysis guide · Grad school application guide