Peter Shaw

Staff Research Scientist at Google DeepMind

Google Scholar · Twitter · Bluesky · LinkedIn

Current Work

I work on Gemini post-training, focusing on improving capabilities related to coding, research, and various scientific workflows. My work spans evaluations, data, and infrastructure for agents and reinforcement learning.

I'm also interested in foundational questions related to training recipes and generalization, with implications for safety and reliability. I'm especially motivated by the potential of AI to help address key challenges in biology, medicine, and scientific discovery more broadly.


Published Work

See Google Scholar for a complete list of publications.

Effective Reasoning Chains Reduce Intrinsic Dimensionality
Archiki Prasad, Mandar Joshi, Kenton Lee, Mohit Bansal, Peter Shaw
ICML 2026 (Spotlight)

Bridging Kolmogorov Complexity and Deep Learning: Asymptotically Optimal Description Length Objectives for Transformers
Peter Shaw, James Cohan, Jacob Eisenstein, Kristina Toutanova
ICLR 2026

ALTA: Compiler-Based Analysis of Transformers
Peter Shaw, James Cohan, Jacob Eisenstein, Kenton Lee, Jonathan Berant, Kristina Toutanova
TMLR 2025 · Code

ProtEx: A Retrieval-Augmented Approach for Protein Function Prediction
Peter Shaw, Bhaskar Gurram, David Belanger, Andreea Gane, Maxwell L Bileschi, Lucy J Colwell, Kristina Toutanova, Ankur P Parikh
MLCB 2025 (Oral) · Code

AgentRewardBench: Evaluating Automatic Evaluations of Web Agent Trajectories
Xing Han Lù, Amirhossein Kazemnejad, Nicholas Meade, Arkil Patel, Dongchan Shin, Alejandra Zambrano, Karolina Stańczak, Peter Shaw, Christopher J. Pal, Siva Reddy
COLM 2025 · Code

Robust Preference Optimization through Reward Model Distillation
Adam Fisch, Jacob Eisenstein, Vicky Zayats, Alekh Agarwal, Ahmad Beirami, Chirag Nagpal, Peter Shaw, Jonathan Berant
TMLR 2025

BAGEL: Bootstrapping Agents by Guiding Exploration with Language
Shikhar Murty, Christopher D. Manning, Peter Shaw, Mandar Joshi, Kenton Lee
ICML 2024

Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking
Jacob Eisenstein, Chirag Nagpal, Alekh Agarwal, Ahmad Beirami, Alex D'Amour, DJ Dvijotham, Adam Fisch, Katherine Heller, Stephen Pfohl, Deepak Ramachandran, Peter Shaw, Jonathan Berant
COLM 2024

From Pixels to UI Actions: Learning to Follow Instructions via Graphical User Interfaces
Peter Shaw*, Mandar Joshi*, James Cohan, Jonathan Berant, Panupong Pasupat, Hexiang Hu, Urvashi Khandelwal, Kenton Lee, Kristina Toutanova
NeurIPS 2023 (Spotlight) · Code

QUEST: A Retrieval Dataset of Entity-Seeking Queries with Implicit Set Operations
Chaitanya Malaviya, Peter Shaw, Ming-Wei Chang, Kenton Lee, Kristina Toutanova
ACL 2023 · Outstanding Paper Award · Code · Dataset

Pix2Struct: Screenshot Parsing as Pretraining for Visual Language Understanding
Kenton Lee*, Mandar Joshi*, Iulia Turc, Hexiang Hu, Fangyu Liu, Julian Eisenschlos, Urvashi Khandelwal, Peter Shaw, Ming-Wei Chang, Kristina Toutanova
ICML 2023 · Code

Evaluating the Impact of Model Scale for Compositional Generalization in Semantic Parsing
Linlu Qiu, Peter Shaw, Panupong Pasupat, Tianze Shi, Jonathan Herzig, Emily Pitler, Fei Sha, Kristina Toutanova
EMNLP 2022

Improving Compositional Generalization with Latent Structure and Data Augmentation
Linlu Qiu*, Peter Shaw*, Panupong Pasupat, Paweł Krzysztof Nowak, Tal Linzen, Fei Sha, Kristina Toutanova
NAACL 2022 · Code

Compositional Generalization and Natural Language Variation: Can a Semantic Parsing Approach Handle Both?
Peter Shaw, Ming-Wei Chang, Panupong Pasupat, Kristina Toutanova
ACL 2021 · Code

Exploring Unexplored Generalization Challenges for Cross-Database Semantic Parsing
Alane Suhr, Ming-Wei Chang, Peter Shaw, Kenton Lee
ACL 2020 · Code

Generating Logical Forms from Graph Representations of Text and Entities
Peter Shaw, Philip Massey, Angelica Chen, Francesco Piccinno, Yasemin Altun
ACL 2019

Self-Attention with Relative Position Representations
Peter Shaw, Jakob Uszkoreit, Ashish Vaswani
NAACL 2018 · Code

* Indicates equal contribution.